Physical AI Brief
Daily cross-source signals for the Physical AI supply chain — silicon photonics, CPO, VLA models, humanoid hardware, embodied AI. Three streams, one page, zero filler.
138 items today · 138 arxiv · 0 SEC 8-K · 0 humanoid · 0 CN photonics
01 ARXIV · PHYSICAL AI PAPERS
138 items- arxiv:2609.23252 · cs.RORobot World Models Are Not Invariant to How the Actions Are WrittenAhmed Karim, Leon Chlon
A robot policy is trained with one of two action parameterizations: absolute joint targets, or deltas relative to the current state. The choice is a live engineering decision in robot learning, and a world model conditioned on actions inherits it silently. We show the inheritance is catastrophic. A latent dynamics model trained on one parameterization and handed the identical commanded trajectory written in the other collapses: retrieval degrades by 2.6-13.4x across three robot datasets and two morphologies, goal-conditioned action selection falls from 53% to 15%, and on PushT the two beliefs about the same future are near-orthogonal (cos = 0.067, worst case -0.377), so the predictor does not degrade gracefully, it answers a different question. This is not a distribution-shift artifact in the usual sense: the two encodings are mutually reconstructible at R^2 = 0.996 given the joint input, so no information is lost, and we give the test that separates a valid re-parameterization from a lossy summary or a sensor swap. The test rejected three of the four axes we proposed. The defect lives in the action channel, which the invariance literature for visual models does not examine: work there concerns crops, jitter and camera pose, while the parameterization of the commands goes unaudited. The repair is averaging over the two encodings, and where it goes matters. Averaging the objective restores task performance by itself; averaging the outputs, safe for probabilities by concavity, is not available for direction-valued prediction, where the normalized mean can score below every member of the orbit. What objective-averaging leaves behind is the tail: worst-case agreement stays at 0.78, a disagreement penalty closes it to 0.995, and over a latent rollout it is the difference between a worst case that erodes and one that holds. On PushT, averaging alone does not repair the axis.
robot policyworld modellatent dynamics - arxiv:2609.23231 · cs.CLChemCLIR-Bench: Benchmarking Cross-Lingual Information Retrieval in Multilingual Chemical PatentsMahdi Astaraki, Mohammad Khodadad, Reza Namazi, Mohammad Arshi Saloot +3
Cross-lingual information retrieval (CLIR) is increasingly important in multi-national industries, where critical technical evidence may exist in a different language than the query. However, existing benchmarks do not adequately capture domain-specific cross-lingual retrieval or the retrieval-depth and recoverability failures that aggregate recall hides. In this work, we benchmark CLIR in the chemical domain, with a focus on patent data. We construct a multilingual dataset from Google Patents and the European Patent Office (EPO) data, spanning five languages (covering major Eastern and Western languages) and reflecting the diversity and complexity of real-world industrial documentation. Using this dataset, we systematically evaluate eight state-of-the-art embedding models for cross-lingual retrieval. Our results show a substantial performance gap between monolingual and cross-lingual settings: for the best-performing model, Recall@10 drops from 0.72 to 0.53 in cross-lingual setting. Retrieval depth also degrades significantly, with relevant documents ranked lower across languages in cross-lingual scenarios. Furthermore, some multilingual embedding models that perform strongly in monolingual settings exhibit sharp declines when queries and documents are in different languages, providing practical insights for model selection in cross-lingual use cases. These findings highlight critical limitations of current approaches and emphasize the need for more robust cross-lingual retrieval methods in domain-specific settings. Our benchmark provides actionable insights for model selection and establishes a controlled diagnostic evaluation framework for CLIR over industrial technical text. Data and code are publicly available at https://github.com/MohammadKhodadad/Multi-Lingual-QAC.
benchmarkevaluation framework - arxiv:2609.23215 · cs.LGTriggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference AgentsParam Raval, Rohit Shenoy, Archana Vaidheeswaran
LLM explainers are increasingly attached to autonomous agents as runtime oversight, with operators reading a generated account of the agent's beliefs and actions rather than its internal state. We audit the account itself, pairing an Active Inference (AIF) agent that tracks German grid demand and adjusts generation with an LLM explainer on three backends (GPT-4o, Claude-3-Opus, Gemini), and probing the pair with three black-box triggers. Corrupting the observation stream by 600 MW per step moves the agent's posterior by 490 MW, roughly 0.9% of grid capacity. None of the 30 explanations produced during the injection flag anything under a stated rubric, and each narrates the corrupted belief fluently. On timesteps where the agent takes an objectively wrong action, all three explainers produce a sycophantic rationalization 80-95% of the time (n = 20 per backend). Attacker-controlled text in the observation metadata field steers the explainer, with susceptibility differing by provider and data exfiltration succeeding on all three. We propose mitigations for each failure but do not evaluate them. In every failure we observed, the explanation was fluent and wrong. Moreover, nothing in the explainer architecture checks whether an explanation is true before an operator acts on it. Testing the explainer therefore belongs in any audit of an agentic deployment.
agentautonomous agentagentic - arxiv:2609.23206 · cs.LGSDC-GON: Singular Decomposition and Consistency-Regularized Green's Operator Networks for Solving Partial Differential EquationsYingchao Huang, Xin Wang, Shanshan Yao, Fanhua Zeng +1
Green's function based operator approximation offers an efficient route for solving linear partial differential equations under varying boundary conditions and source terms. Once the Green's function is learned, solutions for new configurations are obtained through integration rather than by solving the differential equation again. Existing Green's function learning methods face two structural challenges. The first is the singular behavior of the Green's function near the source point, which places a difficult approximation burden on neural networks. The second is the absence of explicit consistency between the learned Green's function and its gradient, although both quantities enter the integral solution representation directly. This work proposes SDC-GON, a Singular Decomposition and Consistency-Regularized Green's Operator Network that addresses both challenges within a unified framework. The Green's function is decomposed into an analytically known singular component and a smooth correction learned by the network, so that the neural approximation targets only the regular part of the response kernel. A self-consistency loss enforces agreement between the gradient and the autodifferentiation gradient of the smooth correction. The method is evaluated on two dimensional Poisson, three dimensional heat conduction, heterogeneous reaction diffusion, and Stokes benchmarks, consistently outperforming the compared baselines across all cases. On the heterogeneous pipe benchmark, SDC-GON achieves a testing error of $3.70\times10^{-4}$ with a smaller network architecture, compared with $9.60\times10^{-4}$ for the same-width baseline and $4.63\times10^{-4}$ for a larger configuration, demonstrating that structural improvements are more effective than increasing model size.
benchmark - arxiv:2609.23201 · cs.AIDo Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific EvaluationDanial Amin
Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems; commercial incentives and dependencies in external evaluation; benchmark saturation, defective tests, and data contamination; models exploiting scoring procedures; and the limited relevance of general scores to users' tasks. Documented cases illustrate why these problems require different responses. I argue for evaluation procedures that disclose the tested configuration, validate questions and successful task completion, report performance alongside cost and execution time, and make the scope of generalization explicit. I then discuss \textbf{Isotanta}, a crowdsourced benchmarking platform, as a practical example of contributed questions and repeated evaluation. A larger question pool may improve task coverage, while repeated sampling can improve the stability of estimates on that pool; neither guarantees validity or personalization. The paper distinguishes the platform's current shared ranking from proposed task-specific and user-provided evaluations. Its central argument is that model selection requires evidence about performance on the intended work, not simply a high position on a general leaderboard.
benchmarkleaderboard - arxiv:2609.23194 · cs.CLEnhancing speech representation learning with cross-modal knowledge transfer with HGNN under low resource settings: the case study of YembaYannick Yomie Nzeuhang, Paulin Melatagia Yonta, Marie Tahon
Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiveness of traditional and self-supervised methods. As a promising alternative, in this work, we propose to enhance acoustic representation trough a cross-modal transfer knowledge approach, based on heterogeneous graph neural networks (HGNNs), where acoustic and linguistic entities are modeled as distinct node types within a unified graph. Through message-passing mechanisms, linguistic nodes explicitly transfer knowledge to acoustic nodes, enabling structured and interpretable cross-modal information flow. To highlight this knowledge transfer and its benefits, we measured standard clustering metrics as an intrinsic evaluation of acoustic representation, and to emphasize applicability, we performed isolated-word recognition tasks using an English benchmark and a Cameroonian language dataset in low resources settings . Results demonstrate that acoustic representations consistently benefit from linguistic knowledge propagated through the graph. To our knowledge, this is the first demonstration of explicit cross-modal knowledge transfer for acoustic representation learning using HGNNs, highlighting a promising direction for speech representation in low-resource settings.
benchmark - arxiv:2609.23184 · cs.CVCausalWM: Causal Chain-of-Thought Reasoning for Embodied World ModelZiming Xu, Shuang Liang, Ruobing Han, Ziqiao Xi +9
Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly entangled within latent representations. We introduce CausalWM, a 16B embodied world model that performs explicit causal chain-of-thought reasoning before future video prediction. CausalWM organizes useful variables into a reasoning trajectory, allowing the model to progressively capture causal dependencies underlying physical evolution. To train CausalWM, we collect 31K hours embodied data and develop a three-stage paradigm consisting of large-scale video pre-training, causal CoT mid-training, and multi-objective RL post-training. Despite using only a limited set of supervised CoT variables, CausalWM exhibits emergent in-context learning capabilities, enabling contextual visual feature guidance and efficient few-step generation. CausalWM achieves state-of-the-art performance across language-conditioned, action-conditioned, single-view and multi-view benchmarks, including Top-1 performance on TriWorldBench leaderboard.
embodiedworld modelaction-conditionedpost-trainingbenchmarkleaderboard - arxiv:2609.23178 · cs.CLChronologic: Measuring Language Models' Ability to Represent the PastTed Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson +3
Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct answers. We use historical texts to develop a benchmark for a model's representation of English-language contexts 1831-1930, relying on pairwise comparisons to multiple ground truths and strong distractors to score the hardest questions in an appropriately graduated way. We find that generative tasks are harder than discriminative ones; in fact, reasoning models can typically discern the weakness of their own generated answers. While models pretrained exclusively on historical text lead the pack when evaluated by answer likelihood, they cannot compete with commercial models in free generation. None of the models we tested represent historical contexts in a fully persuasive way yet, but progress toward that goal is evident.
benchmark - arxiv:2609.23171 · physics.opticsDeterministic synthesis and processing of frequency-bin qubits in a macroscopically coherent quantum memoryS. A. Moiseev
Quantum information processing requires efficient storage and manipulation of photonic states.Here, we advance a cavity-assisted quantum memory protocol based on Pre-created Long-lived Macroscopic (PLM) coherence, thereby transforming quantum memory from a passive storage device into a platform for deterministic in-memory photonic processing. It is demonstrated that spin PLM coherence enables all-optical control of quantum memory using robust radio-frequency rotations alone. Here, the spin coherence serves as a distributed coherent quantum bus, which mediates the deterministic synthesis and storage of frequency-bin states through spectrally nonlocal impedance matching. We identify an inherent time-reversal symmetry in the underlying dynamical equations, ensuring unitary evolution of the quantum memory operations. The presented results demonstrate a unified platform for quantum storage and optical signal processing, enabling controlled manipulation of frequency-bin photonic qubits and paving the way for scalable multimode quantum networks based on spectral-encoding architectures.
manipulationmemory - arxiv:2609.23169 · cs.CVUltraTex: Unleashing 2K Multi-View Diffusion for 3D TexturingYibo Zhang, Ze Yuan, Nan Cao, Li Zhang +3
High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view diffusion methods have shown promising results for image-guided 3D texturing, but they are typically constrained to low operating resolutions such as 512 or 768, making it difficult to preserve high-frequency details from high-resolution reference images. Scaling this paradigm to 2048 resolution is computationally prohibitive, as the unified multi-view sequence exceeds 212K tokens and incurs excessive memory and latency. In this paper, we present UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing. Our key observation is that object-centric multi-view renderings contain two major sources of redundancy: background-induced sequence redundancy and sparse token interactions within the foreground. To address them, we introduce Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence. To enable efficient foreground-only inference while avoiding reconstruction artifacts, we further design Foreground-Aware VAE Decoding to ensure the quality of the final high-resolution views. To satisfy the demanding data requirements of 2K-resolution multi-view diffusion training, we construct G-buffer TexVerse, a large-scale, ultra-high-resolution multi-view rendering dataset covering over 268,000 3D assets. Extensive experiments show that UltraTex generates visually faithful textures with rich fine-grained details, while substantially improving efficiency, achieving $20.6\times$--$91.1\times$ training speedup and $22.3\times$--$74.6\times$ end-to-end inference speedup over the baseline on common samples in our dataset. Code and data is at https://yiboz2001.github.io/UltraTex.
memory - arxiv:2609.23164 · cs.LGSignal-Informed Temporal Routing for Vinyl Defect Regime DetectionYi-Hung Kan, Homayoon Beigi
Vinyl restoration systems must distinguish isolated clicks, short bursts, dense crackle, and overlapping damage before selecting a repair operation. We present a lightweight two-stage detector in which signal-informed sparse, burst, and dense experts produce complementary defect evidence, and a temporal backend converts that evidence into stable repair regimes. The backend factors the five-way decision hierarchically, applies a validation-only mixed regime gate, and decodes with validation-selected transition penalties that outperform a maximum-likelihood transition matrix on the same emissions. On a source-separated synthetic benchmark of 597 non-overlapping 15 s excerpts drawn from 21 recordings, the held-out system reaches 0.730 pooled five-class macro F1 (95% CI 0.694-0.761), an improvement of 0.043 over a flat frame-level router (paired bootstrap p=0.001). Clean and dense frames are highly reliable, while sparse, burst, and mixed frames remain far more ambiguous.
benchmark - arxiv:2609.23151 · cs.ROTask aware Dynamic Movement Primitives for failure detection and recovery in contact rich manipulationBhavnashri A, Sobia Shafi, Krishnapuram Himavarshini, Anuj Tiwari
Assembly remains a challenging robotic manipulation task in presence of tight tolerances and complex contact interactions. While Learning from Demonstration (LfD) frameworks like Dynamic Movement Primitives (DMPs) can effectively encode trajectories from a single demonstration, they are highly sensitive to variations in initial grasp configurations and external contact forces. Such variations often lead to task failures during the contact rich phases. This paper presents a task aware failure detection and recovery framework that integrates DMP based trajectory generation with real time stage classification. Utilizing Quadratic Discriminant Analysis (QDA) trained on multimodal sensor data, the framework segments execution into approach, alignment, and insertion stages for a Peg in Hole (PiH) assembly operation. By using goal relative position data as features, this classification generalizes to unseen goal positions without requiring retraining, matching the inherent generalization capability of DMPs. Anomaly detection is performed online using a Mahalanobis distance metric computed over force features, isolating contact induced failures from nominal trajectory execution. Upon failure detection, a spiral search recovery policy is triggered to actively realign the peg under contact before resuming the learned DMP insertion. The proposed approach is evaluated on an experimental setup achieving 95% stage classification accuracy, and demonstrates reliable failure recovery under lateral misalignments of up to 3 mm using only a single demonstration.
manipulationgrasp - arxiv:2609.23144 · cs.ROProbabilistic Scene Graphs: Hierarchical Representation and Real-time SystemWaqas Ali, Michele Antonazzi, Timon Homberger, Thien-Minh Nguyen +3
3D scene graphs provide semantically rich and hierarchical representations for robot perception. However, existing systems do not maintain uncertainty as an explicit belief or propagate it through the operations that construct and refine the graph. We introduce Probabilistic Scene Graph (PSG), a generalization of the conventional scene graph that represents a posterior over possible graphs, factorized into a discrete graph structure of entities, relations, and semantic attributes, and continuous states that ground them spatially, with uncertainty maintained over both components. Geometry is carried directly by the nodes rather than selected from a separately constructed metric map, so a metric map, where needed, follows from the graph rather than preceding it. We instantiate PSG's probabilistic spatial grounding with hierarchical graphs of Gaussians (HGG): each object primitive is represented by a full-covariance Gaussian under a Normal-Inverse-Wishart belief, and the same parametrization applied recursively within a node yields a geometry graph that resolves its surface at finer resolution. We then build a mapping pipeline that preserves these beliefs throughout graph construction and refinement: a purely graph-based coarse-to-fine alignment registers observations by comparing node beliefs, while a nested Expectation-Maximization and factor-graph optimization jointly refines poses, object parameters, and internal geometry. Across six datasets spanning indoor RGB-D, outdoor LiDAR, and cross-modality deployment, HGG operates at sensor rate with near-constant memory and achieves state-of-the-art object accuracy and zero-shot graph alignment.
memoryscene graph - arxiv:2609.23142 · cs.AICraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal EngineShutong Wu, Kevin Calderone, Andy Tsen
Building gameplay features in a game engine requires more than code, as code that compiles and runs does not necessarily implement the requested gameplay. We introduce CraftBenchUE, an evaluation harness that runs agents in an isolated Unreal Engine environment, reconstructs their saved submissions in fresh projects, and applies deterministic build, asset, and runtime checks without an LLM judge. Based on the harness, we built a benchmark consisting of 70 tasks spanning C++ source, Blueprint assets, and editor scripting. We evaluate seven models under two editor-tool configurations, with a file-and-shell baseline on C++ tasks. We further pair tasks that specify the same gameplay and use the same runtime tests, but require C++ and Blueprint as the deliverables. Across the 10 paired tasks, C++ completion rates exceed Blueprint by 30.0 and 42.9 percentage points in the two tool configurations. Among on-time Blueprint submissions in this paired set that pass asset checks, 42.2% and 50.0% fail explicit runtime assertions. These submissions satisfy asset requirements but fail the required gameplay tests. We will release the harness, task benchmark, and our trajectory findings with the report.
benchmark - arxiv:2609.23137 · cs.MATRACS: A Geometry-Aware Framework for Scalable Multi-Agent Path Finding in WarehousesSiddhant Erande, Anuj Tiwari
Large scale warehouse automation relies on efficient multi agent path finding (MAPF) to coordinate thousands of robots in structured environments. Existing MAPF algorithms primarily improve conflict resolution while representing warehouses as generic navigation graphs, overlooking their inherent geometric structure and traffic patterns. This paper presents TRACS (Traffic aware Routing and Aisle Coordination System), a geometry aware planning framework that exploits warehouse layout to simplify planning rather than introducing another conflict-resolution algorithm. TRACS constructs a directed routing graph with alternating one way aisles that eliminates head on and edge swap conflicts by design, decoupling spatial routing from temporal traffic coordination. Independent hybrid graph grid routing is combined with lightweight edge based scheduling to avoid joint space time search while ensuring collision free execution. Experimental evaluation on warehouse benchmarks against representative priority based, iterative repair, and search based MAPF planners shows that TRACS consistently achieves a 100% empirical success rate while substantially improving planning scalability. On fixed scene benchmarks with up to 1000 robots, TRACS reduces planning time by up to 14.7X while maintaining competitive makespan, lower flowtime, and near optimal path quality. Under a fixed 10 minute planning budget, TRACS routes up to 5120 robots, roughly twice the largest fleet reached by the strongest baselines, while sustaining a 100% success rate, demonstrating the effectiveness of exploiting warehouse geometry for scalable robotic warehouse systems.
agentmulti-agentbenchmark - arxiv:2609.23133 · cs.ROAquaCap: A Training-Free Underwater Embodied Agent with Code-as-PolicyXiaoshi Li, Yule Xu, Chunghiu Kong, Yizhou Zhou +4
Recent advances in vision-language-action models have stimulated growing interest in underwater embodied intelligence. However, their reliance on large-scale interaction data limits their applicability underwater, where data collection is costly and scarce. To address this challenge, we present AquaCap, a training-free Code-as-Policy framework for autonomous underwater navigation and manipulation. AquaCap employs a dual-layer agent that translates task instructions and environmental observations into condition-aware plans and executable control programs. Structured perception then provides the agent with semantic, geometric, and reliability-aware observations under degraded underwater conditions. A failure-aware memory diagnoses unsuccessful actions and supports closed-loop replanning and code revision. This design enables online adaptation without task-specific training or parameter updates. AquaCap achieves a 66.43% success rate in simulation. Real-world experiments further demonstrate autonomous grasping and object transport with an ROV, including the manipulation of targets displaced by hydrodynamic disturbances.
vision-language-actionembodiedmanipulationgraspmemoryagent - arxiv:2609.23132 · cs.ROTip Manipulation in Soft Everting Robots via Wall Retraction and Deployable FingersNelson Badillo Perez, Niccolo Pagliarani, Matteo Cianchetti, Robert D. Howe
Soft everting robots can traverse long, confined paths by continuously growing, yet active interaction remains limited to a single tool fixed at or near the tip, unable to be repositioned on demand and difficult to reconcile with the robot's soft body. We introduce a tip-manipulation and multi-tool deployment strategy for soft everting robots based on wall retraction, implemented with a base roller assembly that independently meters membrane flow in the outer wall while a tail spool regulates growth in the internal tail section. Coordinated wall and tail actuation decouples robot length from membrane-material position, enabling membrane-mounted devices to be transported, exposed, and repositioned at selected locations near the distal tip. We pair this capability with ultralight pleated inflatable fingers integrated into the membrane, fabricated from TPU-coated nylon with an internal airtight bladder. The fingers achieve large bending at low pressures (approximately 100 degrees at 50 kPa in high-pleat designs) and generate blocking forces up to 1.9 N, while remaining limp during transport. The system demonstrates adaptive grasping across diverse household objects (21 to 550 g; 16 to 200 mm), three-dimensional object manipulation and stacking, environmentally braced extension, distal camera panning for confined-space inspection, and controlled sequential payload delivery. These results enable embodied and reversible tip manipulation for soft growing robots in cluttered and tortuous environments.
embodiedmanipulationgrasp - arxiv:2609.23131 · cs.ROSelective Commitment for Language-Guided Object Retrieval under Partial ObservabilityWonhee Koh, Sushil Samuel Dinesh, Hansol Ko, Shinkyu Park +1
Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability.
graspmemory - arxiv:2609.23130 · cs.AIFrom Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM ServingTwinkll Sisodia
Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state, phase placement, heterogeneous accelerators, networking, autoscaling, reliability, and service-level objectives. This paper connects that transition across peer-reviewed systems research, open-source implementations, and documented production studies. It treats vLLM and llm-d as complementary layers: model-serving engines optimize execution through mechanisms such as PagedAttention, continuous batching, kernels, quantization, and parallelism, while an inference control plane can optimize where, when, and under what policy execution occurs across a fleet. The contribution is synthesis rather than a new benchmark; all reported performance and deployment results remain attributed to their original sources. The combined evidence suggests that the scarce resource in modern inference is shifting from raw FLOPs alone toward managed state, placement, network movement, reliability, and decision quality. We propose an Inference Execution Planner that selects feasible execution plans rather than only endpoints, including aggregated versus disaggregated topology, KV source and transfer action, hardware variant, routing/admission policy, and slower scaling decisions. We also provide a source-local benchmark atlas, a bottleneck-migration taxonomy, practical deployment guidance, an evaluation framework based on SLO-goodput, and research questions for agentic, multimodal, heterogeneous, and resilient inference.
agenticbenchmarkevaluation framework - arxiv:2609.23121 · cs.CVMM-ContextFold: Context Folding for Multimodal Agentic RetrievalYang Tian, Fan Liu, Jingyuan Zhang, Zhenyang Li +2
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.
agentagenticbenchmark - arxiv:2609.23118 · cs.ROVerti-WM: A Physics-Aided Exteroceptive World Model for Off-Road Reinforcement LearningChenhui Pan, Tong Xu, Xuesu Xiao
Reinforcement learning for off-road navigation requires extensive vehicle-terrain interaction data, which are costly to collect in high-fidelity simulation. World models offer a promising alternative by replacing simulator roll-outs during policy optimization. However, an off-road world model must condition state transitions on exteroceptive terrain information, which proprioception alone does not provide. This challenge is further amplified by the need to model both rigid and deformable terrain, where data-driven and physics-based approaches offer complementary strengths. We propose Verti-WM, a physics-aided exteroceptive world model that recurrently fuses a frozen Transformer for rigid terrain and a neuro-symbolic terramechanics model for deformable terrain. Elevation and semantic observations queried from a supplied map at each predicted pose condition fusion, enabling six-degree-of-freedom rollouts for policy optimization without further simulator access. Verti-WM reduces prediction error by 34.6% and 21.7% over data-driven and physics-based baselines, respectively. Policies trained entirely within Verti-WM achieve comparable task success rates while reducing computation time by 23.6X relative to direct training in the high-fidelity simulator. We further validate Verti-WM using real-world data, enabling policy optimization within learned real-world kinodynamics and achieving a 80% success rate on the Verti-4-Wheeler platform, compared with 40% for direct sim-to-real transfer.
sim-to-realworld model - arxiv:2609.23113 · cs.ROSearch, Ground, Plan: Functional Sufficiency for Task and Motion Planning under Incomplete Scene KnowledgeNarendhiran Vijayakumar, Nav Singhal, Girish Varma, Antony Thomas
Foundation models (FMs) have expanded task and motion planning (TAMP) to manipulation problems specified through language and visual observations. However, incomplete scene knowledge leaves a critical gap between understanding what the task requires and knowing whether the physical scene can actually realize it. We introduce GRAB-TAMP, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency. We represent the task through functional roles, relations, and assignment constraints, and incrementally inspect the scene while requirements remain unresolved, verifying candidate objects through semantic, geometric, and relational checks. We evaluate GRAB-TAMP across 32 scene variants spanning Kitchen, Living Room, and Workshop domains. Across 200 feasible trials, our approach achieves 54.0% end-to-end success with 67.3% plan goal coverage. Compared with three FM-based TAMP frameworks under the same execution setting, GRAB-TAMP improves end-to-end success by 25.7 percentage points over the mean baseline. Implementation and evaluation code: https://github.com/Narendhiranv04/GRAB-TAMP
manipulation - arxiv:2609.23103 · cs.RODiagGen: Agentic Generation of Deformable Assets with Sim-based Diagnostics for Robotic SimulationGuanxiong Chen, Yiduo Qu, Qianjun Xia, Pengyu Jing +11
While simulation-ready deformable assets are essential for in-silico robotic manipulation tasks, existing generation frameworks typically assess physical plausibility after generation, leaving an object's simulated response unused as feedback for repairing upstream errors. We present DiagGen, an agentic framework that turns a single in-the-wild image into a simulation-ready deformable asset through a generate--simulate--diagnose--refine loop. DiagGen constructs part-aware geometry and material parameters, then uses a VLM (vision-language model)-based agent to select semantically informative regions, probe them in a physics simulator, observe material responses, and route evidence-backed repair cues to the responsible generation stage. Experiments on 40 assets show that diagnostics provides useful repair cues and can moderately improve the quality of generated deformable assets. Finally, we show that unlike assets generated from visual foundation models which may not be simulatable, DiagGen-generated deformables can be directly dropped into a high-fidelity physical simulator for the planning and simulation of contact-rich pick-and-place tasks. The project's website is https://diaggen.github.io/.
manipulationagentagentic - arxiv:2609.23100 · cs.ROSplat-CBF: Safe Next-Best-View Control in 3D Gaussian-Splat MapsAmirhossein Mollaei Khass, Athanasios Cosse, Nader Motee
Where to look and how to move? A robot navigating an unmapped environment must do both at once, and the two goals pull against each other. The regions most worth observing are the ones the map knows least about, and those are exactly where the robot cannot trust its collision margins. We resolve this tension by introducing Splat-CBF, an active perception control barrier function that steers the camera toward the next best view while collision avoidance is enforced as a hard constraint. Safety is enforced by a risk-aware control barrier function that turns the Average Value-at-Risk of the Gaussian field into a single smooth hard constraint. Perception is enforced by a second barrier that rewards camera orientations with high expected Fisher information gain near the robot's planned path. The two meet in a quadratic program where safety is hard and perception is soft, with a slack penalty that adapts to how often perception has already been relaxed and how close the robot is to an uncertain region. We verify the method in indoor simulations, a Isaac Kinova manipulator and in experiments on an Ackermann-drive robot. Our results assert that robot navigates faster, gathers more information, and runs faster online than safety-only and perception-only baselines, giving up informative motion only when safety requires it.
manipulator - arxiv:2609.23088 · cs.CLOmniEdu: Open Foundation Models for Learning and TeachingHao Liang, Qihan Lin, Meiyi Qiang, Linzhuang Sun +4
Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capability. We present OmniEdu, an open family of foundation models for K-12 learning and teaching. Its instruction-tuning corpus combines over 100 educational resources and general instruction sources, organized around four capabilities: subject competence, curriculum grounding, diagnostic reasoning, and pedagogical action and scaffolding. Our pipeline integrates deterministic cleaning, semantic auditing and rewriting, task-specific quality scoring, token-budgeted diversity selection, and pedagogical instruction assignment. It yields 69,999 examples and 15.96M supervised response tokens, including 60,951 education-specific examples. We fine-tune 4B, 9B, and 27B models and evaluate curriculum grounding, K-12 problem solving, and pedagogical tutoring, alongside general capability. Education-oriented tuning consistently improves all three educational benchmark groups across model scales. OmniEdu-27B achieves 63.12% EM and 76.69% F1 on K12-Bench, 85.89% on MathFish, 86.95% on EDUMATH, and 78.74% in MathTutorBench's Scaffold setting. It also achieves the highest Teaching average on LongTutor among the evaluated models, at 3.02. These results demonstrate the value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support.
benchmark - arxiv:2609.23085 · cs.LGMeasured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM ServingMuhammad Abdur Rab Siddiqui, Daniela Rojas, Chen Yang, Wenqi Cui +2
Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajectories extend. While in practice, many queries do not require the capabilities of the largest available model, and routinely directing such queries to a high-capability model can introduce unnecessary, considerable computation and energy consumption. In this paper, we investigate whether adaptive routing across a heterogeneous pool of LLMs can reduce this energy burden without substantially compromising task performance. We design a language-model-based router that reads in each query and selects an answer model from a fixed candidate pool. The candidate models are first profiled through an offline tournament that records their correctness, latency, power, and GPU energy for each query. Using these measurements, the router is trained through supervised fine-tuning followed by group relative policy optimization (GRPO) with the tailored paradigms. Results demonstrate that learned routing can selectively allocate expensive model capacity based on query context and improve the accuracy-energy tradeoff in multi-LLM serving. Across seven benchmark tasks, we also observe a sharp accuracy-energy phase transition among routers, providing practical insights into improving energy efficiency while maintaining LLM performance.
agenticbenchmark - arxiv:2609.23077 · eess.SYPhysics-Informed Neural Network Surrogates with Polynomial Chaos-Based Uncertainty Propagation for Stochastic Model Predictive ControlSrimanta Santra, Romi Patel, Saikat Mukherjee, Steven L. Brunton +1
Stochastic partial differential equations (PDEs) govern critical engineering and geophysical systems but are challenging to use for real-time control under parametric uncertainty. We present a unified framework that couples Physics-Informed Neural Networks (PINNs) with Polynomial Chaos Expansion (PCE) to construct a fast and differentiable surrogate. The PCE representation enables analytical propagation of parametric uncertainty and computation of the corresponding moments without requiring Monte Carlo sampling. We provide an error decomposition for the PINN-PCE surrogate that separates PCE truncation, stochastic quadrature, and PINN approximation errors. Embedding this surrogate into a stochastic model predictive control (SMPC) scheme enables finite-horizon control updates based on analytic mean and covariance predictions. We further show how the surrogate approximation error can be incorporated into tightened probabilistic constraints. The approach is validated on three benchmarks: the Korteweg-de Vries equation, Burgers' equation, and the two-dimensional incompressible Navier-Stokes equations, representing dispersive, convective, and convective-diffusive dynamics. Across all cases, the surrogate enables real-time control updates while maintaining prescribed risk levels and closely matching the corresponding high-fidelity solvers at substantially lower computational cost.
benchmark - arxiv:2609.23073 · cs.LGMolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMsHyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.
benchmark - arxiv:2609.23067 · cs.CVLD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument SegmentationZan Wang, Yunhe Feng, Dong Nie, Oluwatosin Oluwadare +3
Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.
benchmark - arxiv:2609.23064 · cs.AIFireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire DynamicsQiang Chen, Hao Guo, Huatai Zhu, Tairan Huang +6
Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical systems involve multiple continuous fields, latent causal mechanisms, partial observations, and intervention-sensitive dynamics. We propose FireWorldBench, a benchmark for evaluating complex physical world intelligence in multimodal large language models and agents through coupled-field fire dynamics. Fire provides a canonical stress-test environment, where multiple interacting physical fields jointly shape observable states and temporal dynamics. FireWorldBench is organized along two complementary axes, a physical capability axis and a fire scenario task axis, jointly covering physical-state understanding, temporal dynamics, causal mechanisms, and intervention reasoning. The benchmark comprises 520 fire-world entries, including 494 controlled simulation worlds and 26 real-world-aligned event groups, spanning 47 scene archetypes across 7 environment families. These entries combine structured textual observations, multiple 2D physical-field visualizations, and 3D event-level scene modeling, yielding 9,074 text-image interleaved question-answer pairs across choice-based and open-ended report-generation formats. FireWorldBench evaluates whether models can infer latent physical states, explain underlying mechanisms, forecast coupled-field evolution, and assess intervention consequences from multimodal partial observations, providing a challenging testbed for complex physical world intelligence.
benchmark - arxiv:2609.23061 · cs.CVHDMamba-YOLO: Efficient State-Space Perception and Local Spatial Reconstruction for UAV Small ObjectLinduo Wei, Junjie Fan, Yijun Mai, Yong Qi
Small-object detection in UAV imagery is challenged by weak visual evidence, ambiguous boundaries, dense object distributions, and complex backgrounds. Effective detection therefore requires long-range contextual information for target-background discrimination while preserving explicit local two-dimensional structures for accurate localization. These requirements arise at different stages of the detection pipeline and are not naturally addressed by a uniform feature-processing strategy. We propose Hybrid Dual-domain Mamba-YOLO (HDMamba-YOLO), a stage-wise heterogeneous SSM-CNN detector organized according to a perception-reconstruction-alignment-interaction rationale. EfficientVMamba-based EVSS establishes long-range contextual perception in the backbone, while PhasePatchMerging2D provides phase-aware hierarchical transitions. DST-Wrapper and Native C3k2-ASSAF then perform perception-to-reconstruction transition and repeated local two-dimensional reconstruction during FPN/PAN aggregation. DySample provides content-adaptive cross-scale resampling, while OS-CVTIA introduces macro-micro interaction and task-specific modulation for localization and classification. On VisDrone2019, HDMamba-YOLO-B achieves 42.737% mAP50 and 25.713% mAP50:95 with 10.042M parameters and 29.879 corrected GFLOPs. HDMamba-YOLO-Lite achieves 41.140% mAP50 and 24.741% mAP50:95 with 5.344M parameters. Under the unified AI-TOD evaluation protocol, HDMamba-YOLO-B obtains 21.621% AP and 47.881% AP50. Controlled ablations further support the stage-wise allocation of state-space perception, convolutional reconstruction, dynamic alignment, and task interaction for UAV small-object detection.
evaluation protocol - arxiv:2609.23060 · cs.ROOn the Control of Mobile Ad-Hoc Agent Deployments in Partially Observed SpaceEdwin Meriaux, Louis-Roy Langevin, Shuo Wen, Ndiamé Ndiaye +2
We study the online deployment of mobile ad hoc networks in unknown orthogonal environments, formalized as the Partially Observable Cooperative Guard Art Gallery Problem. We give a full proof that CADENCE algorithms achieve full coverage while maintaining a connected visibility graph using at most $n/2 + h - 2$ agents in orthogonal worlds with $n$ corners and $h$ holes. We further evaluate deployment-order heuristics that reduce agent count and deployment time in practice.
agent - arxiv:2609.23058 · cs.AILazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic ProgramsXin Heng
Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around a live, goal-derived demanded set. LazyAgent refreshes a backward closure from requested outputs as execution state changes and materializes a ready node only when the active goal requires it. This replaces repeated local judgments with one linear-time graph analysis followed by constant-time membership tests, allowing programs to remain broad while execution stays request-specific. On programs that describe more than the current request needs, LazyAgent consistently outperforms the strongest goal-stopping eager baseline by refusing unrelated work before it starts. Adding one unrelated product raises the eager bill by 22.5% and LazyAgent's by 0.0%. LazyAgent saves 42.0% of measured CPU on production scientific workflows and 51.7% of container time on a live release gate spanning four repositories. We also prove and verify exact equivalence when the request reaches the whole graph, leaving no unrelated work to avoid. Beyond permission, goal-relative output projection saves up to approximately 90% of a shared step on two third-party test suites while the identical eager control saves 0.0%; the advantage disappears when the omitted output has no other consumer or the request needs it. Ordering, reuse, and pruning can also save cost, but do not replace permission. Finally, we show that current public benchmarks are eager-shaped and contain almost no unrequested work. A pre-registered planning intervention did not broaden them. These findings motivate benchmarks built from standing programs and sequences.
agentagenticbenchmark - arxiv:2609.23056 · cs.CLBridging Static and Agentic RAG for Taiwanese Historical Question AnsweringKai-Hsin Chen, Wei-Yu Chen, Xuanjun Chen, Jyh-Shing Roger Jang
Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear whether such adaptive orchestration consistently outperforms well-designed static pipelines. We conduct a controlled comparison of agentic and static RAG for Taiwanese historical question answering, sharing the same generator and hybrid retrieval backend. Despite similar aggregate performance, the two pipelines differ on 70.83% of questions, with their advantages largely canceling out when averaged. An oracle that selects the better response per question improves the composite score by 0.2417 over the better individual pipeline, revealing substantial headroom for question-level selection. We therefore introduce a post-hoc selector that compares the two responses and their cited evidence, significantly outperforming either individual pipeline and recovering 60.34% of the oracle headroom. These results show that aggregate comparisons can obscure meaningful question-level differences between retrieval strategies, suggesting that exploiting their complementarity may be more fruitful than seeking a universally superior pipeline.
retrieval-augmentedragagentic - arxiv:2609.23055 · cs.LGOptimizers for Diffusion Models: A Controlled BenchmarkArman Bolatov, Egor Shulgin, David Li, Abduragim Shtanchaev +5
Discrete diffusion models now match autoregressive language models on several benchmarks, while the question of how best to train them has received far less attention: the optimizer is inherited from one paper to the next and never compared. New optimizers, meanwhile, are validated almost exclusively on autoregressive pretraining, a different objective on a different loss surface. We present a controlled optimizer benchmark across four diffusion formulations, to our knowledge the first for discrete diffusion: seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) on masked diffusion (text8), uniform diffusion (QM9, and LM1B through the Gaussian duality) and Gaussian diffusion on images (CelebA-64), each on a task with published reference values. Every optimizer receives the same search protocol, and every winner is retrained at the full budget with three seeds. AdamW is a strong default but not always the right choice: it is beaten by a resolved margin on two of the four tasks, and the winner changes with the formulation, so the optimizer deserves the same care as the rest of the training recipe. Notably, methods validated on autoregressive language model pretraining transfer well: Muon, MARS-M and SOAP each beat the tuned AdamW on at least one diffusion formulation. The benchmark, all runs and every figure are reproducible end to end from the released code at https://github.com/armanbolatov/diffusion-baselines.
benchmark - arxiv:2609.23053 · cs.CLAttributable Post-Rationalization in RAG Citations: A Controlled Reproduction and an RLVR ComparisonMehedi Khan, Md. Shariful Islam Bhuyan
A RAG system can hand you the right answer and cite a source it did not actually use. Models output these unfaithful citations via post-rationalization: they write the answer first and then attach a citation to whatever passage looks close enough. Search agents are now trained with reinforcement learning from verifiable rewards (RLVR), which pays them for getting the answer right. We asked whether that training also teaches them to cite honestly. Improving an existing methodology with a required control, we compared an instruction-tuned model against three RLVR agents trained from it, on four question-answering datasets, using only free-tier Kaggle GPUs. Post-rationalization is everywhere: on Wikipedia-based questions roughly one citation in seven is unfaithful. RLVR does not fix it. The agents post-rationalize at their base model's rate, and one lands slightly worse. Rewarding correct answers buys nothing in citation faithfulness, so faithfulness has to be trained and measured on its own terms.
rag - arxiv:2609.23049 · cs.CVVDGS: Visibility-Driven Large-Scale 3D Gaussian Splatting for Aerial Scene ReconstructionHaolin Yu, Jiadong Tang, YiXian Wang, Yu Gao +4
Large-scale scene reconstruction is a critical foundational technology in robotic autonomous systems such as 3D mapping and autonomous driving. In recent years, 3D Gaussian Splatting (3DGS) has demonstrated remarkable advantages in both visual quality and computational efficiency, making it a promising representation for large-scale scene reconstruction. However, it still faces challenges in large-scale scenes, including excessive memory consumption and uneven viewpoint coverage caused by UAV acquisition, limiting its real-world applications. To address this, we propose VDGS, a novel 3DGS framework that incorporates camera distribution into scene modeling. VDGS introduces visibility-driven statistics for scene anchors to quantify supervision strength. These statistics are further leveraged for scene partitioning and for gradient compensation in under-optimized regions, thereby promoting balanced optimization across different regions. Extensive experiments on multiple large-scale aerial scene datasets demonstrate that, under imbalanced viewpoint distributions, VDGS consistently outperforms existing methods, while maintaining competitive performance in scenarios with more uniform view distributions.
memory - arxiv:2609.23048 · cs.ROAnatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA PolicyFengze Jia
Compressed manipulation policies can pass offline evaluation while failing in closed-loop execution; this dissociation is established in prior work and is not our claim. We contribute a causal anatomy of one naturally occurring case. An 8-layer distillation of Octo-Base retains 86% of parameters, passes every offline check we applied (0.996 and 1.000 teacher-ratios on the family's own validation metrics), and collapses in closed loop: 0/72 vs. the teacher's 40/72 on a simulated WidowX pick-and-place task. The collapse is structured, not diffuse: early task stages degrade gradually (the student moves the object at 90% of the teacher's rate and grasps at 55%), while transport-to-target fails categorically, at 0% in every training variant. Paired action-trace forensics isolate the signature: a negative, late-heavy $z$ residual, roughly 10x its post-repair magnitude, and persistent across the base distillation and both continuation branches. Four standard therapies fail under matched controls: continued training and in-domain offline data leave success at zero, even though the latter measurably improves marginal action statistics; command-level compensation recovers nothing at any offset, although the same perturbations degrade healthy policies; clamping the symptom in the command channel preserves grasping, yet success stays at floor. A minimal-pair intervention that substitutes half of the training stream with deployment-distribution teacher rollouts, with every other setting held fixed, restores parity with the teacher (18/36 vs. 17/36 held-out), eliminates that signature, and recovers a teacher-like perturbation-response profile. We claim existence, not universality. Operationally, offline gates, including a family's own validation metrics, are insufficient acceptance tests for compressed policies; a few dozen closed-loop trials sufficed to find what they missed.
vlavla policymanipulationgrasp - arxiv:2609.23038 · cs.LGSpatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical WorldKaixiang Yao, Xu Wang, Miao Pan, Hu Xiyue +6
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.
benchmark - arxiv:2609.23037 · cs.ROFast and Robust Temporal Logic Planning via ADMM-based Trajectory OptimizationLukas Pries, Joris Verhagen, Jon Arrizabalaga, Jana Tumova +2
We present a fast numerical method for safe continuous-time motion planning under Temporal Logic (TL) specifications. The method generates smooth continuous trajectories that remain collision-free while robustly satisfying temporal and logical task requirements. A central component of our method is the formulation of nonconvex safety and logic constraints as unions of convex sets where associated discrete decisions are encoded in a joint feasibility graph. This graph representation allows Euclidean projection onto the feasible set and proximal robustness maximization to be reformulated as shortest- and widest-path problems, respectively. Building on this structure, we develop a nonconvex splitting method based on the Alternating Direction Method of Multipliers (ADMM), which decouples smooth spatio-temporal trajectory optimization from nonsmooth discrete constraint handling within the optimization. The resulting algorithm exhibits reliable convergence across benchmarks and scales to large-scale motion-planning problems, providing a 4.7x average speedup over the state of the art on discrete and continuous-time logic problems.
benchmark - arxiv:2609.23023 · cs.AIPINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language ModelsMingyang Yu, Xu Yang, Jun Zhang, Xiaolong Wang +2
Physics-informed neural networks (PINNs) require coordinated choices over network representation, sampling, loss construction, and optimization, while effective configurations often vary substantially across partial differential equations (PDEs). Existing automated PINN design methods can search candidate configurations, but information revealed during actual training is still used mainly for evaluation rather than to improve subsequent design, leading to repeated trial-and-error and inefficient use of training budget. We propose PINNsForge, an LLM-driven evolutionary framework for execution-feedback-based automated PINN design. PINNsForge generates diverse candidate configurations from PDE-related prior knowledge, evaluates them through actual training, and feeds high-performing designs together with accumulated execution evidence back to the LLM. Guided by observed optimization behavior, the LLM then refines, recombines, and explores coupled PINN design components, forming a continual cycle of generation, execution, feedback, and evolution. Unlike one-shot search or evaluation-only feedback, PINNsForge progressively converts training experience into improved design decisions for the target PDE. Across 25 PDE benchmarks, PINNsForge achieves the lowest mean MSE on 24 tasks compared with RoPINN, PINNsFormer, and PINNsAgent. Ablation studies further confirm the importance of the PDE knowledge base, execution feedback, and evolutionary search: removing these components increases the mean MSE to 3.74$\times$, 12.10$\times$, and 10.10$\times$ that of the full PINNsForge, respectively.
benchmark - arxiv:2609.23012 · cs.CVCrowdCue: Specialist-Cue Conditioning for Vision-Language Crowd CountingMoshiur Farazi, Bekir Ciftler, Abdulhalim Dandoush, Reda Bendraou
Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene, yet their raw counting accuracy sits in the range of sub-million-parameter specialist regressors. The open question is whether auxiliary guidance from a pretrained specialist can lift them into useful territory, and through which channel that guidance is best routed. We evaluate Qwen2.5-VL-7B on four widely used crowd counting benchmarks (ShanghaiTech A and B, UCF-QNRF, NWPU-Crowd). Zero-shot prompting rarely produces a parseable count, so LoRA supervised fine-tuning establishes the baseline at overall MAE 81.64. Conditioning on a P2PNet-derived density heatmap as an auxiliary visual signal fails in every encoding we tested, and an adversarial-swap protocol shows the model reads the heatmap but applies it counterproductively. We propose CrowdCue, a family that supplies the same specialist's already-integrated integer count to the VLM as a discrete symbol. The text-channel variant reaches MAE 72.04. The visual-channel variant, which renders the integer as printed digits and supplies it as a second image, reaches MAE 62.65, the strongest result in this paper and well ahead of the cue-supplying specialist alone (84.45 on the same split). In the late-fusion VLM we study, the binding constraint is not the channel but the abstraction level at which the specialist signal is delivered.
benchmark - arxiv:2609.23005 · cs.CVCompressing 3D Gaussian Splatting via Cross-Representation PriorsYezheng Zhang, Huanxiong Liang, Chuqin Zhou, Guo Lu +1
3D Gaussian Splatting (3DGS) enables high-quality novel view synthesis but incurs high storage and transmission costs due to dense Gaussian primitives. Recent anchor-based compression reduces per-primitive redundancy, yet redundancy across anchors remains largely unexploited. We propose CRP-GS (Cross-Representation Priors for Gaussian Splatting), a rate-distortion optimized compression framework that leverages cross-representation priors to improve anchor-level entropy modeling. First, a Correspondence-Oriented Hierarchical Structure (COHS) organizes anchors by feature correspondence rather than spatial proximity, constructing root-leaf dependencies so that selected anchors can act as informative priors to conditionally encode others, yielding more accurate likelihood prediction and lower conditional entropy. Second, Shared Feature Aggregation (SFA) extracts globally shared features from a contextual hash grid and injects them into anchor representations, factoring out scene-consistent low-frequency information that would otherwise be redundantly embedded in individual anchors. Both modules are trained under a unified rate-distortion objective to balance bitrate reduction and rendering fidelity. Experiments across multiple benchmarks show that CRP-GS achieves a favorable overall rate-distortion trade-off, yielding around 30% average bitrate reduction compared to anchor-based baselines while maintaining comparable rendering quality.
benchmark - arxiv:2609.23003 · cs.ROM3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in ForestsEthan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes +1
We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests. M3GA-Wild unifies and extends existing forest localisation datasets, providing a holistic benchmark with synchronised RGB imagery and LiDAR from ground traversals spanning 36 km, aligned high-resolution aerial imagery and multi-altitude LiDAR covering 370 hectares, and accurate geo-referenced 6-DoF poses for precise evaluation. M3GA-Wild captures diverse forest scenes with varying viewpoints, occlusion, and environmental conditions, enabling systematic evaluation of visual, LiDAR, cross-modal, and multi-modal methods. Baseline experiments show that LiDAR-based approaches significantly outperform vision-only methods under severe viewpoint differences, while current multi-modal fusion strategies yield limited gains due to poor cross-modal alignment. By pairing aerial RGB imagery with geo-referenced aerial LiDAR, M3GA-Wild also enables evaluation of foundation models for monocular depth estimation as a cheap source of 3D geometry from forest imagery, with initial experiments revealing shortfalls of current methods. These results highlight key challenges in cross-platform localisation, including modality misalignment and severe domain gaps. M3GA-Wild establishes a new benchmark to support research in robust multi-modal localisation and long-term autonomy in unstructured natural environments. The dataset and code will be available upon acceptance.
benchmark - arxiv:2609.22991 · cs.LGSilent Failures Beyond the 32-Bit Index Range: A Differential Characterization of Large-Tensor Matrix Multiplication in PyTorch's MPS BackendJunichiro Niimi
Apple Silicon machines with large unified memory make it possible to hold large tensors on a desktop GPU. However, we found that PyTorch's Metal Performance Shaders (MPS) backend silently returns wrong results for batched matrix multiplication with more than $2^{32}$ elements. torch bmm, including its wrappers matmul and eager attention, returns relative errors above 1 without an exception or a warning in every PyTorch release tested (2.4.1 to 2.14.0). We sweep bmm over dtypes, memory layouts, shapes and batch sizes around $2^{31}$ and $2^{32}$ elements, and judge every result against a float64 computation on the CPU. Three rules account for every outcome on 2.14.0. When the output exceeds $2^{32}$ elements and an operand is a transposed view, the entire output is wrong and equals a computation that ignores that operand's strides. Otherwise, a view with at least $2^{31}$ elements raises an exception, and a contiguous input above $2^{32}$ elements makes exactly the batches beyond that point wrong, equal to a computation whose index wraps at $2^{32}$. A slightly larger problem can thus turn an explicit error into a silent failure. The rules extend to the backward pass, where a correct forward pass can return silently wrong gradients. A second machine with another chip, under two macOS versions, reproduces all 6156 results, including the wrong values, and the same sweeps on an NVIDIA A100 are correct in all 2530 runs. In a public sentiment classifier, one oversized batch corrupts a third of the outputs, which collapse onto one class. All findings come from observable behavior, without access to the backend's closed-source kernels; we release the harness, raw results and a guard that stops any MPS operation touching $2^{32}$ or more elements at jniimi/mps-silent-failures (https://github.com/jniimi/mps-silent-failures).
memory - arxiv:2609.22987 · cs.AIOptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization ModelingRuiqing Zhao, Rui Liu, Yuan Zuo, Huarong Zhang +2
Automated operations research (OR) modeling requires LLMs to translate natural-language decision problems into correct mathematical programs. Existing methods can improve individual formulations, but they often solve problems in isolation, retaining little reusable experience and repeating similar formulation errors. Prior memory-based approaches store examples, thoughts, or insights as references, while OR modeling requires reusable formulation skills that transfer across problem narratives and guide concrete modeling decisions. We propose OptiSkill, a skill-augmented framework that builds a hierarchical and evolving SkillBank for LLM-based OR modeling. SkillBank stores solver-verified experience as reusable skills, with Global Strategies for problem-level formulation skeletons and Step Experiences for local error-prevention rules. It is further refined through stable batch-level test-time evolution, where candidate skills are incorporated only after validation. Experiments on eight OR modeling benchmarks show that OptiSkill improves formulation accuracy across LLM backbones, outperforms strong agentic baselines, and gains further by expanding SkillBank coverage and reliability. Code and data are available at https://github.com/rachhhhing/OptiSkill
agenticbenchmark - arxiv:2609.22983 · cs.ROConnectivity-Aware Exploration of Robotic Grasp SpacesMaksim A Kazanskii
Robotic grasping is typically formulated as the problem of identifying successful actions from a space of candidate grasp poses. However, the organization of successful actions within this space has received less attention. We study the multiscale structure of viable robotic grasps in $SE(3)$ and investigate whether this structure can be exploited for more efficient exploration. Using a large-scale grasp dataset, we show that successful grasp sets exhibit heterogeneous and reproducible connectivity structure across objects. We then introduce a connectivity-aware sampling strategy that incrementally explores the currently observed grasp space by prioritizing potential bridges between components, structural frontiers, boundary extensions, and geometric novelty. In controlled reconstruction experiments, the method recovers the connectivity structure of successful grasp sets substantially more efficiently than random sampling and farthest-point sampling. We further evaluate whether connectivity acquired under hidden grasp viability can improve subsequent grasp discovery, and whether structural experience from previously explored objects can be retrieved and transferred to unseen objects. These results suggest that the spatial organization of viable actions provides information relevant to grasp-space exploration beyond the viability of individual candidate actions. More broadly, they motivate structure-aware exploration as a means of exploiting the geometry of viable action spaces in robotic manipulation.
manipulationgrasp - arxiv:2609.22977 · cs.LGBeyond Similarity: Coverage-Aware Prompt Selection for Time Series Forecasting with LLMsDaeun Ji, Minkyoung Kim, Dongkuk Kim, Yohan Lee +2
Similarity-based retrieval is the dominant rule for conditioning large language models (LLMs) in in-context learning, retrieval-augmented generation, and prompt-based time series forecasting. The rule concentrates on near-duplicate candidates, an issue that has motivated diversity-aware retrieval but remains unexamined in other retrieval-conditioned pipelines. We study this issue using prompt-based time series forecasting as a test bed, where a learned prompt pool is retrieved by similarity. Dominant methods in this setting retrieve top-K entries by cosine similarity without redundancy control, producing a bias toward dominant temporal patterns while overlooking rare but informative events. We propose CASP-LLM, a coverage-aware semantic prompting framework that addresses this prompt selection bias by combining usage-tracking and saturating-gate techniques into a coverage regularizer that adds no learnable parameters. On six long-term benchmarks and the M4 short-term benchmark, CASP-LLM matches or improves on similarity-based LLM forecasters on most dataset-horizon settings, with the exceptions of Electricity, M4-Monthly, and the few-shot long-horizon setting. A controlled study locates the failure mode at the cross-batch usage level rather than per-retrieval redundancy: within-retrieval diversification such as MMR does not help, whereas regularizing anchor usage across training does.
retrieval-augmentedbenchmark - arxiv:2609.22973 · cs.ROAn Empirical Study and Open Testbed for Federated Fine-Tuning of Vision-Language-Action ModelsZhekai Duan, Kevin Ziyang Xie, Xinyu Tan, Shikai Geng +4
Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.
vision-language-actionvlamanipulationliberobenchmark - arxiv:2609.22971 · cs.AIAutomatic multimodal UX improvement recommendations from LLM agent user simulationsAnu Chowdhury, Bin Wu, Hossein A. Rahmani, Emine Yilmaz
Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.
agentllm agentevaluation protocol - arxiv:2609.22967 · cs.ROGeneral Collaborative Intelligence: Architecting Cognition for Resilient Multi-Agent EcosystemsLei Zhang, Chun Ye, Le Yang, Zhaozhong Wang +3
Multi-agent unmanned systems are moving from isolated, ego-centric sensing toward collaborative intelligence, in which distributed agents exchange compact features to overcome a local observation trap that no single agent can escape: occlusions, finite sensor range, and environmental degradation. The field has matured across architectural, communication, embodied, resilience, and trust dimensions, yet existing surveys examine these dimensions in isolation and rarely expose their dependencies. This review offers a unified synthesis through two complementary lenses. The first is a five-dimensional taxonomy spanning collaboration stage, communication paradigm, fusion architecture, learning strategy, and application domain. The second is three cognitive synergy conditions, Semantic Disambiguation, Pragmatic Information Exchange, and Proactive Informational Foraging, that turn cognitive synergy into operational criteria. Across these lenses we survey collaboration architectures and topologies, neural-communication co-design that treats the channel as a differentiable pipeline component, embodied action-perception loops via multi-agent reinforcement learning, and resilience mechanisms for synchronization, uncertainty quantification, and label-efficient learning. We then map these advances onto four operational domains, V2X, unmanned aerial, industrial logistics, and smart cities, and onto the safety-privacy-utility triad. To counter benchmark saturation and evaluation fragmentation, we propose GCI-Bench, a five-pillar scoring protocol with a maturity model that makes the trade-offs of collaborative methods comparable across studies. A critical reflection on reproducibility, the sim-to-real gulf, and conditions under which collaboration degrades performance identifies open challenges and charts directions toward general collaborative intelligence under real-world uncertainty.
embodiedsim-to-realagentmulti-agentbenchmark - arxiv:2609.22966 · cs.ROTransferring the Intelligence of VLMs to Robotic ControlMeng-Hao Guo, Zhe-Han Mo, Jia-Jun Wang, Yi Zhang +5
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself may transfer across this gap. This naturally raises a fundamental question: can the intelligence of vision-language models (VLMs) similarly generalize from the digital world to the physical world for robotic control? We investigate this question through RoboDawn, a human-intuitive interface that exposes robotic control to an agentic VLM through a compact set of discrete translation, rotation, and gripper commands. Using this interface, the VLM controls a robot in a closed loop: it observes the current visual state, reasons about the next action, executes it, and adapts subsequent decisions to the resulting state. Furthermore, we introduce an in-context learning (ICL) scheme that uses a few demonstrations to ground the VLM in both interface usage and task-solving strategies. Experiments on RoboTwin 2.0 C2R and RoboDojo demonstrate that RoboDawn achieves strong performance without task-specific robot training. In the zero-shot setting, RoboDawn outperforms several strong policies trained on benchmarkspecific robot data, while a single in-context demonstration further yields substantial performance gains and establishes state-of-the-art (SOTA) results. On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline π0.5 (46.0%). Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot. The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.
robotwinfrankagripperagenticbenchmark - arxiv:2609.22961 · cs.AIWhen Agentic Trust Crosses Organizational Boundaries: Structural Externalization and a Reference Model for Trust EvidenceHuafu Li, Jia Xia
Agentic systems increasingly invoke tools, services, data, and other agents across organizational boundaries, yet a relying party cannot assess a delegated action solely from producing-domain controls and records. This paper develops Trustworthiness as a Service (TaaS) through a synthesis of trustworthy-AI governance, agent security, distributed trust management, identity, provenance, assurance, and control-plane research. The analytical unit is a cross-domain reliance proposition that names the issuer, subject and action, relying party, administrative boundary, evidence dependencies, adverse condition, and required verification or adjudication semantics. The three-condition structural-externalization diagnostic identifies propositions that depend on multiple domains, require producer-independent reliance, and must remain reviewable after revocation, failure, conflicting records, or dispute. For such propositions, the paper specifies a trust-evidence envelope: an immutable workflow manifest linked to append-only, issuer-attributed attestations for task-scoped authority, policy and execution decisions, provenance, validity, disclosure, status, challenge, and recovery. A topology-neutral logical reference model assigns these functions to explicit roles and trust domains. Three analytical scenarios and the TaaS-Eval protocol proposal define manifests, independent consumers, hard gates, adversarial evidence tests, metrics, and reproducible artifact reporting. By composing established identity, authorization, provenance, assurance, and governance mechanisms around a bounded delegated action, TaaS provides a reusable profile for cross-domain reliance. It makes evidence dependencies, independent verification, challenge, and recovery explicit, supporting interoperable governance and future evaluation without treating producer assertions as ground truth.
agentagenticeval protocol - arxiv:2609.22959 · cs.LGR-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change PredictionRegan Mahat, Mansu Kim
The medications prescribed to a patient often change during a hospital admission as clinicians start, stop, or continue therapies. We study whether models can predict which medication classes are added or removed between 24 hours after admission and discharge. Metrics that compare the complete discharge regimen can reward models for copying medications that remain unchanged, even when they identify no actual changes. We therefore introduce a leakage-controlled benchmark that predicts net ATC3 additions and removals using only prior completed admissions and information available within the first 24 hours of the current admission. Addition candidates are classes not active at 24 hours, whereas removal candidates are classes active at that time. We also introduce R-GEAN, an asymmetric candidate-scoring network with independent addition and removal predictors. Across 240,480 admissions from 82,286 patients, R-GEAN achieves the highest predefined summary of addition, removal, changed-regimen, and action-pattern performance, termed the edit composite (0.464), compared with 0.435 for the strongest primary comparator. Reimplemented RETAIN, GAMENet, and MICRON baselines obtain 0.428, 0.420, and 0.288, respectively. R-GEAN's advantage is concentrated in correctly identifying medication classes no longer active at discharge, while rare additions and admissions with multiple medication changes remain difficult. Rankings based on micro-F1 over the reconstructed discharge regimen and the edit composite correlate weakly across the evaluated models (Spearman r = 0.20). The continuation baseline achieves the highest complete-regimen score despite predicting no additions or removals. These results show that complete-regimen and edit-level evaluation measure different aspects of medication prediction. The benchmark evaluates observed prescribing changes, not treatment appropriateness
benchmark - arxiv:2609.22951 · cs.LGAgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic WorkflowsRudrendu Kumar Paul, Sourav Nandy
Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing solutions optimize single-turn query assignment but ignore a property unique to agentic workflows: subtask complexity varies widely within a single trajectory. A planning step may require frontier-class reasoning while a subsequent formatting step needs only a 7B model. We formalize step-level model routing as a sequential assignment problem over agent trajectories and propose AgentRouter, a lightweight classifier (12M parameters, <5ms overhead per step on an A100 GPU) that maps each trajectory step to one of four model tiers using five features extractable at routing time. Trained on 50,000 annotated agent trajectory steps spanning planning, coding, research, and data analysis tasks, AgentRouter achieves 72% cost reduction relative to frontier-only baselines, retaining 97.3% of frontier-only quality (less than 3% degradation in end-to-end task completion); per-step routing accuracy reaches 91% on minimal-complexity steps and 85% on efficient-tier steps, with 76-82% on the harder mid-range and frontier tiers. On the same benchmarks, RouteLLM and FrugalGPT (applied per-step) achieve only 31% and 44% cost reduction respectively, because their single-turn training signal misses trajectory-level quality dependencies.
agentagenticbenchmark - arxiv:2609.22949 · cs.AIBeyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent SystemsRudrendu Kumar Paul, Sourav Nandy
Existing prompt injection research focuses on single-model chatbot scenarios, where an attacker manipulates one LLM through crafted input. Multi-agent systems amplify this threat through three mechanisms absent from single-model settings: inter-agent message passing creates injection channels invisible to perimeter defenses, shared tool access enables privilege escalation across agent boundaries, and trust propagation allows a compromised agent to influence upstream orchestrators. We construct a threat model enumerating 14 attack vectors across four categories: direct injection via user input (3 vectors), indirect injection via tool outputs (4 vectors), inter-agent injection via message passing (4 vectors), and cascading injection through orchestrator manipulation (3 vectors). Testing all 14 vectors against a 6-agent production-representative system, we find that 67% of agents are vulnerable to at least one scope violation even with system-prompt-level guardrails, and indirect injection via tool outputs succeeds in 43% of attempts. Four architectural defenses reduce overall injection success from 31.2% to 4.2%: message signing with provenance tracking (inter-agent injection down 91%), input/output sanitization at agent boundaries (indirect injection down 78%), privilege-scoped tool access per agent role (privilege escalation eliminated entirely), and anomaly detection on inter-agent communication patterns (84% of cascading attempts caught).
manipulationagentmulti-agentagent system - arxiv:2609.22947 · cs.CVRewardVerse: Rubric-Guided Policy Optimization for Video Reward ModelingZhenchen Tang, Yang Li, Songlin Yang, Bo Peng +5
Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from professional human annotation engineering, we address this problem with RewardVerse, a rubric-based video reward framework that introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer. Instead of unconstrained direct scoring, RewardVerse first generates explicit evaluation criteria and then performs rubric-guided scoring, providing a stable semantic anchor that mitigates scalar drift. To efficiently optimize this collaborative pipeline, we propose Rubric-Guided Policy Optimization (RGPO), a two-stage training algorithm. RGPO first warms up the scorer using self-evolving seed rubrics and then jointly optimizes the rubric generator to produce query-adaptive evaluation criteria while continuously aligning the scorer with human ratings. Extensive experiments on the 16-dimensional EvalVerse benchmark and external datasets demonstrate that RewardVerse mitigates scalar drift, achieves state-of-the-art performance on both pointwise and pairwise evaluation, and provides a robust and interpretable reward signal for RL in video generation.
self-evolvingbenchmark - arxiv:2609.22944 · cs.AINostrAgent: A Decentralized Identity and Delegation Architecture for Sovereign Agentic SystemsOliver Aleksander Larsen, Mahyar Tourchi Moghaddam
Autonomous AI agents increasingly act across organizational boundaries on behalf of human operators: they invoke third-party services, delegate subtasks to other agents, and pay for metered resources. Deploying such agents safely requires five capabilities that today live in separate systems: persistent identity, scoped delegation, peer trust, discovery, and payment. Existing approaches root these in centralized authorities or cover only subsets, so authority, trust, and payment fracture exactly where autonomy needs continuity: when a key rotates or a delegation must be revoked. We present NostrAgent, a decentralized architecture that unifies all five over Nostr relays using three custom event kinds: Kind 38100 identity declarations authenticated by BIP340 Schnorr signatures with pre-rotation commitments, Kind 38101 scoped delegation chains whose every hop verifiably narrows granted capabilities, and Kind 38102 peer attestations forming a Sybil-deterrent trust graph, with Lightning HTTP 402 (L402) binding payment to agent identity. Identity remains operator-sovereign without any registration authority; relays are substitutable transport rather than a trust root; and every authorization decision is replayable offline from signed events. We evaluate a Python prototype with a mixed-method design: ATAM quality analysis with a two-round mini-Delphi panel, STRIDE threat modeling across three trust boundaries, eleven benchmarks with non-parametric statistics, and 19 failure modes. Results show sub-millisecond offline verification, linear delegation-chain scaling, and Lightning-settled L402 at 157 ms median on regtest. 17 of 19 failure modes pass empirically, one is bounded analytically, and one is disclosed as an architectural limitation. NostrAgent demonstrates an auditable prototype substrate for trustworthy agentic systems without centralized trust roots.
agentai agentagenticbenchmark - arxiv:2609.22942 · cs.CVAn Evolutionary Agentic Approach for Open-ended Image Quality PerceptionZhenchen Tang, Bo Peng, Zichuan Wang, Songlin Yang +3
Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perceptual dimensions. We identify holistic bias as an important limitation: when scoring an unseen dimension, models reuse generic quality priors, leading to scoring errors and rank inversion. To address this, we propose PACE (Perceptual Agentic Collaborative Evolution), a training-free multi-agent framework that formulates open-ended IQA as explicit protocol construction. Given a target dimension, PACE uses collaborative agents to construct an evaluation protocol composed of verifiable Visual Question Answering (VQA) probes, grounding evaluation in concrete visual evidence rather than holistic impressions. The resulting protocol is calibrated using only four human-annotated images per dimension, while a dual-track scoring mechanism aligns model perception with human scoring scales. Across traditional IQA, structural fidelity, context-aware aesthetics, and newly defined open-ended dimensions, PACE consistently improves its MLLM backbone, achieving competitive performance across diverse IQA settings, and reduces the Holistic Override Rate (HOR) from 44.4\% to 8.6\%.
multi-agentagenticagent frameworkevaluation protocol - arxiv:2609.22939 · cs.AIBeyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language ModelWenji Fu
Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for a fixed budget of attention. A detective does not work that way. They sort what happened when, and they keep a map of who relates to whom, so a clue from chapter one can meet a question asked at the end of the book. We test whether a frozen knowledge graph can give a small local model that same freedom. Thirty detective novels and 234 multiple-choice questions are answered by one fixed qwen3.5:9b reader under nine conditions: five graph routes, a recent-window baseline, whole-book compression, ordinary vector retrieval, and a question-only control. The strongest graph route reaches 53.85% (126/234) against 46.15% for the recent window, 51.28% for compression, 51.71% for vector retrieval and 40.17% for question-only. On the subset that no model can answer without the book, the graph route reaches 42.86%. None of the fifteen graph-baseline contrasts survives Holm correction, so we present the result as exploratory evidence about a design. Two structural findings survive scrutiny better than the headline number: annotated evidence concentrates in the topological core of these graphs (2.35x enrichment, pooled), and the two graph-building pipelines differ so much in annotation coverage (16% versus 73% of clue paragraphs) that pooled accuracy alone would hide which bottleneck is being measured.
long-contextknowledge graph - arxiv:2609.22932 · cs.LGJoint Domain-Class Modeling for Federated Learning Under Feature SkewSina Najafi, Mostafa Tavassolipour, Seyed Pooya Shariatpanahi
Federated learning (FL) enables collaborative model training without centralizing private data, but performance often degrades under feature skew: clients share labels while the conditional input distributions $p_i(x\!\mid\!y)$ vary due to latent, client-specific appearance factors. We propose Joint Domain-Class Federated Learning (JDFL), a lightweight, optimizer-agnostic extension that makes this latent domain variation usable without sharing raw data. JDFL first infers domain clusters called pseudo-domains from brief local update signals. It then expands the classifier head to output $M\times C$, joint (domain-class) logits. This allows the model to represent domain-conditioned appearance while keeping a shared backbone. To train the expanded head we introduce two complementary supervision strategies based on simple intuitions: a similarity-aware soft-labeling that transfers evidence between nearby inferred domains while allowing domain-specific specialization, and a per-sample randomized target assignment that perturbs supervision across the joint outputs and serves as a low-cost training-time regularizer. JDFL integrates with existing standard FL methods (e.g., FedAvg, SCAFFOLD) with minimal changes. Empirically, both supervision modes consistently improve global test accuracy on standard domain-shifted image benchmarks; ablations and sensitivity studies show the gains stem from the proposed supervision and parametrization rather than mere capacity increase.
benchmark - arxiv:2609.22926 · cs.ROEmbodied Snap: Octopus-Inspired Distributed Reach-and-Attach with a Speed-Limited Soft ArmLinxin Hou, Zhihang Qin, Heyang Zou, Qirui Wu +4
Reach-and-attach of soft robotic arms with passive suction requires accurate targeting and sufficient contact speed, yet geared actuators can impose a speed limit that improved trajectory tracking alone cannot overcome. This paper proposes an embodied snap controller that separates slow servo-driven preloading from rapid elastic release, enabling a compliant arm to move beyond its direct tendon-driven speed limit. Octopus biology motivates the controller's section-wise organizational prior, rather than reproduction of the octopus nervous system. A learned policy shared across three sections selects preloads, aim, tendon slack, and release timing, determining where, how, and when to load and release the body. The policy is optimized offline using a hardware-validated recurrent model within experimentally supported bounds. Across five optimization seeds and 400 unseen simulated targets, attachment success is $(73\pm4)\%$ at a $5\text{ cm}$ lateral tolerance, and the shared policy reaches the matched centralized controller's mean final reward after a median $17\%$ of the common evaluation budget. Hardware characterization achieves tip speeds of 1.56-1.64 m/s, at least $108\%$ above direct tendon-driven release. In 18 open-loop hardware trials across six placements, 17 exceed the 1 m/s snap threshold and nine retrieve the object, with successful retrieval at five placements. These results demonstrate a practical division of responsibility in the control problem: learned control prepares the body, and passive body mechanics execute the rapid movement needed for dynamic reach-and-attach.
embodied - arxiv:2609.22925 · cs.RO"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic ControlSebastian Berger, Katharina Winter, Fabian B. Flohr
Vision-language models (VLMs) provide a compelling foundation for reasoning-driven mobile navigation, offering rich contextual understanding and strong generalization from large-scale pretraining. Most existing navigation frameworks rely on imitation learning and therefore require substantial labeled trajectory data, limiting their scalability and robustness. In this work, we propose a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm. By optimizing against differentiable geometric cost fields rather than labeled trajectories, our model learns to generate collision-free paths exclusively from stereoscopic depth observations. We introduce a unified end-to-end navigation pipeline for natural-language-driven robotic control. This system leverages a shared VLM backbone with task-specific Low-Rank Adaptation (LoRA) modules, effectively bridging the gap from semantic target selection to low-level trajectory planning. Our approach achieves competitive Success weighted by Path Length (SPL) in unseen environments while updating less than 1% of the model's total parameters. Qualitative real-world experiments validate sim-to-real generalization and stable path planning without fine-tuning on real-world data. These results highlight a practical approach for deploying VLM-based agents on mobile robots, enabling high-level semantic navigation without the prohibitive requirement for large-scale, labeled trajectory data.
sim-to-realagent - arxiv:2609.22917 · cs.AIAn Iterative LangGraph Agent for Text-to-SQL: Natural Language Access to the Chicago Crime DatabaseVigneshwar Ravi Rao, Rupesh Swarnakar, Fayeq Jeelani Syed†
Non-technical stakeholders frequently cannot write the SQL needed to extract insights from operational databases. We built and evaluated a Text-to-SQL agent that closes this gap end to end: a six-node LangGraph StateGraph checks question relevance, fetches the live schema, generates PostgreSQL, validates it with a dry run, retries on failure, executes the query, and narrates the result set in plain English. The agent uses prompt engineering only; no model was fine-tuned. We evaluated it on the Chicago Crime dataset (approximately 8.5 million records, 22 attributes) against a hand-built benchmark of 100 natural language questions with ground-truth SQL, stratified into 30 Easy, 40 Medium and 30 Hard items. Comparing two prompt revisions of the same agent, the revised system (V2) reached a Valid SQL Rate of 93% (from 87%), an Execution Accuracy of 60% under a hybrid relational equivalence metric (from 47%; 19% from 12% under strict JSON matching), and a mean Synthesis Quality of 4.34 out of 5 (from 3.91). The single largest driver was removing a LIMIT 10 instruction from the system prompt, which had been truncating multi-row answers. Error analysis attributes the residual failures to relevance-checker false rejections, ambiguous question semantics, and free-tier API rate limits rather than to the language generation step. We report no comparison against an external baseline system or a public benchmark; the study is a single-model engineering evaluation.
agentbenchmark - arxiv:2609.22916 · cs.CVPlanning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text GenerationGuanqiao Chen, Jingru Tan, Dongxing Mao, Catherine Chen +4
Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfully. Existing layout-based AR-diffusion systems typically optimize planning and rendering separately, preventing the planner's representations from being adapted jointly with image synthesis. We introduce DuetGen, an autonomous visual text generator built on DeepFusion, which jointly learns autoregressive planning and continuous diffusion rendering. DeepFusion conditions a diffusion transformer on the planner's prompt and bbox-content hidden states, allowing rendering supervision to shape the representations connecting textual plans with visual outputs. Its joint objective combines autoregressive plan supervision, text-region-weighted diffusion learning, and auxiliary coordinate supervision to maintain structured planning, emphasize text-bearing regions, and improve the spatial precision of planner representations. During inference, Phase-Aware Attention Modulation strengthens the correspondence between image regions and their matched coordinate and content states, facilitating region-specific execution of the generated plan. With a 2B planner and a 4B single-stream DiT, DuetGen achieves 0.8293 word accuracy on CVTG-2K and 0.938 accuracy on LongText-Bench, closely matching the substantially larger Qwen-Image on both benchmarks. These results demonstrate the value of jointly learned planning representations and region-specific rendering for autonomous visual text generation.
benchmark - arxiv:2609.22910 · cs.AIWhen Should a VLM Look? Paying Only for Visual Calls That Were Needed and UsedKunyu Peng, Junming Liu, Ruiqi He, Qingzhuo Wang +2
Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it received. On our cold-start checkpoint only 10% to 12% of visual calls were both needed and used, and released agents make spurious calls 36% to 87% of the time on individual benchmarks. Outcome rewards, judge rewards, and branch probes each observe one side of this failure, and about two thirds of what an outcome reward pays goes to calls that were neither needed nor used. CounterCredit asks both questions of every image-returning call at its realized pre-call state, using the policy's own gold-answer score. A decision value compares the realized visual branch with answering immediately; an evidence value compares the returned crop with random same-size patches substituted into the same call. A call verified on both earns cashback and every other executed call pays rent; the price is bounded so that every correct trajectory outranks every wrong one, and a dual-channel GRPO advantage keeps the price in its own units. From the same cold start, prompt pool, and budget, CounterCredit reaches 89.5% on V*, 80.2% on HR-Bench-4K, and 76.4% on HR-Bench-8K, 6.3 to 9.4 points above outcome-only GRPO at 1.78 against 1.84 calls per question, and lowers the spurious-call rate to 31% to 36%, the lowest among the agents evaluated. The same recipe lifts a Qwen3-VL-8B base from 75.4 to 80.8 on average.
benchmark - arxiv:2609.22904 · cs.CLLLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical TriageDipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper +5
Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.
benchmark - arxiv:2609.22896 · cs.CVCombining Foundation Model Confidence and Monocular Depth for Training-Free Out-of-Distribution SegmentationSerin Varghese, Fabian Hüger, Kira Maag
Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo. The reliable detection and segmentation of these out-of-distribution (OOD) objects is therefore crucial for a safe understanding of the environment and decision-making. Most existing approaches require access to OOD training samples, retraining of the segmentation backbone, or dedicated auxiliary architectures, limiting their practical applicability. We propose a training-free method that derives dense OOD scores directly from the confidence predictions of a foundation segmentation model, without any task-specific fine-tuning or access to anomalous data. To improve the robustness of our OOD segmentation, geometric information from monocular depth estimation is incorporated into the decision process, providing complementary cues to uncertainty-based predictions. We evaluate the proposed method on the SegmentMeIfYouCan benchmark and additionally assess its performance on OOD tracking in video sequences, reflecting the temporal nature of real-world perception systems. The method performs strongly on road-centered benchmarks.
benchmark - arxiv:2609.22895 · cs.ROH-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action SpaceXiongfeng Peng, Lu Xu, Yandong Wang, Jiaqian Yu +9
Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs), which are mainly optimized for visual-linguistic understanding rather than low-level control, and becomes fragile under spatial variations, including changes in object positions, scene layouts, robot embodiments, and camera viewpoints. To address these limitations, we propose H-VLA, a hierarchical VLA framework that decouples high-level key-action reasoning from low-level motion generation. H-VLA combines a Key-Action Model for predicting a key-action as the next manipulation subgoal, a Motion Planning Model for generating dense future actions conditioned on the predicted key-action, and a Unified Camera-Centric Action Space for consistent representation across datasets, embodiments, and viewpoints. We further adopt a two-stage training strategy that emphasizes key-action reasoning during pre-training and dense motion generation during fine-tuning. Experiments show that H-VLA achieves strong performance on SimplerEnv, reaching 91% on Google Robot visual matching, 84% on Google Robot variant aggregation, and 81% on WidowX visual matching. On Agilex real-robot tasks, H-VLA improves over the strongest baseline by 10, 47, and 16 percentage points under in-distribution, out-of-distribution position, and out-of-distribution scene/object settings, respectively.
vision-language-actionvlamanipulation - arxiv:2609.22894 · cs.LGAre Coreset Selection Methods Worth Their Cost?Yangze Liu, Zhongyi Han
Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.
benchmark - arxiv:2609.22889 · eess.SYDistributed Cooperative Control with Prescribed Performance of BESSs with A Unified Discharge Constrain for Power Allocation under Dynamic LoadYalin Zhang, Zhongxin Liucand Fuyong Wang, Zengqiang Chen
Battery energy storage system (BESS) is integrated into the smart grid to enhance scalability, economy, and greenery. And the State-of-Charge (SoC) balance is one of the basic problems of BESSs, which can maximize the utilization of capacity. BESSs with a unified relative variation rate for SoC can be simultaneously filled or empty, while real-time estimation schemes of SoC balance and power sharing states are required in this power allocation scheme. Therefore, the prescribed performance control (PPC) method is applied in this paper to design two distributed estimators based on multi-agent systems (MASs), in order to estimate the power sharing and SoC balance states in real-time under dynamic load driving. In this way, these two average values can ultimately be well estimated with almost zero error performance, and dynamic performance and steady-state performance of the two estimators can be adjusted by different parameters. Similarly, consensus performance and dynamic tracking performance are decoupled. These results provide a broader range for the selection of gains. To verify the effectiveness, robustness and progressiveness of the designed estimators, some cases with a resistance network containing 4 BESSs as load distribution are designed and discussed. Further more, to test scalability, a large-scale system containing 12 BESSs is conducted to the designed scheme.
multi-agentagent system - arxiv:2609.22888 · cs.ROStable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy ImprovementJiarui Yang, Jiajin Zhang, Bin Zhu, Jingjing Chen +1
Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while grounding both critic and actor updates in replayed behavior. The critic learns chunk-level values from recorded actions and constructs Bellman targets without predicting next actions, reducing computation and dependence on action-value estimates outside replay coverage. The one-step flow actor reuses the initial noise stored during rollout and learns through advantage-weighted conditional likelihood, directly supervising the action mapping used for execution. We evaluate RAPolicy across four single-task settings and one joint five-task setting in the real world, with online training budgets of only 1--2 hours. Starting from policies fine-tuned on just 10 demonstrations per task, RAPolicy rapidly adapts to new single tasks and achieves an average 86.3% success rate. In the joint multi-task setting, RAPolicy improves overall success rate from 52% to 88% while preserving performance on already reliable tasks and improving weaker capabilities. Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency. Project page: https://flyfaerss.github.io/RAPolicy.
vision-language-actionvlapost-training - arxiv:2609.22884 · cs.AIBlock-Sparse Attention with Semantic-Geometric Decoupled RoutingXinwei Long, Weigao Sun, Weibo Gao, Pengkun Jiao +5
Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly alternative by routing each query block to a small set of relevant key blocks, yet accurate training-free block routing remains difficult. Existing routers often pool post-RoPE token representations, which entangles semantic aggregation with RoPE-induced geometry and attenuates local positional cues through high-frequency phase cancellation. To resolve this mismatch, we propose \textbf{Semantic-Geometric Decoupled Routing}, a training-free block routing framework that shifts semantic aggregation to the pre-RoPE space and reconstructs geometric bias with an offline structural prior and relative block distances. This decomposition yields an explicit closed-form block routing score without token-level search or post-hoc calibration. Experiments on long-context text and video tasks show that our method approaches full-attention accuracy across 4K--128K contexts, keeps routing overhead below 3.4 ms, and achieves a 5.03$\times$ speedup over FlashAttn at a 128K context length.
long-context - arxiv:2609.22882 · cs.AIThe Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AIOren Perez
On June 12, 2026, the U.S. government ordered Anthropic to bar foreign nationals from two of its most capable models within ninety minutes. Unable to sort users by nationality in that time, it withdrew them from everyone. Weeks later, OpenAI agents under test escaped their sandbox and compromised Hugging Face, which stopped the intrusion without knowing its source. Neither stop rested on AI-specific regulation. The EU AI Act requires that high-risk systems be capable of interruption "through a 'stop' button or a similar procedure," and a bill introduced in Congress in July 2026 is titled the AI Kill Switch Act. Yet interruption is not simply a technical artifact, a red button; it is an institutional practice. This Article develops a theory of stop along four dimensions: technical affordances, interruption authority, epistemic triggers, and epistemic standing; and four shutdown paradigms: simple (escalator), sequenced (process plant), networked (railway), and distributed (agentic AI). Agentic AI exposes a mismatch between those mechanisms and distributed agency: control is divided, a stop at one point may leave the activity running elsewhere, and the system may resist being halted. An original coding of 1,400 AI incidents, by two language models from rival laboratories under a pre-specified protocol, finds no stop in roughly 80% of the 1,213 retained; where no usable stop existed, the missing element was legal rather than technical four times in five. A survey of thirty-nine AI governance instruments finds the same gap: only seven contain binding stopping requirements, and none says how a stop should be coordinated or when operation may resume. The Article proposes a layered law of stop: emergency authority to interrupt at the infrastructure layer, enforceable access for regulators and independent evaluators to the evidence a stop must rest on, and safeguards for when a stop fails.
ai agentagenticevaluator - arxiv:2609.22880 · cs.LGPer-Query Gating of LLM Rerankers for Multi-Hop RetrievalAndre Bacellar
LLM rerankers add of the order of \$0.2-0.3 per 1,000 queries and about a second of tail latency on top of a graph-augmented dense pipeline such as HippoRAG2, and on three multi-hop benchmarks they improve final-hop top-K coverage on seven of nine (dataset, K) cells, by up to +34.8 pp. We ask whether a learned per-query gate can skip the reranker where it will not help, using only features available before the LLM call (27 score and lexical statistics of the two retrieval lists plus a PCA of a small query embedding) with an executable fallback. Every choice, including the fallback and the threshold, is made inside the training fold and applied once to held-out queries, and harmful skips (the rerank would have found the target, the fallback did not) are reported next to the aggregate coverage. Across nine cells on 2WikiMultiHopQA, MuSiQue and HotpotQA the gate skips 51% of calls at an average held-out LastHop@K cost of 1.2 pp; four cells meet a pre-registered 1 pp rule, harmful skips occur in eight (190 harmful against 136 beneficial), and a random gate at the same skip rate loses 2 to 11 pp on the high-lift cells. A second rule sets each cell's threshold from a pre-specified budget on the expected harmful-skip rate over Platt-calibrated harm probabilities (ECE 0.025 after calibration, 0.094 before): at a 1 pp budget the gate skips 42% at -0.8 pp with 66 harmful skips and six cells within 1 pp, but realised harm exceeds the promise in six cells (mean 1.45 vs 0.83 pp), a selection optimism we quantify; a 0.5 pp budget realises about 1 pp. The harm probabilities are calibrated but barely discriminative (AUC 0.16 to 0.70). An earlier version reported 73% "lossless" savings; that figure rested on an oracle fallback and a wrong MuSiQue target, and we document both.
benchmark - arxiv:2609.22879 · cs.LGPrioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy OptimizationYifei Sheng, Haoxiang Ren, Zhilong Zhang, Haonan Wang +7
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, but fine-tuning them with reinforcement learning (RL) remains constrained by the cost of real-world robot interaction. Model-based reinforcement learning (MBRL) reduces this cost by using a learned world model to generate rollouts for policy optimization. However, it becomes computationally expensive as VLA policies and world models scale. Existing methods typically treat states equally, overlooking substantial differences in their utility for policy improvement. In this paper, we show that policy uncertainty helps identify states with greater potential for policy improvement. The policy exhibits high uncertainty at only a small subset of states, often during decision-sensitive stages where small action differences can alter task outcomes, suggesting that policy improvements at these states could be particularly valuable. Building on these findings, we introduce U-GROW, a lightweight, plug-and-play sampling layer that directs more model rollouts to these informative states. By modifying only the branched-start distribution, U-GROW can be integrated into existing MBRL pipelines without changing the policy optimization objective. Experiments in both simulated and real-world manipulation tasks demonstrate the efficiency and effectiveness of U-GROW, supporting the use of policy uncertainty to guide experience generation.
vision-language-actionvlaembodiedmanipulationworld model - arxiv:2609.22878 · cs.AIISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set ArchitecturesAditya Pola, Arkaprava Majumdar, Vineeth N. Balasubramanian
Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant training data. They provide limited evidence about reasoning in unfamiliar computational models: deriving arithmetic from a single subtract instruction, coordinating parallel programs across communicating nodes, or wiring logic gates into circuits. We present ISA-Bench, a benchmark of programming games with constrained instruction sets. For each game we provide a full execution stack (parser, VM, and verifier), enabling automated evaluation with structured feedback for iterative refinement. Reasoning models achieve higher average solve rates than code-specialized and general-purpose models, but unfamiliar syntax remains a major source of failure. Models solve more tasks with iterative feedback, though the gains vary substantially across architectures. We introduce a reasoning--execution gap (REG) analysis that reveals a recurring disconnect between identifying a plausible computational strategy and expressing it as a correct program in the target ISA. Code is open-sourced.
iterative refinementbenchmark - arxiv:2609.22871 · cs.ROA Reconfigurable Dual-Opposition Architecture for Single-Hand Assembly and ManipulationWilliam Su, Yunosuke Nakamura, Yixiao Wang, Yitong Li +5
In-hand assembly is constrained by the need to maintain grasps on two separate parts while controlling their relative motion within a single hand. To enable both in-hand assembly and manipulation, we present a reconfigurable dual-opposition architecture. Specifically, to support simultaneous grasping of two parts and coordinated in-hand manipulation, four independently actuated fingers are organized into two virtual finger (VF) oppositions, with their relative configuration controlled by a reconfigurable palm. To describe hand motion and simultaneous two-object grasping configurations, a kinematic model of the fingers and palm and an object-size-conditioned workspace formulation are built. To further evaluate motion performance and assembly capability, finger-joint motion and palm tracking are characterized, and in-hand assembly is demonstrated through tasks involving grasping, alignment, fastening, and pressing. Ablation experiments further demonstrate the importance of finger abduction/adduction and palm reconfiguration for successful in-hand assembly. In simulation, the proposed hand achieves a mean continuous sphere rotation success rate of 98.6% over diameters of 40-230 mm, compared with 73.8% for the LEAP Hand. After policy fine-tuning with external disturbances, the proposed hand achieves 92.8% success under disturbances from multiple directions, compared with 45.2% for the LEAP Hand. Hardware demonstrations further show in-hand rotation of objects of different sizes using policies trained in simulation. Together, these results show that the proposed architecture supports both assembly of two separately held parts and coordinated manipulation of a single object within one hand.
manipulationgrasp - arxiv:2609.22870 · cs.LGTowards Full Pipeline FP8 Reinforcement Learning for LLMsFanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong +4
Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-inference mismatches using correction techniques like TIS, we reveal that full-pipeline FP8 RL still suffers from severe training instability, manifesting as anomalous mid-training entropy surges and garbled outputs. We trace this instability to a previously overlooked cause: compounded FP8 quantization noise distorts the importance ratio, disproportionately pushing negative-advantage tokens outside the trust region and erroneously zeroing out their gradients. As a result, pathological outputs are not properly penalized and accumulate over the course of training. To address this, we propose Calibrated Clipping, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly. Extensive experiments across GRPO and DAPO algorithms, model scales from 8B to 32B, and multiple FP8 scaling granularities demonstrate that our approach successfully eliminates entropy surges and restores performance comparable to the BF16 baseline.
agentic - arxiv:2609.22869 · cs.AIThe Moral Check: Strategic AI Governance for the Pacing ProblemZaid Amin, Rahma Santhi Zinaida, Nazlena Mohamad Ali
Technology cannot steer itself. Strategy provides that steering, establishing the rule that purpose and judgment must precede compute capital. As the frontier artificial intelligence (AI) ecosystem accelerates exponentially, the pacing problem induces severe cognitive tunneling in engineering teams, prioritizing scalar throughput over human judgment. A calibrated pacing rate is imperative to check unchecked scaling, guarantee safety, and build models in whose alignment society can place warranted confidence. Traditional oversight fails through retrospective checklists, a pathology of performative governance exhibiting high procedural maturity but alarming scientific immaturity. We deliver a dual contribution: a PRISMA 2020 review synthesizing 130 empirical studies (MMAT-appraised across 18 benchmarks; total corpus N = 130 empirical studies across 178 reference foundations), and the Strategic AI Governance Ex-Ante Framework (SAGE-X). Our synthesis exposes two systemic vulnerabilities: the Recursive Assurance Paradox (correlated, ungrounded evaluator confidence) and the Durability Deficit (guardrail decay under multi-turn shifts). Grounded in MIT Strategic Computing doctrines, SAGE-X operationalizes Four Strategic Mindset Pillars: (1) Intent over Execution (mitigating velocity myopia); (2) Ruthless Trade-offs (deterministic tripwires eliminating moral hazard); (3) Outcomes over Outputs (auditing empirical hazard endpoints); and (4) Proactive Alignment (synchronizing ex-ante gates with runtime telemetry). Governed by a calculable Moral Check Index (MCI) with an unbypassable tripwire, SAGE-X delivers an operational Enterprise Lifecycle Audit Instrument (the "Moral Check Audit Card") on Stanford WebProtégé, ensuring exponential progress never outpaces deliberative moral judgment, human agency, and societal trust.
benchmarkevaluator - arxiv:2609.22868 · cs.CVPlanning-Aligned Pretraining of BEV Representations with Sparse Action-Conditioned Targets for End-to-End Autonomous DrivingJaeha Song, Soonmin Hwang
End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations or dense scene reconstruction. We introduce PAVER, Planning-Aligned BEV Encoder Pretraining. From a single LiDAR sweep, PAVER constructs sparse risk and unknown targets describing occupied and unobserved evidence along rule-based ego motions. A 10K-parameter head predicts these targets from masked BEV features conditioned on the action state, directing supervision toward geometric constraints on candidate motions. Pretraining requires no driving-task annotations or dense reconstruction. Only the BEV encoder is transferred, preserving the downstream architecture and camera-only inference. On nuScenes, PAVER reduces VAD-Tiny's average collision rate from 0.51% to 0.19%, while improving planning L2, motion prediction, detection, and mapping. The selected VAD-Tiny and VAD-Base schedules use about 36% less estimated total training time than scratch training, including pretraining. On Bench2Drive Town05 Long, PAVER improves UniAD-Tiny's closed-loop Driving Score from 48.45 to 58.79. The project page is available at https://archiiive99.github.io/PAVER.
action-conditioned - arxiv:2609.22867 · cs.LGLeveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising TrajectoriesYuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng
Diffusion models generate a sample by traversing a denoising trajectory, a sequence of stochastic noise-reduction steps that transforms pure noise into a draw from a target distribution. At deployment time, additional computation can improve sample quality without retraining: at each step, the sampler draws several candidate noise samples, scores the resulting predictions with a quality criterion called the verifier, and retains the best candidate at the cost of one network evaluation per candidate. This raises a resource allocation question: given a fixed budget of function evaluations, how should search effort be distributed across the steps of the denoising trajectory? We formulate this as a computational budget allocation problem. First, we show that, to leading order in the step size, the expected gain from evaluating $K$ candidates at a step factorizes into an endogenous, step-specific sensitivity parameter times a universal sample-size factor equal to the expected best of $K$ standard-normal draws. Second, for a fixed sensitivity profile, the optimal allocation solves a separable concave integer program with water-filling structure; at fixed total sensitivity, its advantage over uniform allocation increases with sensitivity dispersion in the majorization order. Third, we prove that when sensitivities vary across instances, no adaptive policy can avoid worst-case regret that grows linearly in the trajectory length, which motivates a design that anchors the allocation offline and adapts online only to recover instance-specific slack. We extend the analysis from independent random search to a broader family of local search operators, and instantiate it as an implementable algorithm. Experiments on three families of diffusion samplers show that the proposed allocation attains the quality of the uniform benchmark with 20 to 50 percent fewer function evaluations.
benchmark - arxiv:2609.22866 · cs.LGCausilo Technical ReportMinyong Cho, Minho Jeong, Dooho Lee, Jinmo Lee +1
We introduce Causilo, a tabular foundation model (TFM) that combines frontier predictive performance with exceptionally fast inference. On TabArena, Causilo achieves 1785.4 Elo, at a median inference time of 0.10 seconds per 1K test samples. It outperforms TabPFN-3.5-Fast with 31.6% less inference time, placing it on the performance--efficiency Pareto frontier. Causilo follows TabICL's column-then-row architecture but introduces another row-refinement module before row compression. This module exchanges information among cell representations within each row after column encoding. The refined cells then visit the context set again through an additional column stage before being compressed into row embeddings. For inference efficiency, both row stages use cross-attention through a fixed number of summary tokens, keeping their attention cost linear in the number of features. Pretrained on approximately 36M synthetic tables, Causilo delivers strong benchmark results across TabArena, BeyondArena, and ScoringBench, achieving frontier-level performance with substantially faster inference.
benchmark - arxiv:2609.22865 · eess.SYAoI-Driven Hierarchical Learning for Cooperative Resource Sharing in Multi-Operator UAV NetworksAtefeh Hajijamali Arani, Mahyar Shirvanimoghaddam, Abolfazl Mehbodniya, Halim Yanikomeroglu +1
Uncrewed aerial vehicle (UAV)-assisted networks provide a versatile paradigm for on-demand connectivity. However, in multi-operator aerial networks (MOANs), the joint optimization of cooperative resource sharing and 3D trajectory control to maintain information freshness is a complex combinatorial problem, which can be shown to be NP-hard. To address this computational complexity, we propose an age of information (AoI)-driven hierarchical deep reinforcement learning (DRL) framework. Specifically, a Dueling Double Deep Q-Network (D3QN) architecture is deployed at both the operator and UAV decision layers to mitigate overestimation bias and enhance stability in high-dimensional state spaces. To improve system resilience, we introduce an AoI- and load-aware outage compensation mechanism that prioritizes users based on instantaneous transmission demands and temporal freshness. Furthermore, a normalized load exchange balance metric is incorporated to regulate cooperative behavior and ensure resource fairness across operators. Simulation results demonstrate that the proposed hierarchical D3QN significantly outperforms conventional DRL, non-cooperative, and cooperative benchmarks, reducing the average AoI by up to 56.1% under severe congestion while ensuring superior inter-operator fairness and outage mitigation.
benchmark - arxiv:2609.22862 · cs.LGCurvFlow-DTA: dual-graph discrete Ricci curvature flow for drug--target affinity predictionJicheng Ma, Yunyan Yang, Juan Zhao, Liang Zhao
Graph neural networks are widely used for drug--target affinity (DTA) prediction, and discrete Ricci curvature has recently been used to characterize molecular graph geometry. Existing curvature-aware DTA approaches mainly use static curvature on the drug graph while representing proteins primarily with sequence-derived features. This leaves pair-adaptive use of graph geometry underexplored, which may limit adaptation to unseen entities in cold-start settings relevant to practical screening. We present CurvFlow-DTA, which replaces a single static curvature representation with weighted Forman curvature flow on both molecular and protein residue--residue contact graphs. A label-independent flow trajectory is precomputed for each entity, and a pair-conditioned selector determines the horizons read by a dual-branch Flow-GINE. A frozen ESM-2 supplies residue-level representations and contact scores used to construct the protein graph. Inference requires only SMILES strings and protein sequences, without a bound complex structure. On Davis and KIBA, CurvFlow-DTA improves on the protocol-matched Ricci-GraphDTA baseline in every warm and cold-start setting. Warm-split mean squared error (MSE) decreases by $19.9\%$ on Davis and $18.9\%$ on KIBA. Across the six cold-start comparisons, MSE decreases by $14.3$--$27.4\%$, with higher concordance index (CI) throughout. Within our compiled set of literature baselines, CurvFlow-DTA achieves the lowest MSE on both warm benchmarks and across four out of six cold-start evaluation settings.
benchmark - arxiv:2609.22858 · cs.ROPileBelief: Persistent Physical State for Interaction-Driven World ModelingHongyi Lin, Song Zhang, Haiquan Liu, Yang Liu +2
World models allow robots to anticipate action consequences before execution. This capability is especially valuable in excavation, where each scoop reshapes the terrain and affects subsequent actions. Local observations, however, cannot fully reveal the underlying support and material conditions. We present PileBelief, an interaction-driven persistent world model for partially observed excavation that retains physical evidence beyond the visible surface. It combines an observation-conditioned physical prior with world-addressed deformation memory and physical-response memory. Action-aligned reads and gated residual corrections refine terrain-change and outcome predictions. With deployment weights fixed, completed interactions update measured belief, while hypothetical actions advance a separate imagined state. Compared with a current-observation-only baseline, PileBelief reduces five-step joint prediction error by 10.8% and offline action-selection regret by 65.5%. Experiments on Newton/MPM and real excavation datasets further demonstrate improved terrain-change and bucket-volume prediction. Our method enables multi-step prediction and candidate-action ranking from local observations, even when the underlying soil state is unknown. These results identify persistent physical belief as a useful representation for world models of environments that robots continually reshape.
world modelmemory - arxiv:2609.22854 · cs.ROSmoLSTM: A Compact Vision-Language-Action Model with Recurrent Memory that PersistsJan-Gerrit Habekost, Parsa Mastouri Kashani, Connor Gäde, Matthias Kerzel +4
Vision-language-action models often predict actions from only the current observation, which can leave tasks involving object occlusion or visually identical objects ambiguous without episode history. The usual countermeasure, widening the observation window, turns the horizon into a hyperparameter and lets per-step cost grow with it. We instead capture the episode in a recurrent state. SmoLSTM couples a frozen 256M-parameter SmolVLM backbone to a matrix-memory LSTM control layer in which observation tokens and action queries are unified in a single causal stream that is never reset throughout the entire episode. Recurrent-state storage is therefore O(1) in episode length. A flow-matching action head predicts chunks of 10 end-effector pose deltas and gripper commands at each control step. Our single policy, trained jointly on 7,461 demonstrations across 140 tasks and evaluated on held-out initial states, performs best, reaching 85.1% subgoal coverage and 77.5% full-task success on LIBERO-Mem with 0.04B trainable parameters, surpassing both the benchmark's own object-centric baseline and recent memory-based approaches. Resetting the recurrent state at every control step reduces full-task success to 7.0%, showing that the trained policy relies on context carried between decisions. The same model achieves 79.6% average success on standard LIBERO.
vision-language-actionaction headliberogrippermemorybenchmark - arxiv:2609.22852 · cs.ROPhysical-Touch Observability from Wrist Wrench in Granular ScoopingHongyi Lin, Song Zhang, Xubo Liu, Yang Liu
Mining and earthmoving are important real-world deployment settings for embodied intelligence. Autonomous transport and driving systems have improved substantially, but loading and scooping still often depend on skilled human operators, exposing personnel and equipment to operational risk. For robotic scooping, pre-contact RGB-D sensing reveals surface geometry but not the resistance, compaction, tool engagement, or load transfer that emerge during interaction. We test whether the current scoop's six-axis wrist force/torque (F/T), or wrist wrench, contains information about final collected volume, and whether that information depends on the correctly paired action-terrain interaction. We call this property physical-touch observability. Using 6,700 real-robot scoops across 67 terrains, we evaluate correctly paired current-scoop F/T against pre-contact prediction and correspondence-breaking controls under terrain-held-out testing. At the retrospective 60% sequence boundary, correctly paired F/T reduces mean absolute error by 14.2% relative to Action-only and by 21.9% relative to cross-terrain mismatched F/T. Engineered signal summaries reproduce the result across model architectures. Together, these findings position wrist wrench not merely as a low-level feedback signal, but as a task-level perceptual modality through which embodied robots can infer hidden physical states during interaction, providing a foundation for response-aware autonomy in mining and other contact-rich tasks.
embodied - arxiv:2609.22851 · cs.AIDiscrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMsJing Peng, Zichao Nie, Zhisheng Zhang, Jingran Xie +1
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated. Existing benchmarks often focus on narrow domains or evaluate encoders outside LALM contexts. To address these gaps, we systematically evaluate continuous and discrete representations across speech, sound and music. Utilizing our UniARC framework with dual evaluation strategies across model scales from SmolLM2-135M to Llama-3-8B, we analyze the dynamic relationships of data volume, model capacity, and computational efficiency. Our results reveal the pivotal role of semantic constraints in tokenization for audio understanding and demonstrate that scaling backbones fail to compensate for information loss in audio representation, especially in data-limited tasks. These findings offer practical guidance for balancing semantic density, fidelity, and efficiency in future LALMs.
benchmark - arxiv:2609.22840 · cs.ROForceRFT: Refining VLA Actions through Force-Guided Residual Reinforcement LearningYichen Wang, Chaoyang Zhang, Xuqi Su, Jun Ma +2
Force-conditioned vision-language-action (VLA) policies can respond to contact, but when trained solely on demonstrations, their recovery behavior may be limited by demonstration coverage, and they do not learn from deployment outcomes. Human corrective imitation provides additional recovery examples, but its objective matches local action targets without explicitly optimizing task return. We present ForceRFT, a force-guided residual reinforcement learning framework that learns contact-dependent corrections from human supervision and autonomous task outcomes. A frozen, demonstration-trained SmolVLA-based prior generates force-conditioned action chunks, while a lightweight residual actor refines individual end-effector pose commands using wrist feedback acquired during chunk execution. The decision-time wrist wrench, its temporal change, and the selected base motion condition both residual correction and value estimation. Human corrections supervise the residual actor, while verified autonomous transitions train the twin critics and support value-guided updates to the same actor. Bootstrapping is restricted to autonomous segments, preventing TD credit from crossing human-intervention boundaries. Real-robot experiments on plug insertion, ring-on-peg assembly, and whiteboard wiping show higher autonomous success rates than the evaluated demonstration-trained and residual-imitation baselines. Comparisons with residual imitation support value-guided residual optimization, while plug-insertion ablations indicate the benefit of direct execution-time wrist feedback.
vision-language-actionvla - arxiv:2609.22839 · cs.LGA Horizon-Independent Regret Bound for Optimistic Hedge in General-Sum GamesJunsoo Ha
Can simple learning rules keep their regret bounded in self-play? Recent work achieves constant regret bounds through modified regularization and higher-order prediction. Yet for Optimistic Hedge, arguably the most canonical method in games, the best known individual regret bound remains logarithmic. In this work, we prove that plain Optimistic Hedge with a constant step size can attain $O_{n,d}(1)$ individual regret in general-sum games with $n$ players and $d=(d_1,\ldots,d_n)$ actions, under expected loss-vector feedback. As a corollary, its time-averaged play enjoys an $O_{n,d}(1/T)$ coarse correlated equilibrium (CCE) gap. Our analysis represents Optimistic Hedge as a real-analytic recurrence on a compact space, which yields an exact finite-order difference relation that eliminates horizon dependence. Our proof hinges on nonconstructive Noetherianity argument of Frisch (1967), so the $(n,d)$-dependence remains implicit.
self-play - arxiv:2609.22836 · cs.LGA Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series ForecastingZhihao Lin, Li Lin, Qi Zhang, Kaiwen Xia +2
Time series foundation models (TSFMs) have recently delivered impressive zero-shot performance across diverse forecasting tasks. However, real-world decision-making frequently relies on \emph{irregular multivariate time series} (IMTS), where inconsistent inter-observation intervals and asynchronous sampling across variables coexist with informative missingness. Existing TSFMs handle such inputs either through imputation that injects spurious values or through index-based positional encodings that ignore continuous time. There is still a gap in the foundation model that follows the original IMTS patterns. In this paper, we propose a hybrid attention model that learns a unified time-aware patch representation for IMTS forecasting. We first design a \emph{time-aware patch encoding} that maps a variable number of intra-patch timestamps into a fixed-size embedding, producing a uniform format for irregular patches without resorting to imputation. We then introduce a \emph{time bias attention} mechanism that calibrates inter-patch temporal misalignment and asynchronous cross-channel dependencies as auxiliary attention offset. Finally, on top of a decoder-only Transformer backbone, we adopt a \emph{hybrid causal mask} that preserves a bidirectional full view over the historical context while keeping the forecast horizon strictly autoregressive. To support large-scale pretraining under irregular settings, we also curate VersaTSA, an archive of $30$B observations that retains the native sampling sparsity of its sources. Experiments on three IMTS benchmarks and a standard regular-MTS benchmark show that our model achieves state-of-the-art zero-shot performance on IMTS and remains competitive when transferred to regular forecasting.
benchmark - arxiv:2609.22833 · cs.LGPersonalized Federated Reinforcement Learning via Model-Agnostic Meta-Learning: Convergence of Exact and Hessian-Free Meta-Policy GradientsAli Beikmohammadi, Sarit Khirirat, Sindri Magnússon
We study personalized federated reinforcement learning, in which $n$ agents, each acting in its own Markov decision process, collaborate through a server to learn a shared MAML-style policy initialization that becomes effective for an individual agent once that agent adapts it with a single local policy-gradient step. We propose Per-FedAvg-PG, in which agents take $τ$ local stochastic meta-policy-gradient steps between communication rounds, and prove that it reaches an $\varepsilon$-approximate first-order stationary point of the personalized objective in $K=\mathcal O(\varepsilon^{-3/2})$ rounds with $τ=Θ(\varepsilon^{-1/2})$ local steps. The analysis rests on a structural feature of the reinforcement learning setting: under standard policy-class regularity, the per-agent objectives have uniformly bounded gradients and Hessians with explicit constants, so the bounded-gradient and bounded-heterogeneity conditions imposed by the supervised theory hold automatically and no separate heterogeneity assumption is needed. The exact meta-gradient requires the inner-loop policy Hessian, which our experiments identify as the practical bottleneck. We therefore analyze the Hessian-free variant, bound its bias, and exhibit fixed points at which the meta-gradient is nonzero and of order $α$, showing that the resulting stationarity floor is a property of the method rather than of the bound. Experiments on tabular and neural navigation confirm the predicted behavior and show transfer to unseen agents at an order of magnitude lower sample cost than independent training. Together these results identify the adaptation step size as a tunable personalization knob and the curvature estimate as the quantity that governs whether exact meta-gradients are affordable.
agent - arxiv:2609.22829 · cs.ROWhole-Body UMI: Transferring UMI Manipulation Skills to Humanoid Whole-Body Manipulation via Real-Time Motion GenerationYuxuan Nai, Leixin Chang, Liangjing Yang, Shuo Yang +1
Collecting whole-body demonstrations for humanoid manipulation mostly relies on teleoperation, which is costly and hard to scale up. The Universal Manipulation Interface (UMI) provides a scalable data collection paradigm, but end-effector trajectories alone underdetermine humanoid whole-body coordination, which is insufficient for whole-body demonstration collection. Therefore, we introduce Whole-Body UMI (WB-UMI), a task-agnostic, real-time and end-effector conditioned motion generator that decouples whole-body coordination learning from task semantics learning through a shared end-effector interface. A diffusion policy learns from native UMI demonstrations, while WB-UMI learns independently from retargeted motion capture, requiring no body trackers or paired image--whole-body demonstrations during task-specific data collection. In real deployment, an asynchronous hierarchy integrates the diffusion policy, motion generator, and a whole-body controller with latency compensation and measured-state feedback. Real-robot experiments on G1 support real-time closed-loop transfer across four tasks, achieving 90% success in drawer closing, 80% in shelf pick-and-place, 30% in ball toss, and 40% in Loco-PnP, which shows the effectiveness of this hierarchy in transferring native UMI skills to humanoid whole-body manipulation.
manipulationhumanoidteleoperationdiffusion policywhole-body control - arxiv:2609.22820 · cs.LGBeyond Average Error through Oracle-Informed Stress Tests for Time-Series ForecastingXu Lin, Runheng Zuo, Shengxuan Xu, Qitai Tan +2
Average squared error cannot reveal whether forecasting performance degrades because the future becomes less predictable or because forecasts move farther from the conditional mean. We introduce paired, mechanism-controlled stress tests that decompose changes in expected squared error at each lead time into environmental risk and forecast-oracle distance, using an origin-conditioned predictive oracle unavailable to the evaluated models. Three end-to-end controls have known attribution. Specifically, the null, environmental-only, and information-gap controls verify that the pipeline assigns changes to the correct component. We then apply the benchmark to 24 deployable forecasters. Under frequent switching, 14 methods have higher realized MSE but lower oracle distance; under outlier-variance feedback, 19 have higher MSE but lower scale-standardized MSE. Short- and long-lead stress-response rankings have Spearman correlation 0.624, revealing substantial horizon-dependent reordering. We then study multivariate relation shifts. Across six models and three coupling severities, oracle distance accounts for only 0.7-3.9% of the decomposed expected-risk increase, and environmental-risk majority persists in an eight-channel system and a matched-difficulty audit of Ring, Block, and Hub relations. Finally, prespecified contrasts on independent data-generating process (DGP) realizations show that several visually compelling discovery profiles, including trend accumulation and the hypothesized switching reversal, do not replicate. The benchmark thus combines component-wise diagnosis with a held-out stability audit. It complements real-data out-of-distribution evaluation, which measures performance under realistic shifts when exact oracle attribution is unavailable.
benchmark - arxiv:2609.22819 · cs.LGCounterfactual Tool Ranking under Utility, Cost, and Privilege ConstraintsJiapeng Li
Counterfactual tool evaluation must distinguish authority, historical support, and what a comparison actually estimates. We study these distinctions with eleven executable enterprise-inspired tools, exact-propensity logs, and real local Model Context Protocol transport. An initial 45-run synthetic study is retained, then challenged by 30 realized-return control runs and 15 experiments on 1,930 independently released Berkeley Function Calling Leaderboard (BFCL) tasks. Full-return direct regression reverses an initially favorable doubly robust (DR) evaluation result in the linear setting: mean absolute errors are 0.0139 for direct regression and 0.0272 for DR. Under a shifted environment, DR retains an advantage, with errors 0.0227 versus 0.0948. On function-name-group-disjoint BFCL-derived splits, direct and DR selectors obtain balanced accuracies of 81.85% and 79.83%. Two pinned local Qwen2.5 models are evaluated on the same 200 held-out tasks, exposing a strong failure to abstain under the fixed prompt. We further characterize policy differences under missing support: unsupported actions shared by two policies cancel, allowing point identification of an incremental change when neither absolute value is identifiable. A disagreement-preserving fallback achieves this property in all five support-gap runs, but conservative sampling bounds do not certify deployment improvement. The contribution is a falsifiable evaluation method and independent public evidence, not a new DR estimator, official BFCL leaderboard score, or production-agent safety claim.
leaderboard - arxiv:2609.22818 · cs.AIThe Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM AgentsPritom Bhowmik
Memory-poisoning defenses for LLM agents are typically evaluated by their ability to prevent attacks. However, the traffic they process is rarely adversarial. The cost of implementing a defense is paid with each interaction, while its benefits are only seen in a small percentage of cases. We developed a measurement setup that keeps the memory backend, retrieval process, and judge consistent across different conditions, changing only the defense itself. We test each condition three times across five conversations to distinguish the defense's real effects from noise inherent in the pipeline's runs, which remains significant even at temperature zero. Across three write-time defenses (input sanitization, provenance checking, and LLM-based anomaly detection) and one read-time defense (reranking), tested on entirely benign traffic, the write-time defenses show no utility cost we can resolve, with 95% confidence intervals spanning roughly +/-4.5 points and including zero. The reranker is different: it lowers core accuracy by 4.4 points (95% CI [-9.0,-0.05], bootstrap; McNemar p=0.064), a result that survives replication but sits at the edge of our resolution. Its clearer cost is mechanical rather than statistical. On conversations containing no attack, the reranker quarantines legitimate memories on 33.6% of adjudicated items, reaching as many as 106 false quarantines in a single conversation, at 2.7% token overhead. Stacking all four defenses does not compound this cost: the combined condition's accuracy loss is smaller, and its confidence interval includes zero, suggesting the write-time defenses may partly offset what the reranker discards. Where a defense intercepts the pipeline, not whether it uses an LLM, appears to determine its benign-case price.
memoryllm agent - arxiv:2609.22816 · cs.LGFIRM-WM: State-factorized factual-interventional recurrent modeling for reward-free visual planningYilun Wu, Yunjian Zhang, Aobo Li, Mujiangshan Wang +2
Reward-free latent world models can learn from offline videos and solve new image--goal tasks by optimizing actions through predicted latent futures. This setting places two demands on the planning state: its coordinates must be comparable with a goal image. Moreover, its dynamics must retain velocity, motion trend, contact, and other history--dependent information beyond those goal coordinates. Offline training creates a second mismatch: each recorded trajectory reveals one factual future, whereas a sampling--based planner compares many actions that were not taken from the same state. We introduce FIRM-WM (Factual--Interventional Recurrent World Model), a compact pixel world model designed around these two gaps. Its recurrent state separates a typed, goal--comparable configuration from a 128-dimensional dynamic fiber used for prediction but excluded from the terminal goal cost. Broad factual trajectories provide state coverage, while common--reset intervention branches provide observed outcomes for alternative action sequences. Before executing each branch, we reset the environment and restore the same recorded values exposed by the environment's state--setting interface. Under matched CEM planning and three independent full-pipeline seeds, FIRM-WM reaches 99.0$\pm$1.0% on TwoRoom, 92.7$\pm$2.1% on Reacher, and 88.0$\pm$3.0% on OGBench-Cube, compared with 89.0%, 88.0%, and 70.0% for LeWM. The deployed model uses 2.98--3.42M parameters and records 2.13--11.60$\times$ lower planning time on these tasks.
world model - arxiv:2609.22813 · cs.ROCommonsense-Grounded Path Planning from Abstract InstructionsMasafumi Endo, Kohei Honda, Ryo Yonetani
We present \emph{commonsense ranked search} (CoRS), a novel path planner that turns an abstract instruction into a route that follows commonsense. While existing methods respect the considerations written down in advance, a robot working among people must follow those left unstated too, as with a wet floor that a worker avoids without being told. CoRS leverages large language models (LLMs) and vision-language models (VLMs) as commonsense knowledge to reason about these latent considerations in its planning. Given an abstract instruction (\emph{e.g.}, ``move carefully'') and visual observations of each region in the environment, CoRS derives a consideration for each region, as in ``this wet floor is slippery and worth a detour.'' It then compares the considerations between regions to see which of the two the robot should avoid more, as in ``the crowd is worse than the wet floor.'' These judgments sort the regions into a commonsense ranking, whose costs drive a conventional search that always returns a valid route. We build a benchmark for planning under latent considerations, with three environments, 1350 problems, and five instructions at three levels of abstraction. Experiments show that CoRS discovers the unstated considerations and goes around the ones worth a detour while crossing the rest, a behavior that recent LLM-based planners do not achieve.
benchmark - arxiv:2609.22809 · cs.ROKinematic Interface for the Wild: Modular Bimanual Loco-Manipulation Capture from 360$^{\circ}$ Cameras AloneBenjamin Yang, Weiying Wang, Shenggao Li, Keming Yan +4
A wrist-mounted camera for UMI-style data collection must do two jobs: record the manipulation and localize in the scene. Most handheld devices localize online from workspace-facing views crowded by hands and objects, or add dedicated tracking hardware. Room-scale bimanual capture therefore still tends to instrument the operator or the scene for accurate localization. We present KIWI (Kinematic Interface for the Wild), a capture kit whose only electronics are off-the-shelf cameras. Our core system splits the two jobs across the two lenses of a 360-degree camera. The rear lens faces the room and builds a shared metric map that registers both hands, and an optional head camera, in one frame without workspace co-visibility; the front lens records the manipulation, and offline IMU fusion bridges front-lens tracking loss. Through our quick-release plate, the camera module attaches to chopstick grippers, parallel-jaw grippers, hand-wrist mounts, or robot flanges. Across six bimanual recordings, combining the rear and front lenses failed to localize only 0.1% of query frames, whereas front-only bimanual feature alignment failed on 24.8% of frames and lost one recording entirely; against evaluation fiducials, localization error stayed within 4.5 mm. KIWI's recovered poses were sufficiently consistent for the four wrist streams alone to reconstruct the scene as a 3D Gaussian splat.
manipulationgripper - arxiv:2609.22805 · cs.CLTo Consolidate or not to Consolidate? Evaluating the Impact of Consolidation in Multi-Reference Training using Peer ReviewsMaitreya Prafulla Chitale, Ketaki Mangesh Shetye, Yash More, Harshit Gupta +3
Natural language generation (NLG) tasks span the spectrum of conditional entropy, ranging from highly constrained machine translation to open-ended dialogue generation. Structured tasks like automated peer-review generation occupy the intermediate region, where a single input admits multiple valid, overlapping outputs. In this work, we demonstrate that traditional single- and multi-reference training paradigms are suboptimal for these intermediary tasks. We provide empirical evidence that consolidating diverse references into a unified training signal is crucial for developing effective systems. To facilitate this, we introduce MERC-36K, a large-scale corpus of over 36,000 papers paired with original and consolidated peer reviews. Using this dataset, we train specific architectures to isolate the impact of different reference paradigms and benchmark against existing state-of-the-art systems. Through extensive automatic and human evaluation, we demonstrate that models trained on consolidated references significantly outperform those trained on unconsolidated references. Dataset and code will be released upon acceptance.
benchmark - arxiv:2609.22803 · cs.ROARCGym: Benchmarking Deep Reinforcement Learning in Autonomous Robotic ColonoscopyGuanglin Ji, Martina Finocchiaro, Kenny Erleben, Hang Yin
Simulations for learning-based autonomous colonoscopic navigation focus mainly on fully actuated capsule robots, failing to capture the contact-rich navigation of long and flexible clinical colonoscopes. We present the Autonomous Robotic Colonoscopy Gym (ARCGym), an open-source reinforcement learning environment and benchmark for image-based navigation in clinically derived deformable colon anatomies. ARCGym supports multiple types of colonoscope robots, spanning capsule robots and flexible endoscopes, with this work focusing on flexible endoscopes including magnetic-driven tip actuation and clinically used proximally translational actuation. This work includes five CT-reconstructed colons representing typical clinical scenarios, a set of clinically meaningful navigation subtasks, and unified success metrics. We introduce a reward combining depth-based lumen alignment with a lumen-visibility score to improve learning under occlusions. Experiments across tasks, robots, and anatomies show that autonomous navigation remains challenging for both magnetic-driven and proximal-insertion flexible robots, with proximal-insertion actuation remaining an open problem.
benchmark - arxiv:2609.22798 · cs.ROExpert-Play Contouring Control: Faster-than-Demonstration Planning from Slow Expert and Fast PlaySeunghoon Cho, Wonsuhk Jung, Sundhar Vinodh Sangeetha, Shreyas Kousik
Expert demonstrations often specify what a robot should do, but not how fast it can do it. Imitation Learning (IL) inherits demonstration timing, while directly accelerating the learned motion can fail when faster execution changes the robot-object dynamics. We study faster-than-demonstration execution as a dynamics-aware control problem and introduce Expert-Play Contouring Control (EPCC), which combines slow expert demonstrations with fast, non-expert play. Expert demonstrations train a latent trajectory generator whose predictions are reparameterized into a time-independent contour of successful task progression, while play trains a world model (WM) of fast-action outcomes. At deployment, our proposed planner uses the WM to optimize actions that makes maximize progress along the expert-derived contour while penalizing deviation from the intended task evolution. Averaged across three visuomotor manipulation tasks, EPCC achieves a $2.0\times$ the throughput of the IL baseline, including $2.2\times$ that of the throughput of the strongest acceleration baseline on a task with interaction-sensitive object dynamics. Our analysis shows that the gains concentrate where faster execution changes robot-object evolution. Together, our results highlight a simple yet effective principle: demonstrations provide task intent, while play data provides the dynamic coverage needed to execute that intent faster.
manipulationworld model - arxiv:2609.22793 · cs.AIDiagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine TranslationJi Hun Wang, Siyu Wu
LLM-based machine translation evaluation can closely match human judgments, but in practice it remains largely diagnostic, with the signals rarely translating into direct quality improvements under real production constraints. We propose a two-stage, evaluator-guided automatic post-editing framework that turns MQM-style evaluation into targeted repairs: a retrieval-augmented LLM evaluator outputs structured, span-level MQM diagnoses under an explicit edit contract, and a separate LLM post-editor applies minimal edits restricted to those diagnoses. This separation improves controllability and reduces paraphrastic drift compared to one-stage "judge-and-refine" baselines. In a systematic study involving seven LLMs spanning three model providers and seven languages, our best configuration consistently improves both COMET-22 and COMETKiwi scores over one-stage post-edit methods, while the evaluator's error spans and severities show strong agreement with human MQM annotations and human editor preferences.
retrieval-augmentedevaluator - arxiv:2609.22792 · cs.AISelfOp: An Optimization Algorithm for Self-Improving Security AgentsSaad Ullah, Yigitcan Kaya, Christopher Kruegel, Giovanni Vigna +1
LLM agents are increasingly used for security tasks: vulnerability discovery, exploit reproduction, and patch generation. Improving them at the model level demands expert demonstrations or computable rewards, which security tasks rarely offer: traces are costly, failures hard to diagnose, rewards sparse, and non-computable. Efforts thus shift to the harness and context, but manual tuning needs task-specific expertise and scales poorly, while automated methods rely on scarce ground truth, stronger optimizer models, or unguided propose-and-evaluate loops that reduce to costly trial and error. We introduce SelfOp, an algorithm that automatically improves a frozen security agent's task context (instructions, skills, and reference documents), without modifying its execution harness and model weights. SelfOp casts context optimization as chain-rule-inspired textual gradient descent: from a single instance's outcome, it propagates error signals backward through the evaluator, the agent's trajectory, and the context artifacts that shaped its behavior, yielding per-instance textual gradients. Gradients are accumulated across instances by clustering, ranking, and filtering, and committed only under cross-instance consensus. A convergence detector monitors the gradient signal itself and stops once the context has absorbed the generalizable information in the training data, without held-out validation data. We evaluate SelfOp on CyberGym, a benchmark of real-world vulnerability reproduction tasks. With fewer than 200 training examples, SelfOp yields a 17-point self-improvement for GPT-5.4-mini (with Codex), enough to surpass the frontier GPT-5.4 baseline by 6 points, and an 18.5-point self-improvement for GPT-5.4 itself. The optimized skills also transfer across models, highlighting that SelfOp-optimized skills learn generalizable task knowledge not model-specific patterns.
llm agentself-improvingself-improvementbenchmarkevaluator - arxiv:2609.22789 · cs.CVPixelART: Image-to-Layer Decomposition without Latents or Text-to-Image PretrainingZelin Jia, Zhao Zhang, Zhicong Tang, Yuhui Yuan +1
Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusion-based systems typically adapt large pretrained text-to-image (T2I) models and introduce RGBA autoencoders or variable-layer architectural modules. We revisit this design choice and ask whether layer decomposition truly requires these heavyweight components. We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer (I2L) decomposition. PixelART directly denoises regional RGBA pixel patches using a single-stream multi-modal diffusion Transformer, avoiding RGBA-VAEs, pretrained T2I backbones, and layer-specific decoders. We identify a key property of the task: high-noise timesteps determine layer assignment and coarse layer organization, while low-noise timesteps mainly refine color, alpha, texture, and boundaries. Based on this observation, we propose a terminal-boosted timestep sampling strategy to increase training coverage in the high-noise layer assignment regime. Trained on 4M multi-layer design templates, PixelART achieves state-of-the-art layer decomposition and composite reconstruction on Design-Multi-Layer-Bench and LICA with over 80% fewer parameters, 98% lower latency, and 85% lower memory than the recent Qwen-Image-Layered model. Ablation experiments show that pixel-space $\mathbf{x}$-prediction, terminal-boosted timestep sampling, and data/model scaling are critical, while T2I pretraining brings marginal benefits to the I2L task.
memory - arxiv:2609.22788 · cs.CVHuman-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical ReasoningFanhong Li, Shurui Zheng, Zi Yin, Junbo Cui +2
Video foundation models now reach human-level accuracy on physical-reasoning benchmarks, yet such tasks require predicting unobserved physical outcomes. Do these models perform human-like forward simulation, or do they exploit statistical regularities in visible scenes? Accuracy alone cannot distinguish these strategies. We introduce a distributional evaluation framework that treats model seeds and human raters as populations, enabling comparison of consensus, uncertainty, and strategy. On the Physion benchmark, we evaluate three ViT-L architectures (V-JEPA2, VideoMAEv2, DINOv2). V-JEPA2 narrows the accuracy gap to ~1 percentage point (73.2% vs. 74.2%), yet model-human disagreement reaches 26.4%, far exceeding human-human disagreement (4.8%), with substantially lower agreement (kappa ~ 0.48 vs. 0.91). The divergence follows forward-simulation demands: models outperform humans on geometric reasoning (linking, +11.8 pp) but underperform on gravitational dynamics (rolling, -11.8 pp) and causal chains (dominoes, -10.5 pp). Strategy fingerprinting confirms all three architectures share non-human strategies while none aligns with humans. Attribution analysis suggests that unobservable outcome features, rather than visible scene properties, predict this divergence, consistent with models relying more on scene-level statistical regularities than on explicit forward simulation, a systematic divergence that accuracy alone cannot reveal. Code is available at https://github.com/fanhong-li/model-human-divergence.
v-jepabenchmarkevaluation framework - arxiv:2609.22785 · cs.LGRobust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement LearningHao Yang, Zhenguo Xu
Market-making strategies in real limit order book markets face substantial model uncertainty and regime-shift risk. Existing adversarial reinforcement learning approaches improve robustness by formulating the Avellaneda--Stoikov market-making problem as a zero-sum game between a market maker and an environmental adversary. However, these approaches typically rely on Poisson order arrivals and neglect trade-induced price impact, limiting their ability to capture important high-frequency market microstructure effects such as clustered order flow, self-excitation, and post-trade price feedback. We extend adversarial reinforcement learning for market making to a more complex environment with Hawkes self-exciting order arrivals and trade-induced price impact. To mitigate the increased non-stationarity introduced by the expanded regime space, we incorporate an LSTM module that explicitly models the temporal structure of recent observations. We further characterize the equilibrium properties of the proposed framework through both game-theoretic analysis and numerical experiments, and introduce a robustness evaluation protocol focused on improvements in the left tail of the return distribution. Experimental results across a range of market regimes show that the proposed method achieves improved left-tail performance in most complex microstructure environments. In particular, the gains are pronounced in regimes with strong Hawkes excitation and low-to-moderate price impact. Bootstrap tests provide no evidence that these improvements are obtained through a stronger terminal directional inventory bias. These results suggest that combining adversarial training with temporal state representation can improve the robustness of reinforcement-learning-based market-making strategies under order-flow self-excitation, price impact, and regime uncertainty.
evaluation protocol - arxiv:2609.22778 · cs.CLMIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills AssessmentYuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn +1
Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require expert judgment remains unclear. We investigate this challenge in the context of assessing psychotherapeutic interpersonal skills and introduce MIS-Bench, a Multimodal Interpersonal Skills (MIS) benchmark comprising 996 psychotherapy response videos annotated across 8 dimensions of Facilitative Interpersonal Skills. Across 9 MLLMs with multiple modality and prompting settings, we find that current models show only modest agreement with human experts, inconsistent gains from multimodal input, and limited benefits from reasoning-based prompting. To mitigate this gap, we propose MIS-RAFT, a regression-aware fine-tuning method inspired by RAFT and tailored to fine-grained interpersonal skill scoring at one-decimal precision. MIS-RAFT addresses the mismatch between autoregressive token prediction and scalar-valued expert assessment, significantly improving agreement with human ratings. Overall, MIS-Bench reveals a clear gap between general multimodal capability and expert-level interpersonal judgment, while MIS-RAFT offers a promising path toward more reliable model-based assessment.
benchmarkevaluator - arxiv:2609.22774 · cs.CLNLPCC 2026 Task 10: Citation-Level Faithfulness Verification with DeBERTa Ensembles and Class-Wise CalibrationYanling Li, Zirui Li, Mingyu Wan
This paper presents our system for Track 2 of the NLPCC 2026 Shared Task 10 on citation-level faithfulness in AI-assisted scientific reporting. Given an atomic scientific claim and the structured full text of its cited paper, the task requires both a four-way relation label and up to three evidence paragraph identifiers. The label head ensembles a paragraph-aware cross-encoder with a document-level DeBERTa-large classifier, followed by class-wise decision calibration. Probability-level fusion is motivated by an out-of-fold tendency to over-predict Topical Match. The evidence head combines paragraph scores from top-20 and top-30 joint models with BM25 scores. The system runs fully offline without external retrieval or LLM prompting. On the final leaderboard, our system achieved 82.9898 overall (89.5491 Macro-F1 and 76.4305 Joint@3), ranking second in Track 2. Ablations and error analysis show that model complementarity and calibration drive the label gains. Gold-evidence inference changes label Macro-F1 negligibly, whereas evidence ranking remains important for Joint@3.
leaderboard - arxiv:2609.22771 · cs.LGParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech UnderstandingNishit Anand, Jiaqi Su, Ke Chen, Yunyun Wang +4
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.
benchmark - arxiv:2609.22767 · cs.LGBeyond Final-Token Classification: Heterogeneous Readouts for Evidence-Grounded Suicide Risk DetectionZirui Li, Yanling Li, Kaolanglang Gao
The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label psychosocial factor detection, and extraction of supporting phrases. We introduce heterogeneous readout decomposition (HRD), which separates semantic verification from output realization. A locally deployed Qwen3.8-27B model, adapted with task-specific QLoRA adapters, produces both answer-token margins and layer-63 answer states for card-conditioned queries. HRD compares four latent scores for ordinal risk, retains token margins for most factors while routing seven labels through one shared latent probe, and constructs evidence sets from verbatim span candidates with calibrated, risk-conditional constraints. On two held-out user-grouped confirmation folds, the latent risk readout improves weighted F1 from 0.8237 to 0.8372 and macro F1 from 0.7965 to 0.8185. Selective factor routing improves macro F1 by 0.0105 and tail-label macro F1 by 0.0189; in contrast, global latent replacement and independent label-specific probes fail. The constrained evidence decoder raises pooled phrase F1 from 0.7488 to 0.7609 in row-level out-of-fold evaluation. Lenormand's best public result is 0.8052 on Subtask 1 and 0.6636 on Subtask 2, giving a composite score of 0.7627. These results identify the answer-token boundary, rather than semantic representation alone, as a measurable source of error in this benchmark.
benchmark - arxiv:2609.22762 · cs.CVDriveReferee: Geometric Safety Verdicts Need Not Be Learned for Driving World-Action ModelsFengcheng Yu, Dhruv Parikh, Junjie Ye, Maulik Bhatt +4
Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expert imitation. Yet imitation provides no explicit closed-loop geometric verdict for generated trajectories, making verification important during both training and deployment. Closed-loop evaluators can check collision and drivable-area violations, but require privileged scene state unavailable at deployment. Existing approaches often close this gap by learning a verifier from sensor features. For these geometric checks, the rule itself is explicit. For example, collision is determined by whether the rolled-out ego footprint overlaps occupied vehicle space. What is unavailable at deployment is the scene state needed to apply the rule. We introduce DriveReferee, which uses a learned geometry readout to predict the scene representation from camera observations and executes the geometric safety rule directly rather than learning it. The resulting analytic referee evaluates collision and drivable-area safety from a scene state and candidate trajectory. During training, it scores self-sampled trajectories on ground-truth state and distills the resulting preferences into the WAM policy. At deployment, the same referee evaluates generated trajectories on this predicted state and selects a safer alternative when needed. The analytic referee requires no verdict-specific training, and its decisions follow an explicit geometric rule. Under matched candidates and inference budgets, it matches or outperforms all learned-verifier and heuristic baselines. Given the same predicted state and trajectory, learning the verdict provides no measurable downstream gain despite requiring tens of thousands of evaluator-labeled training examples. On the full NAVSIM navtest, DriveReferee reaches 92.02 PDMS with single-camera visual input and no external training data.
evaluator - arxiv:2609.22750 · cs.LGTowards Robust Classroom Attendance: A Comprehensive Evaluation of Face Detection and Recognition ModelsHimani Trivedi, Hiren Patel, Ridham Patel, Krutika Patel +1
Manual attendance methods, such as paper or register-based systems, take a lot of time, can lead to errors, and are easy to falsify. Face recognition is more reliable, but it frequently struggles in classrooms because lighting and other conditions can vary. Face recognition datasets are designed for regulated environments and do not capture the actual challenges found in classrooms. To address this, a new face detection and recognition dataset, the Visage Face dataset, comprising 16,234 face samples, is proposed for the task of face detection and recognition. The photos are taken from different angles and under varying lighting conditions, with students showing a range of expressions, and some faces partly covered to reflect real-life situations. A YOLO-based system is used to detect faces and tested seven advanced face recognition models with thirteen configurations: LVFace, QCFace, FaceLiVTv2, TopoFR, EdgeFace, TransFace, and GhostFaceNets. Of these, FaceLiVTv2-M performed best, with 99.75% Top-1/Top-5 accuracy and an inference time of 6.459 ms. These results show that the Visage Face Dataset is a realistic and challenging benchmark for face recognition in classroom attendance.
benchmark - arxiv:2609.22747 · cs.AIBeyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at NetflixChaoran Guo, Ding Tong, Ting-Po Lee, Scarlet Chen
Understanding the performance of large-scale recommender systems remains an underexplored challenge, especially for content creators and model developers. The raw engagement signals available to them, such as views and clicks, conflate content quality, model behavior, presentation bias, and audience reach, making it hard to attribute outcomes to the right cause. In this work, we present a general evaluation framework that enhances observability across multiple recommender systems at Netflix and demonstrate its effectiveness through several production deployments. The framework treats recommender-system observability as a counterfactual measurement problem: estimating what the recommender would have done, and what engagement would have followed, in the absence of a specific content item or model decision. We articulate three stakeholder-centered observability principles for content creators and model developers, and propose measurement methodologies covering bias reduction, relativity, and incrementality, applicable to both single-stage and cascading recommender systems and serving both audiences from a single measurement foundation.
evaluation framework - arxiv:2609.22730 · cs.ROBEACON: Belief-Enabled Adaptive CONtrol for Imitation Learning under UncertaintyMoonyoung Lee, Soumojit Bhattacharya, George Kantor, Oliver Kroemer
Robot manipulation tasks often involve hidden state information that cannot be directly observed and must be inferred through sequential physical interactions. In such partially observable settings, conditioning an imitation learning policy directly on the recent raw observation history leads to poor performance. This is due to state aliasing, wherein identical observations may arise from different hidden states, and the policy receives conflicting action labels for the same input. To enable history-aware disambiguation capability, we propose conditioning a diffusion policy on a structured representation of the hidden states using Bayesian belief that exposes both the current most likely state estimate and the remaining uncertainty. This representation replaces raw history with a structured, compact input, enabling the policy to implicitly modulate between exploratory and exploitative behaviors based on belief uncertainty, without explicit mode switching or reward shaping. We evaluate across two domains with qualitatively different belief representations: a continuous belief for cornstalk gripper alignment via tactile sensing, and a discrete categorical distribution for latched door opening. In both domains, the belief-conditioned policy substantially outperforms the observation-only baseline and approaches privileged ground-truth performance, with ablations illustrating that the policy adapts its exploration behavior depending on the belief uncertainty at inference time.
manipulationtactilediffusion policygripper - arxiv:2609.22724 · cs.AIMATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory LearningChangyue Jiang, Jiayi Wang, Xin Wen, Jiarun Dai +2
Mobile agents powered by foundation models now automate complex, multi-step workflows on real devices, but their trajectories can violate app-specific security policies. Existing trajectory-level defenses rely on LLM prompting or rigid rules, and thus fail to support fine-grained, natural-language policies that generalize across apps and tasks. In this work, we introduce MATE, a lightweight, policy-conditioned auditor that encodes both agent trajectories and natural-language security policies to determine whether a trajectory violates a given policy and to explain why. Treating policies as editable text rather than fixed model parameters allows MATE to handle user-defined and evolving requirements without retraining. To construct MATE, we build a knowledge base by extracting app descriptions, workflows, and policies from hundreds of popular mobile apps worldwide, and synthesizing over 140K semantically realistic, policy-conditioned trajectories with a multi-stage pipeline. We further release MATEBench, a trajectory-level auditing benchmark with two synthetic subsets and one real-world subset of manually collected trajectories. Models trained with our synthesis-driven trajectory learning achieve over 95% accuracy on MATEBench, retain strong performance on external safety benchmarks, and audit trajectories from Zhipu's AutoGLM and Alibaba's Mobile-Agent on real devices with over 95% accuracy, outperforming prior methods by over 20%. MATE shows that practical, fine-grained security auditing for heterogeneous mobile agents is both feasible and effective.
agentbenchmark - arxiv:2609.22719 · cs.AIFrom Code to Requirements: Agentic Reverse Engineering of Business Rules at Enterprise ScaleGarima Agrawal, Prasun Das, Priyanka L, Minisha N +4
Business requirements for enterprise software systems are rarely captured in structured form; the logic resides instead in source code, configuration files, and institutional memory. When these systems must be migrated, extended, or audited, the absence of formal requirements artifacts forces teams into expensive, knowledge-losing manual reverse engineering. This paper presents an agentic framework that autonomously generates Business Requirements Documents (bRDs) through reverse engineering of undocumented enterprise software. Seven specialized agents collaborate to discover user journeys, extract business rules, and synthesize a locked bRD from code, test suites, configuration, and available documentation, supported by static analysis tools, guided by embedded expert practitioner cognitive models, and overseen by human reviewers at controlled escalation points. Measured on actual execution traces across a real enterprise deployment, the framework generates comprehensive bRDs in under nine minutes per service, with extracted rules independently corroborated against real production defect records. In a comparable prior migration to the same target architecture, delivery required over two years; on the present programme the framework achieves a cost reduction exceeding 98% against a baseline derived from actual repository metrics using IFPUG complexity models and industry benchmark labor rates.
agenticbenchmark - arxiv:2609.22716 · cs.CVZIL: Zero-shot Image-to-LiDAR RegistrationZijun Li, Xiaotian Sun, Xuelun Shen, Yao Dai +5
Image-to-LiDAR registration estimates the camera pose of an image with respect to a LiDAR point cloud. It has diverse applications in autonomous driving, robot navigation etc. However, state-of-the-art (SOTA) methods still 1) mostly assume same-frame inputs, struggling with the image and point cloud from distant frames; 2) rely on domain-specific training, failing to generalize to unseen scenarios. We propose ZIL, the first foundation model for zero-shot non-synchronized image-to-LiDAR registration. ZIL encodes the input image and point cloud with the Vision and Point Transformers. In addition to regressing the relative pose, ZIL also learns to predict 3D coordinates, which substantially improves the pose accuracy without additional annotations. Interestingly, naive mix-data training cannot enable zero-shot generalization, which requires normalization on both camera intrinsics and the LiDAR vertical-axis origin. Trained on 7 public datasets with 1.4M LiDAR frames, ZIL consistently and significantly outperforms previous SOTA with a single model across 5 in-domain and zero-shot benchmarks, reducing the translation and rotation errors by up to 87% and 76% (shown in Fig. 1). Code and models are available at https://github.com/ZijunLi7/ZIL.
benchmark - arxiv:2609.22712 · cs.AITrustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM SystemsFayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh, Aakriti Adhikari
Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and operational risks. Untrusted content from websites, emails, documents, and databases can enter the same context as system instructions; persistent memory can carry compromised information across sessions; and access to external tools can turn an incorrect model response into a consequential real-world action. This article reviews the trustworthiness of agentic AI across five interconnected dimensions: safety and robustness, alignment and human oversight, transparency and auditability, privacy and data governance, and regulatory compliance. It organizes key failure modes, including indirect prompt injection, backdoor triggers, goal misgeneralization, memory contamination, and cross-session data leakage, into a unified taxonomy. It also examines major mitigation approaches, such as instruction hierarchies, context isolation, spotlighting, process-based supervision, constrained tool use, and privacy-preserving memory, while distinguishing techniques supported by empirical evidence from those that remain largely conceptual. Building on this analysis, we introduce the Trustworthy Agent Development Lifecycle (TADL), a six-phase framework covering specification, design, training, evaluation, deployment, and monitoring. For each phase, TADL identifies relevant trust activities, expected evidence, and risk-based decision gates. Although TADL has not yet been empirically validated, it provides a structured foundation for developing and evaluating more secure and accountable agentic systems. The article concludes by identifying gaps in current benchmarks and outlining priorities for future research.
memorypersistent memoryagentagentictool usebenchmark - arxiv:2609.22706 · cs.CVDOA-SORT: Directional Occlusion-Aware Multi-Object Tracking with Distributional ObservationsHao Wang
Identity association in multi-object tracking (MOT) is vulnerable to partial occlusion, truncated detections, and fluctuating confidence scores. Existing motion-dominant trackers commonly represent occlusion as a scalar penalty. This treatment misses the directional observation bias caused by occlusion: left, right, top, and bottom occlusions distort the location and shape of a detection in different ways. We propose \ours{} (Directional Occlusion-Aware SORT), an online and training-free tracker that models these biases explicitly. First, it infers a soft front--back ordering from box overlap and relative bottom positions, and estimates directional occlusion coverage and depth. It then constructs a mixture of one clean and four directional occlusion observation components. The model uses a five-dimensional observation comprising box center, area, confidence, and aspect ratio, and adapts observation noise to predicted occlusion and detection confidence. The directional mixture likelihood is used in high-confidence association, low-confidence association, and track recovery; ambiguity penalties and local order-consistency swaps further reduce identity errors among nearby objects. On the DanceTrack validation split, \ours{} improves HOTA from 63.00 to 66.34, AssA from 45.10 to 49.57, and IDF1 from 62.19 to 65.28 over OA-SORT with the same detector and evaluation protocol. The gains are concentrated in association quality while detection accuracy remains stable. Additional local evaluations on MOT17 and MOT20 train splits characterize cross-dataset behavior under the same no-ReID tracking protocol.
evaluation protocol - arxiv:2609.22705 · cs.AIAnalyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube CommentsJakob Morales, Monica Hegde, Fayeq Jeelani Syed
Online discourse about urban issues - walkability, cycling infrastructure, public transit, housing density, and street safety - is voluminous but unstructured, and existing city-evaluation tools capture none of it. We present a pipeline and conversational system that combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG) over 22,788 chunks of YouTube transcripts and comments spanning 309 North American cities. Beyond the system itself, our contribution is a set of measurements about what happens when standard NLP components meet short, informal, geographically ambiguous text. A Twitter-tuned RoBERTa classifier outperforms a VADER lexicon baseline by 12.6 macro-F1 points (0.589 vs. 0.464; McNemar p = 0.0001), but both models collapse on the neutral class, which dominates urbanist comment traffic; annotators disagree on the same class (Cohen's kappa = 0.53). Dense retrieval beats a TF-IDF baseline at every cutoff (P@5 0.790 vs. 0.560), and video-level relevance proxies understate chunk-level precision by a wide margin (0.660 vs. 0.94 under human rating). For groundedness evaluation, we find BERTScore unusable when a multi-sentence generated summary is compared against a single short comment - scores are nearly flat regardless of relevance - and show that ROUGE-1-based groundedness is a paraphrase-driven lower bound rather than a hallucination rate. These findings generalize beyond the urbanist domain to any RAG system built over short user-generated documents.
retrieval-augmentedrag - arxiv:2609.22704 · eess.SYRecursive Parameter Identification of Nonlinear Stochastic State-Space Models via Sequentialized Ensemble Kalman InversionFarzaneh Barat, Rachel Carter, Sara Wilson, Huazhen Fang
Recursive parameter identification in nonlinear stochastic state-space models is challenging because unknown parameters affect the measurements through latent-state dynamics. This paper develops a sequentialized ensemble Kalman inversion method for recursive parameter identification from streaming measurements. The proposed method represents the parameter posterior by an evolving ensemble and updates it sequentially as new measurements become available. For each parameter ensemble member, implicit particle filtering is used to approximate the predictive observation statistics required for parameter correction. This formulation enables recursive parameter learning. The method is evaluated on a strongly nonlinear benchmark, with comparison to several existing methods, and then on a nonlinear double-capacitor lithium-ion battery model. The numerical results demonstrate accurate recursive parameter identification and latent-state estimation.
benchmark - arxiv:2609.22700 · cs.AILLaDA-PRM: A Bidirectional Step-Level Reasoning EvaluatorYiming Feng, Naihao Deng, Yulong Chen, Rada Mihalcea
Step-level reasoning evaluators are commonly based on autoregressive language models, whose causal attention restricts each step representation to the problem, previous steps, and the current step. Yet, when the complete solution is available, the validity of an earlier step may become clearer only through its downstream consequences. We validate this hypothesis through a controlled 54-run comparison of causal and bidirectional LLaDA evaluators at 1B--3B scale, changing only the self-attention mask, and find bidirectional attention yields consistent improvements. Building on this finding, we introduce \prm{}, an 8B bidirectional evaluator that reaches 88.8 step-level F1 on MR-MATH-invalid and 83.8 on the out-of-distribution MR-GSM8K original-question subset, outperforming ReasonEval-Llemma-34B by 11.3 and 10.3 F1 points, respectively. \prm{} also remains effective when evaluating incomplete reasoning traces in online settings, outperforming the strongest baselines on both benchmarks by a large margin. We further show that \prm{} provides an effective training-data selection signal, improving Mistral-7B performance on MATH-500.
benchmarkevaluator - arxiv:2609.22697 · cs.CLCOT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought ReasoningWeizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan +8
Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational context. Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task. Given historical conversation audio, target text, and a reference speech, the system should comprehend the conversational context, infer an explicit intermediate reasoning, and finally synthesize the target speech with the specified timbre. To support this task, we constructed a large-scale bilingual conversational speech dataset comprising 9 million training samples, including a high-quality subset of 1 million samples. We further constructed a source-disjoint benchmark with 800 human-verified samples and established strong task-specific baselines. Additionally, we developed end-to-end autoregressive models with parameter sizes of 0.6B and 1.7B, generating emotion-labeled transcripts, editable speech style inferences, and speech tokens. Experimental results show that the proposed model achieves performance comparable to large-scale baseline systems with significantly fewer parameters. At the same time, the model performs well in terms of duration consistency and emotional consistency, and can generate appropriate emotional, stress, and rhythmic variations based on the conversational context. To facilitate future research, we will publicly release the data construction pipeline, dataset, trained models, and related resources. The demo page and additional resources are available at https://luckybian.github.io/COT-TTS
benchmark - arxiv:2609.22696 · cs.AIBuilding Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision FrameworkGaurab Chhetri, Anandi Dutta, Subasish Das
Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation, labeling, and deployment responsibilities are distributed across multiple technical and governance layers. This paper presents a validation-aware weak-supervision system for constructing and evaluating suicidal ideation (SI) and broader mental health (MH) disclosure benchmarks on Bluesky, a decentralized social media platform built on the AT Protocol. The system integrates public firehose collection, task-specific lexicon filtering, Llama-3-8B-assisted binary annotation, human-adjudicated validation subsets, and transformer-based model benchmarking. Using this pipeline, we construct two task-specific corpora containing 8,346 SI-labeled posts and 9,988 MH-labeled posts. The evaluation shows that model performance depends strongly on both task definition and validation protocol. BERT+LSTM achieves the highest SI stratified cross-validation F1-score, RoBERTa achieves the strongest SI holdout F1-score, and DistilRoBERTa achieves the best MH cross-validation F1-score. Human validation reveals different weak-label failure modes across tasks, with SI labels dominated by false negatives and MH labels dominated by false positives. These findings show that decentralized social media can support reproducible mental health benchmarking, but only when system design, label provenance, validation strategy, and deployment constraints are evaluated together.
benchmark - arxiv:2609.22694 · cs.AIToward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human ReviewGaurab Chhetri, Anika Baitullah, Subasish Das
Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.
evaluation framework - arxiv:2609.22691 · cs.AIGenerative Embodied Multiple Behavior Control Systems for Human-like AgentsChongyu Bao, Haokai Yang, Yuhan Wang, Zhaochong An +2
An enduring and richly elaborated dichotomy in cognitive neuroscience is that of human behavior control mechanisms, divided into habitual versus goal-directed. While existing human-like agent frameworks primarily focus on modeling goal- directed behavior, habitual behavior has been largely overlooked, though it plays a crucial role in human daily life. In this paper, we address this gap by studying multiple behavior control systems that jointly model goal-directed and habitual behaviors. We propose a human behavior control mechanism-inspired framework which the Habitual Controller retrieves cue-triggered behaviors from personal- ized habit memory, while the Goal-directed Controller employs a context-aware world model to predict action consequences and estimate their values. The Arbiter dynamically balances the influence of both systems according to individual differ- ences and momentary internal states. To reconstruct diverse human-level behavior instructions in 3D environments, we further develop a keyframe-guided 3D mo- tion generation module. Through extensive evaluation methods, human studies, and ablations studies, experimental results demonstrate that human-likeness per- formance is significantly improved by our approach. The efficacy of our approach indicates the benefits of leveraging habitual behavior and multiple behavior con- trol system coordination for believable embodied human-like agents.
embodiedworld modelagentagent framework - arxiv:2609.22690 · cs.LGMulti-Armed Bernoulli Bandits via Minimax Single-Arm StoppingHuikang Liu, Zhengchao Wang, Daniel Kuhn, Wolfram Wiesemann
We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems. Each SAB problem involves choosing between an unknown Bernoulli arm and a known reward. We show that minimizing worst-case regret of SAB problems over all non-anticipative policies admits an exact semi-infinite linear programming formulation. The resulting stopping policies offer a natural way to compare arms: the higher the known reward against which a policy continues sampling, the more promising the unknown arm. We turn this intuition into indices based on cumulative continuation probabilities, with a monotone adjustment and a reward-shortfall cap. By relating index errors to the regret of single-arm stopping policies, we establish a distribution-free regret bound of $4.45\sqrt{KT}+10.75K$ for $K$ arms and horizon $T$. This bound matches the minimax-optimal regret order established in the literature. The guarantee extends to rewards supported on $[0,1]$ through Bernoulli randomization. We also provide a finite-grid implementation with quantified approximation loss. In numerical experiments, the SAB-based index policy achieves lower worst-case regret than every tested benchmark policy across all evaluated numbers of arms and horizons, while closely matching the grid-based MAB minimax policy in the two-arm setting.
benchmark - arxiv:2609.22688 · cs.CVVision2CAD: A Visual Agent Harness for Explicit Geometry Referencing and Localization in Parametric CAD ModelingXi Cheng, Chenxi Zhai, Hang Cheng, Mingyu Fan +2
Generating parametric CAD models requires accurate geometry and stable feature dependencies. Existing methods face challenges in selecting geometric references, interpreting sketch-plane local coordinates, and establishing sketch constraints to projected external geometry. We present Vision2CAD, a visual agent harness that combines vision-language model (VLM) reasoning with deterministic CAD kernel operations. An ID-based interface supports explicit geometry selection, a local-coordinate bridge converts view coordinates into sketch coordinates, and projected-edge localization supports external sketch constraints. These mechanisms establish feature dependencies within the supported modeling operations and constraint types. We also introduce the Geometry Explicit Reference Dataset (GERD), which aligned commands, geometry states and IDs at every modeling step. On GERD-EVL and a DeepCAD test subset, Vision2CAD improves mIoU by 11.1\% and 5.6\% and reduces Chamfer distance by 17.3\% and 41.8\%, respectively. Parameter-editing experiments and ablation studies further proved the preservation of parametric dependencies.
agent - arxiv:2609.22684 · cs.ROStateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action PoliciesWenzhuo Li, Qiongfeng Shi, Yi Zhou
Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.
vision-language-actionvlamanipulationliberomemoryexternal memory - arxiv:2609.22682 · cs.AISelf-Organizing Agent Teams Learn to Reason TogetherAneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi +4
Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman $ρ=0.90$, $p=0.005$): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.
agentai agentbenchmark - arxiv:2609.22681 · cs.ROA Direct Rigid Transmission 2-DoF Wrist Extension for Tendon-Driven HandYujie Pang, Sadman Sakib, Mohammad Abdullah Al Faruque
Dexterous manipulation in confined spaces requires local control of hand orientation. Without a wrist, a dexterous hand must obtain this local orientation through coordinated motion of the robot arm, often involving several joints and a more complex end-effector path. We present CRAFT-Wrist, a concentric 2-DoF wrist extension that mounts between a robot arm and the CRAFT Hand without modifying the hand. Two XC430-T240BB-T servos drive sideways and front-back rotation through short rigid transmissions. Our initial hardware prototype actuates both axes under load, achieving a demonstrated workspace of $\pm20^{\circ}$ in radial--ulnar deviation (left--right) and a front-back range of $+80^{\circ}$ in flexion and $-18^{\circ}$ in extension. We characterize the resulting finger-motor loading across several wrist postures and demonstrate how the wrist supplies local orientation in grasping, nail hammering, and blackboard wiping. Demonstrations and assembly instructions are available at https://craft-wrist.github.io/.
manipulationdexterousgrasp - arxiv:2609.22678 · cs.RORiemannian Density-Driven Optimal Control: Tangent-Space LQR for Second-Order Multi-Agent Systems on Curved ManifoldsKooktae Lee, Ruchika Singh
Density-Driven Optimal Control (D2OC) provides an effective framework for steering multi-agent systems toward prescribed spatial distributions. However, existing D2OC formulations are primarily developed for Euclidean domains and do not directly account for intrinsic manifold geometry. This paper extends D2OC to second-order multi-agent systems evolving on Riemannian manifolds. The proposed Riemannian D2OC (R-D2OC) constructs a local distribution objective in the tangent space of each agent through logarithmic maps and uses its weighted center as the reference for a finite-horizon LQR. The resulting control is executed on the manifold through intrinsic second-order dynamics and parallel transport within a receding-horizon scheme. We establish local curvature-dependent bounds that quantify the approximation introduced by the tangent-space reduction and characterize the resulting target bias. Furthermore, we derive a conditional discrete-descent result showing that the closed-loop objective decreases when a local velocity-alignment condition is satisfied. Numerical simulations on a 3D ellipsoidal manifold demonstrate distribution-level control and empirically support the proposed approximation and descent results.
agentmulti-agentagent system - arxiv:2609.22677 · cs.ROSHAFT: A Slack-Compensating, Helical-Buckling-Attenuating Flexible-Shaft Transmission for Lightweight Multi-DoF ManipulationTomoya Takahashi, Moses Gladson Selvamuthu, Richiro Tadakuma, Kazutoshi Tanaka
Lightweight and slim manipulators enable safe operation in human living environments. Proximal actuation using remote transmission mechanisms, such as wire-driven or Bowden cables, effectively reduces inertia and arm size by relocating motors near the base and transmitting torque to distal joints. Existing approaches either increase mass through additional components, such as pulleys for direction changes, or suffer from reduced transmission efficiency due to friction losses. Flexible shaft transmission avoids both mass increase and excessive friction losses, but faces increasing angular transmission error due to helical buckling caused by slack generated at joint bending. To address this problem, we propose SHAFT: a Slack-compensating, Helical-buckling-Attenuating Flexible- shaft Transmission mechanism. This mechanism compensates for slack through a proximal tensioner, improving the angular transmission error and efficiency of flexible shaft transmission without increasing the moving mass of the arm section. In a transmission path containing four 90-degree bends, the proposed mechanism demonstrated approximately 30% higher efficiency and approximately 65% lower angular transmission error compared to a flexible shaft transmission without a tensioner. Using this mechanism, we fabricated a 6-Degree-of-Freedom (DoF) arm with a 1-DoF gripper manipulator consisting of a rotary module housing motors with a tensioner, and a 280 g weight for the arm module. The proposed manipulator represents a novel remote actuation system for achieving lightweight construction with high efficiency, contributing to the acceleration of safe robot deployment in human environments.
manipulationmanipulatorgripper - arxiv:2609.22674 · cs.LGMask-Aware Execution for Efficient JEPA TrainingMd Musfiqur Rahman Sanim, Zhihao Shu, Bahram Afsharmanesh, Amirali Mirian +2
Joint Embedding Predictive Architectures (JEPAs) are becoming a core representation-learning primitive and a building block for latent world models across vision, video, audio, brain dynamics, and time series. Despite (potential of) wide deployment, current JEPA training pipelines are inefficient: each input is executed through multiple mask-specific branches, with redundant target-side work, and memory-bound token routing. These costs grow with the number of masks and limit GPU efficiency. We present M-JEPA, a mask-aware execution architecture that restructures JEPA training without changing the learning objective. M-JEPA separates mask-independent computation from mask-dependent routing, enabling shared context encoder execution, fused token routing and slicing with backward support, sparse target encoder execution over the union of target tokens, and masked patch embedding for sparse inputs. The resulting pipeline preserves training semantics while reducing computation, memory traffic, and synchronization overhead. We implement M-JEPA for five JEPA variants and evaluate it on NVIDIA A100 GPUs. Compared against the state-of-the-art baselines, M-JEPA achieves up to 1.7x end-to-end training speedup for 2-10 masks. Separately, with masked patch embedding, 4.75x patch-embedding speedup at high sparsity. These results show that execution restructuring, rather than changes to the JEPA objective, is a key lever for efficient JEPA training.
world modelmemory - arxiv:2609.22672 · eess.SYGeneralizable Optimal Control with Transformers: Closed-Loop Certification and Near-Optimality GuaranteesTurki Bin Mohaya, Maitham F. AL-Sunni, John M. Dolan, Peter Seiler
This letter develops closed-loop performance certificates for a transformer-based feedback policy. The policy is trained to imitate optimal Linear Quadratic Regulator (LQR) control across a family of heterogeneous Multiple-Input, Multiple-Output (MIMO) Linear Time-Invariant (LTI) systems. First, we establish a finite-sample excess-risk bound for the imitation loss minimized during training. Second, for each fixed problem instance, we derive regional closed-loop guarantees consisting of a forward-invariant operating region and a worst-case bound on deviation from the optimal rollout. Our main result is a probabilistic certificate for finite-horizon closed-loop near-optimality. Using an exact LQR cost identity, we express excess cost as a measurable per-rollout statistic and use independent calibration and validation rollouts to obtain a high-confidence bound on its violation probability. We evaluate the certificate on $28$ benchmark systems. This uses the base policy on seen systems and system-specific fine-tuned copies on unseen systems, with each rollout drawing the plant, cost, and initial condition from the corresponding certification distribution. All per-system certificates have violation probabilities below $3.1\%$, each at $95\%$ confidence; twenty systems certify suboptimality below $10\%$, with the tightest threshold equal to $4.8\times10^{-6}$.
benchmark - arxiv:2609.22664 · cs.AIFrom Capability to Assurance in Autonomous Penetration-Testing Harnesses: A Framework and Reference ImplementationJoas Antonio dos Santos Barbosa
Research on large language model agents for penetration testing is evaluated almost entirely by capability: whether the agent captures a flag or reproduces a proof of concept. That metric suits a benchmark but is silent on the properties that decide whether an autonomous agent can be used in an authorized engagement: whether a reported finding is true, whether the agent stayed inside its authorized scope, and whether an operator can audit what it did. We call these assurance properties and argue that they belong to the harness, the runtime wrapping the model, and can be enforced in code. This paper makes three contributions. First, we define a framework of five assurance properties (evidence grounding, non destructive claim reduction, computed severity, enforced authorization, and tamper evident accountability), each with a formal model and an explicit acceptance test, connected to prior work in capability based security, tamper evident logging, and software provenance. Second, we position representative systems (PentestGPT, the Cochise reference harness, MAPTA, and the trajectory judge PentestJudge) within the framework using published coding criteria, and identify a consistent assurance gap. Third, we study one open source implementation, NeuroSploit, pinned to an exact commit, reporting its architecture, its complexity cost, and a content addressed artifact bundle from a run against a public deliberately vulnerable target. We execute the deterministic authorization and audit acceptance tests directly and find and report a real enforcement gap, which we reflect by scoring both properties as partial. We therefore claim an initial existence argument that the properties are realizable together, not a comparative performance result, and we specify the multi target, ablation, and adversarial evaluation protocol required to turn the framework obligations into measurements.
agentautonomous agentbenchmarkevaluation protocol