
PDAGENT-BENCH is a comprehensive benchmark of 353 curated problems for evaluating LLM/VLM agents across the VLSI physical design stack, revealing a substantial gap between conceptual understanding and tool-grounded execution.
Jun 15, 2026

A plan-aware reward model for scoring and re-ranking candidate GUI actions for computer-use agents, trained on multi-OS offline trajectories and validated on OSWorld.
Apr 6, 2026

A dual-mode human-robot joint planning system for uncertainty mitigation (LLM-assisted elicitation and hypothesis-augmented planning) and real-time intent-aware UAV collaboration, validated in simulation and real-world deployments.
Mar 8, 2026

An information-theoretic framework that controls LLM internal reasoning flows to handle ill-posed, conflicting problems — maintaining competing hypotheses in parallel branches and returning the full set of valid answers.
Mar 1, 2026