CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?ResearchOct 1, 2026
A Quantitative Study of Sustained Focus in Large Language Models via Repetitive Deterministic Prediction TasksResearchOct 1, 2026
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM RoutingResearchOct 1, 2026
Opportunistic Target Selection: Early Directional Commitment for Query-Efficient Black-Box Adversarial AttacksResearchOct 1, 2026
When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMsResearchOct 1, 2026