What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer QualityResearchSep 23, 2026
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and EvaluationResearchSep 23, 2026
ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-TrainingResearchSep 23, 2026
ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific AgentsResearchSep 23, 2026
"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces ItResearchSep 23, 2026
Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete FlowResearchSep 23, 2026
The MODA General Attribute Suite: A Four-Track Evaluation Benchmark for Fashion Attribute ExtractionResearchSep 23, 2026
AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree SearchResearchSep 23, 2026
DroneGround: Open-Vocabulary Drone Payload Characterization Using Synthetic Data and Grounded Vision-Language ModelsResearchSep 23, 2026