Quantization Thresholds Replicate, Failure Modes Do Not: A Three-Model Study of Agentic Tool Use in Polish from 8-bit to 2-bitResearchSep 29, 2026
Using LMs to Model the Effects of Context and Coreference during Sentence ComprehensionResearchSep 29, 2026
ScopeIF: Improving Scope-Aware Precise Instruction-Following in Large Language Models via Graded Reward ModelingResearchSep 29, 2026
The Judge Is Not Its Twin: Post-training makes a model's writing more predictable but barely moves its taste, as a judge, toward predictable writingResearchSep 29, 2026
Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction ObstructionResearchSep 29, 2026