From Tables to Quantified Statements: Evaluating LLM Inference Generation through Executable VerificationResearchSep 22, 2026
Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)ResearchSep 22, 2026
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World TasksResearchSep 22, 2026
LegendBench: A Diagnostic Benchmark for Legend Understanding with Counterfactual InterventionsResearchSep 22, 2026
Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language ModelResearchSep 22, 2026