SyncAI.news, a Varaisys broadcasting
Decomposing Error and Style in Automated Clinical Coding
HS

Han-Chin Shing, Jack Moriarty, Ryan Ware, Afton Marchbanks, Carlyn Canvasser, Stefanie Higgins, Harsh Gupta, Fang Wang, Joseph Paul Cohen

· 1 min read

ResearcharXiv cs.CL

Decomposing Error and Style in Automated Clinical Coding

arXiv:2609.24877v1 Announce Type: new Abstract: In automated clinical coding, where the label space spans tens of thousands of diagnosis and procedure codes, models are currently evaluated against a single gold annotation, treating any deviation as error. But we find when two teams code the same 110 ACI-Bench encounters, they agree on only 73% of codes (Jaccard similarity) for the same note; even after an independent clinical audit removes erroneous codes, agreement rises only to 77%. Is that gap error or something systematic? We model the systematic component as coding style $\psi$, a coder- or site-specific policy over what to code and how much to document, and recast coding as $p(\mathrm{code}\mid\mathrm{note},\psi)$, estimating $\psi$ with a 10-dimension rubric. If style were noise, conditioning on it would do nothing. Instead, across five datasets a model conditioned with a data-matching style raises ICD F1 by up to 26 points and an extreme mismatched one lowers it by up to 21. Four prompt based coding methods spanning 39-49 F1 converge to 52-56 once style is supplied (All p<0.05). Much of what single-gold evaluation charges to model error is recoverable, unmodeled style.

Original source

This story was published by arXiv cs.CL and written by Han-Chin Shing, Jack Moriarty, Ryan Ware, Afton Marchbanks, Carlyn Canvasser, Stefanie Higgins, Harsh Gupta, Fang Wang, Joseph Paul Cohen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News