
JK
Jinwoo Kim
· 1 min read
ResearcharXiv cs.LG
Constraint-Aware Training
arXiv:2610.02909v1 Announce Type: new
Abstract: When generating programs with language models, constrained decoding can apply program analyses to exclude tokens that violate syntax, scope, or typing rules. However, there is a duplication: standard training already teaches the model to suppress the tokens rejected by these analyses. This duplication leads to the question: if we will perform some analysis to filter a set tokens out during inference anyways, can we avoid teaching the model the said analysis altogether during training, and does this externalization lead to more efficient models? This paper defines a general constraint-aware objective satisfying this externalization desideratum and formalizes the benefits of externalization into three concrete theorems about model size and data efficiency. We show, through a controlled synthetic experiment, that the theorems survive training dynamics: constraint-aware training yields lower prediction loss at a matched parameter count and data compared to ordinary cross-entropy training, motivating training objectives that incorporate the analyses used during generation.
Original source
This story was published by arXiv cs.LG and written by Jinwoo Kim. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


