
TO
Tal Oved, Roi Pony, Oshri Naparstek, Udi Barzelay
· 1 min read
ResearcharXiv cs.CL
Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
arXiv:2609.19799v1 Announce Type: new
Abstract: LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results. We run the analysis over a full grid of seeds and iterations. Our findings suggest that the best way to split a fixed budget between more seeds (width) and more iterations (depth) changes with the strategy, the task, and the total budget. Furthermore, we observe that the ranking of strategies also changes with the budget. On one task the strategy that looks worst at one seed is best at forty seeds. On another the best number of iterations is well below the value common in practice, so extra depth wastes budget that more seeds would turn into score. We provide a measurement protocol that reports the seeds-by-iterations frontier and practical guidance for using it.
Original source
This story was published by arXiv cs.CL and written by Tal Oved, Roi Pony, Oshri Naparstek, Udi Barzelay. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


