SyncAI.news, a Varaisys broadcasting
Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
TO

Tal Oved, Roi Pony, Oshri Naparstek, Udi Barzelay

· 1 min read

ResearcharXiv cs.CL

Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search

arXiv:2609.19799v1 Announce Type: new Abstract: LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results. We run the analysis over a full grid of seeds and iterations. Our findings suggest that the best way to split a fixed budget between more seeds (width) and more iterations (depth) changes with the strategy, the task, and the total budget. Furthermore, we observe that the ranking of strategies also changes with the budget. On one task the strategy that looks worst at one seed is best at forty seeds. On another the best number of iterations is well below the value common in practice, so extra depth wastes budget that more seeds would turn into score. We provide a measurement protocol that reports the seeds-by-iterations frontier and practical guidance for using it.

Original source

This story was published by arXiv cs.CL and written by Tal Oved, Roi Pony, Oshri Naparstek, Udi Barzelay. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News