Note  ·  2026-10-01  ·  RESEARCH

How the benchmark is run, and what it does not show

The Ufinq benchmark runs every week against PySR, gplearn and Operon with one fixed configuration per algorithm. This note is about the rules that make the numbers worth reading, and about the limits that come with them.

A benchmark run by the author of one of the systems in it deserves suspicion. The only useful response is to make the rules checkable and to publish what goes wrong along with what goes right.

The rules

  • One configuration per algorithm. Declared in advance and applied to every dataset and every seed. No per-dataset tuning, for Ufinq or for anyone else.
  • The same budget and the same splits. Every algorithm gets the same time on the same held-out partition.
  • Ten seeds. A single run of an evolutionary search says little. Results are medians over seeds, and the spread is published beside them.
  • Failures are results. Timeouts and non-finite predictions are counted and shown, not dropped.
  • Immutable sweeps. A finished sweep is never rewritten. It has a permanent address and a content hash, so a citation keeps pointing at the numbers it described.

What it does not show

  • Ten datasets is not many. The weekly set is small enough to run every week, which is its purpose. A rank difference on ten datasets is suggestive, not conclusive, and the page says so next to the test statistics.
  • Complexity is not comparable across systems. Each algorithm reports size in its own metric: nodes, tree length, or a weighted cost. A smaller number in one metric is not a smaller formula in another.
  • A fixed configuration is not the best configuration. Any of the four could do better with tuning. The comparison is between defaults a user would actually meet.
  • Ufinq is closed. The other three can be re-run by anyone from the published runners. Ufinq's trials can be checked against the public split, not reproduced.

Where to look

The current sweep, every earlier one, the per-dataset verdicts and the failure counts are on the Benchmarks page, together with the methodology and the data API.