Note  ·  2026-10-01  ·  RESEARCH

Quality you can read: R², discounted for size

A search that trades accuracy against size needs one number to rank candidates by. Ufinq uses description length for the search, and reports it on a scale every reader already knows: R², with the unexplained share inflated by the size of the formula.

Every symbolic-regression system has to decide what "better" means when one candidate is more accurate and another is smaller. The usual answers are a weighted sum, which needs a weight nobody knows how to set, or a front of non-dominated candidates, which postpones the decision and hands it to whoever reads the result.

Ufinq does neither. It ranks candidates by description length: the number of bits it takes to write the formula down, plus the number of bits it takes to write down what the formula gets wrong. A larger formula is better only if the bits it saves in error exceed the bits it costs to state. There is no weight between the two, because both are measured in the same unit.

The trouble with bits

Description length is the right quantity for a search and the wrong one for a reader. Nobody has an intuition for whether 412 bits is good. So the number shown to a reader is a different view of the same ranking. For squared-error targets it is

quality = 1 − (1 − R²) · 4bits / n

where bits is the size of the formula and n the number of rows. It is R² with the unexplained share inflated by the formula's description length. Adjusted R² has the same shape, with the parameter count where this has the code length.

What that buys

  • A familiar scale. Zero is no better than predicting the mean; one is exact. The distance to the plain R² printed beside it is the price of the formula.
  • It converges. As the rows grow, the discount vanishes and quality approaches R². On small data a long formula has to be much more accurate to be worth it; on large data size matters less. That is the behaviour one wants, and it comes from the data, not from a setting.
  • No ranking changes. On the rows a formula was fitted on, the reported scale is a strictly increasing function of the search's own figure. What the search prefers and what the reader is shown never disagree.

What is still open

How to count the bits of a constant. Pricing a constant by its magnitude makes the result depend on the unit the data happens to be measured in, which R² does not. Pricing it by its digits does not have that problem. A benchmark comparison of the two is running; until it reports, the magnitude measure stays the default.