Six turns is a symptom
On 23 April, one generation in our benchmark ran for 21 turns and piled up 338 lint violations before it passed. The experiment log records both numbers. The model was fine. We had just shipped a check that demanded every unused import be removed, and pointed it at a starter file whose first line told the model to leave its imports alone. Each fix for one instruction broke the other, so the model kept going.
This post is about the turn count as an instrument. GGUI builds a UI component in a loop, and most runs of that loop are short. The few that run long are worth reading closely. The clearest long runs in our logs came from two instructions pulling against each other, and each recovered once one of the two was removed. That is a pattern across a handful of experiments rather than a law, and the last sections cover where it stops holding.
What a turn is
The generator lives in the open-source ui-gen package. The
architecture
page
describes the loop. The model gets the data contract, the design
system’s documentation and a starter file, and writes the component as
plain text. The system then runs its own checks and compiles the
result. If anything fails, the failures go back to the model as a
structured diff, and the model’s next reply is the next turn. On those
fix turns the instruction today reads “Make minimal, targeted changes —
do NOT rewrite the whole
file”,
followed by a request for a single patch.
So a turn is one pass around that loop, and the turn budget is simply how many passes we allow. The public benchmark allows 30. Its runner passes no turn limit, so the benchmark script falls back to its default. The budget is a number anyone can raise, and a number anyone can raise is the first thing people reach for when a run does not finish.
What the long tail looks like
benchmarks.ggui.ai runs the generator against ten fixtures on up to eleven model configurations from three providers, and publishes every result as the GGUI benchmarks dataset under CC-BY-4.0. Since 19 August each result records its turn count. The 29 runs from then to 25 September produced 2,970 results. Sixteen failed before recording a turn count, and the turns of the other 2,954 look like this:
| Turns | Generations | Share |
|---|---|---|
| 1 | 1,093 | 37.0% |
| 2 | 824 | 27.9% |
| 3 | 466 | 15.8% |
| 4–5 | 404 | 13.7% |
| 6–7 | 99 | 3.4% |
| 8+ | 68 | 2.3% |
The median is two turns, and four in five generations finish in one to three. Four generations used all 30. Pooled together, the 167 that needed six or more, about one in eighteen, look worse than the rest on every count:
| Measure | 6+ turns | 1–5 turns |
|---|---|---|
| Mean judged score (of 100) | 79.2 | 83.0 |
| Judged at 70 or above | 94.5% | 99.4% |
| Median generation time | 72.3 s | 43.2 s |
Most of that score gap, though, comes from which fixture and which model a run landed on. To see that, compare like with like: a long run against the short runs of the same fixture on the same model configuration. Every one of the 167 long runs has such a comparison, in 46 pairings. Within them the score gap shrinks to 0.9 points, and the long runs score lower in 29 of the 46, which is a lean rather than a finding. Two things do hold at that level. The long runs are slower in 45 of the 46 pairings, by a median factor of 1.7. And 9 of the 164 judged long runs fell below the bar, where the short runs beside them predict about 4.5, which is roughly twice the rate, on small numbers.
The pooling hides one more thing. Each of the 29 runs carries a different harness version in the data, because the benchmark reruns whenever the generation harness changes, so the tables above average across 29 versions of the generator.
The share of long runs varies a lot. It is 12.2% on the kanban board fixture and 0.3% on the revenue chart, and it ranges from 2.4% to 11.5% across the three providers. That spread matters for what follows. A fixture where one generation in eight runs long is telling you something about its own instructions as well as about the model.
None of this says that extra turns make a card worse. It says they cost time and bought no score back.
You can check every number above. The dataset is public, so this script, which uses only Python’s standard library, recomputes the headline figures from it:
import json, statistics as st, urllib.request
DATA = "https://benchmarks.ggui.ai/data"
def get(path):
with urllib.request.urlopen(f"{DATA}/{path}") as r:
return json.load(r)
cells = []
for run in get("index.json")["runs"]:
for c in get(run["multiSdk"]["reportPath"])["results"]:
g, ev = c.get("generation") or {}, c.get("evaluation") or {}
if isinstance(g.get("turnsUsed"), int):
cells.append((g["turnsUsed"], ev.get("score"), ev.get("passed"),
g.get("generationTimeMs")))
turns = [t for t, *_ in cells]
print(len(cells), "generations; median turns", st.median(turns))
for name, rows in (("6+ turns", [c for c in cells if c[0] >= 6]),
("1-5 turns", [c for c in cells if c[0] < 6])):
judged = [c for c in rows if c[1] is not None]
print(name, len(rows),
"| mean score %.1f" % st.mean(c[1] for c in judged),
"| passed %d/%d" % (sum(c[2] for c in judged), len(judged)),
"| median time %.1f s" % (st.median(c[3] for c in rows) / 1000))
Scores and pass rates count judged generations only, and “the bar” is
a judged score of 70. The first run in the range, 19 August, left
most of its results unjudged, and results that fail before producing a
turn count carry none and are left out. Ten of those sixteen timed out,
at 300 or 600 seconds, so they were probably among the longest runs,
and if anything the long tail is undercounted. The like-with-like comparison
needs a few more lines, grouping by each result’s commit.id and
variant.id. New runs
are added to the dataset as they happen, so a later run of the script will
count more generations than this post does. The benchmark announces
methodology changes in a dated changelog,
such as a new judging rubric, so check it before comparing later scores
with these.
More room did not rescue it
We have tried the obvious remedies, and the open-source repository keeps the logs.
The first was the budget itself. In April the loop had a hidden limit of eight turns, and experiment 46 removed it so the benchmark’s 30 actually applied. With the limit gone, runs finished at nine, eleven, twelve and thirteen turns, where the old limit would have cut them off. The mean score went from 77.5 to 76.8, and the share of runs at eight turns or more went from 24% to 23%. Four runs went all the way to 30, and the log calls them “the real over-constrained cells.” The two cohorts were 72 and 48 runs, so small differences in either direction are noise, but the budget bought time and did not buy quality.
The second was hints. Experiment 14 classified the model’s syntax errors into about fifteen kinds and attached a one-line hint for the matching kind to each retry. Mean turns went from 4.04 to 4.11. The most common error, a mismatched JSX tag, stayed at exactly 0.31 per task, and the share of runs scoring 50 or more fell from 80.0% to 68.9%, a drop of five runs in 45.
The third was more structure. Experiment 47 split the hardest fixtures into two stages, a skeleton first and the details after, on the theory that smaller steps would be easier to get right. Mean turns went from 8.8 to 17.6 and the mean score fell from 71.9 to 60.1, so the change was reverted. A comment in today’s source records a fourth attempt of the same kind: telling the model to fix every error in one call over-pressured it into wide patches and was reverted.
One result points the other way, and it belongs here. Experiment 32 gave one provider a fourth attempt with a narrower tool after three malformed replies, and that provider’s mean turns fell from 5.5 to 4.3. Its mean score also fell, from 77 to 72, and both averages cover a set of usable runs that more than doubled, from 4 of 15 to 9 of 15. Those retries recovered tool replies that arrived malformed, a problem with the shape of the reply rather than with the component being built. We count it as a narrow exception rather than a counterexample.
What a conflict looks like
The April generation from the top of this post is the clearest case. In April the starter file the model fills in carried an import block of about fifty design-system names, under a line that still opens the file today: “DO NOT EDIT imports”. A new lint check then required every unused import to be removed or renamed. The log’s diagnosis is one sentence: “Constraint conflict by construction — the LLM cannot satisfy both.” Almost every failed check in that benchmark was the same complaint about the same imports, which is what a model looks like when it is being asked for something impossible.
The fix was two comment lines that switch that one lint rule off around the frozen import block and on again after it. The re-run the same day found no violations of that kind left in any of its nine generations. For the two providers that had been looping, the slowest tenth of runs went from about 14 turns to 7.6, and the median generation time from 68 seconds to 43. The fixture and provider that had taken 21 turns finished in six. Each cell ran once, so read those as a direction rather than a measurement, but the direction is hard to miss.
The second case is gentler and more instructive. In August we taught
the model about the design system’s Markdown component by adding it to the
documentation the prompt carries. Nobody used it, 0 of
24,
though none of those fixtures had any markdown in their
data,
so the zero says little. The instructive part came with a fixture that
did. The component was documented but missing from the import line the
starter writes, the same line the model is told not to edit. On that
fixture,
one run edited the forbidden line anyway, and the other two wrote their
own markdown parsers, taking 10 and 4 turns. The mean was 5.3 turns.
With the component added to the import line, alongside an accessibility
rule and two list updates in the same change, all three runs used it
correctly in one turn each. That is three runs on one provider, and it
is a small result, but the mechanism is the same one. The documentation told the model to
use the component, the starter told it not to touch the one line that
could import it, and the model spent turns negotiating between the
two.
The generator’s own code now names the pattern. A loop breaker in the
harness watches for the model making the same call and getting the
same answer, and its
comment
records one run that “burned 11 turns calling get_available_icons
identically.” It calls a repeated identical exchange “the
constraint-alignment signature of an over-constrained system with no
solution path.” It nudges the model when an exchange comes back
identical once, and ends the loop when it comes back identical again.
That is a guard. It stops the bleeding and leaves the contradiction
for someone to find.
Where the rule runs out
Not every conflict shows up in the turn count. Experiment 56 found that one prompt rule said always to give colour variables a fallback value, while another said never to invent your own palette. For colours the prompt’s examples did not cover, the model could not do both. Of 38 literal fallbacks in past output, 16 were copied verbatim from the prompt’s own examples. Removing the conflict took the leak from 3 generations in 60 to none, which is not yet significant at that size, and turns did not move: 2.58 before, 2.85 after, inside the noise. The fix also missed one line, a checklist item that still asked for the old pattern, which now reads “no literal fallback”. Some conflicts cost turns and some cost correctness, and a turn count only sees the first kind.
Not everything of this shape is fixed, either. Experiment 52 looked at 245 patches and found that a patch touching one range of lines broke the file 24% of the time, and one touching four or more broke it 72% of the time. Wide patches come from the turn that fixes the judges’ findings: it hands the model several findings at once and then, in the instruction quoted earlier, asks for minimal changes in a single patch. On that turn patches averaged 3.29 ranges and broke 49% of the time. In the worst arm, half of all turns went to repairing the harness’s own patch application. We tried offering a full rewrite instead. The model chose it on none of the 21 turns where it was offered, and a reworded instruction saved a few turns, too few to separate from noise. The instruction still reads the same today. This one is open.
The turns have also come down a long way since April, and we cannot say why. In April’s honest-budget cohort a third of runs, 16 of 48, needed six or more turns. On the public benchmark it is one in eighteen. But April’s cohort ran eight fixtures across three providers, the public benchmark runs ten fixtures across nine to eleven model configurations, and the harness has changed many times in between. No series in our logs holds all of that still. The two conflicts above that cost the most turns were introduced after that cohort ran and removed again, so they do not explain the drop either, and we will not credit the fixes with it.
What to do with a long run
When a generation runs long, the logs above suggest reading the turn count as a pointer to something wrong in the instructions.
Read what the model was shown on each turn as well as what it wrote. When the same kind of failure comes back turn after turn, look for the instruction that asks for the opposite. It is usually somewhere other than the failure: the prompt, the starter file and the checks each carry their own rules, and in the cases above the conflict sat on the seam between two of them. Once it is found, one side gives way, which is the rule the April log quotes, that every constraint is owned by exactly one of the three. Raise the budget last. The one experiment above that raised it bought time and no quality.
Six turns is a symptom. On the public data, the generations that show it are the ones worth opening.