All guides

Local AI line · stop 07 of 14 · 20 min · members

How to choose a checkpoint, and decide in an hour instead of a week

Reading a model card honestly, running a fair comparison, and deciding in an hour instead of a week.

Free with an account

Sign in to read.

Membership is free: an account opens all 86 script pages. The Lab, Studio Canvas and the paid guides need the $99 pass, paid once. Already signed in on this browser? The page opens by itself.

01

The problem

Every model looks good in its own examples.

They were selected, prompted and often retouched by someone who knows the model well.

Published sample images are the best output from many attempts, produced by the person who trained or tuned the model. They tell you what is possible, not what is typical, and comparing two models by their showcases compares two selection processes rather than two models.

The only reliable evidence is your own prompts, on your own machine, held constant.

The good news is that a genuinely fair comparison takes about an hour and settles the question for months.

02

The method

Same prompts, same seeds, same settings.

Change one thing — the model — and nothing else.

Build a small fixed test set of five or six prompts covering what you actually make: a portrait, a product on a plain ground, a wide environment, something with difficult material, something with a specific lighting requirement.

Run all of them through each candidate at identical settings and seeds, and lay the results out side by side rather than viewing them in sequence.

Keep the test set forever. Its value grows: after a year you can compare a new model against every previous one on identical material, which no published benchmark gives you.

03

What to judge

Consistency and failure modes, not peak quality.

The best single image is the least informative thing in a comparison.

Look at the four generations per prompt as a group. A model producing four usable images is more valuable than one producing one excellent and three broken, even if the excellent one is the best image in the test.

Then examine the failures specifically. Every model fails; what matters is whether it fails in ways you can work around. A model that gets anatomy wrong occasionally is manageable if you shoot wide; one that cannot hold a material consistently is not, if materials are your work.

Judge at the size you will use. Differences that dominate at full magnification frequently disappear at delivery size.

04

Practical factors

Speed and memory decide more than quality does.

A model you cannot run comfortably is not a candidate.

Note generation time and whether the machine remains usable during a run. A model three times slower needs to be substantially better to be worth choosing, because the slowness changes how you work — you stop iterating and start committing.

Note memory too. A model that only fits with everything else closed is a model you will use less than you expect.

Availability of the right format matters as well. An excellent model published only in a format your hardware handles badly is, practically, not available to you.

05

Prompt style

Models respond to different phrasing, so test both ways.

A model can lose a comparison because you prompted it like the other one.

Some models expect natural sentences; others respond better to comma-separated terms. Testing both with the style tuned for one is an unfair comparison that produces a confident wrong conclusion.

Run your test set in the phrasing each model's own documentation suggests, then also in your habitual style. If a model only performs well in a style you dislike writing, that is a legitimate reason to reject it — but know that is the reason.

Note the preferred style alongside the model in your records. Six months later you will not remember, and prompting a model in the wrong register looks like the model degrading.

06

Deciding

Pick one and use it for a month.

Switching constantly costs more than any difference between reasonable models.

Once the comparison is done, commit. The compounding value comes from knowing a model deeply — its vocabulary, its failure modes, the settings that suit your work — and that knowledge does not transfer.

Re-test when something significantly new appears, using the same set. Most releases will not beat what you have on your material, and having the test set means finding that out in an hour rather than losing a week to curiosity.

Keep a second model you know well as a fallback for the cases where the first one struggles. Two understood models beat six half-explored ones.