Changing a system prompt and eyeballing three example outputs feels like iteration, but it is not measurement — it is a guess wearing the clothes of an experiment.
The rule that keeps it honest
Any prompt change that touches production should run against the same scored evaluation set as the version it replaces, so "this feels better" becomes a number that either moved or did not.