Each model family fails in a signature way. Some pad answers with caveats. Some ignore length limits. Some silently drop the middle of long inputs. Some invent citations with total confidence.
You can't eliminate failure modes — you can only pick the failure mode your task can tolerate. A brainstorm tolerates invention; a legal summary doesn't. A chat reply tolerates padding; an API response doesn't.
The seed prompt is a compliance trap: strict format, strict length, no commentary allowed. Run it on two models and count the violations. That count is real data about who you can trust with structured work.