Every lane got the identical 1,109-word system prompt and the same four unit briefs. The only thing that changed is which model wrote the prose. Scored by validate_native_copy.py — the same gate that runs before Fish reads anything.
No model is crowned automatically. The table catches instruction failures and repeated prose; the side-by-side bodies are the actual decision surface. Read each unit blind for hook strength, native voice, mechanism logic, and emotional continuity.
No. None of the twelve raw outputs are final-framework compliant. GPT followed the dated benchmark brief most closely. Grok produced the best raw native voice. DeepSeek was a clear last.
The benchmark brief itself is stale against the September 2 canon: three supplied hooks are generic instead of naming the 13–18 ADHD-teen-mom lane; it imposes a word cap current doctrine removed; it bans failed-solution prices current doctrine keeps; and it omits the full ADHD Initiation Wall. These are prompt defects, not model defects.
Every final consumer piece also requires a humanizer pass. This test intentionally preserved raw first passes, so it measures writer potential—not publish readiness.
| Model | Framework score | Best at | Biggest failure |
|---|---|---|---|
| Grok 4.5 | 6.5/10 | Native voice + avatar fidelity | Reuses winning lines across units |
| GPT-5.6 Sol | 6.3/10 | Mechanism + brief coverage | Grade creep + most convergence |
| DeepSeek V4 Pro | 4.5/10 | Speed + divergence | Copies banned brief language |
Blunt read: Grok is the best raw writer. GPT is the best mechanism thinker and exact brief follower. DeepSeek is not competitive for final prose without heavy control and editing. Full flags-only audit is saved with the source artifact.
| Model | Units | Clean | Flags | Cloned | Grade | Words | Sec | 4-unit cost |
|---|---|---|---|---|---|---|---|---|
| deepseek-v4-pro | 4/4 | 0 | 30 | 1 | 4.5 | 1131 | 54.2 | — |
| grok-4.5 | 4/4 | 0 | 20 | 4 | 4.4 | 1002 | 129.4 | — |
| gpt-5.6-sol | 4/4 | 0 | 12 | 5 | 6.0 | 958 | 148.3 | — |
Cloned counts sentences a model reused across two of its own units, excluding offer-floor lines that are supposed to repeat. It is the metric that matters: it measures the model's prose attractor, which is the one defect a brief cannot fix. Flags are validator violations — largely instruction-following, fixable by editing the brief. Cost is list price for four units at 8k in / 1.5k out.
Routes: DeepSeek V4 Pro via OpenRouter; Grok 4.5 via the standalone SuperGrok OAuth entitlement stored by Hermes; GPT-5.6 Sol via Hermes-managed ChatGPT Codex OAuth. All received the same imported system prompt and unit briefs. Every successful raw attempt is retained under raw/; only ≥780-word completed responses enter the comparison.