DawnBands · marksman brief · 2026-09-04

Same brief, 3 models, one variable

Every lane got the identical 1,109-word system prompt and the same four unit briefs. The only thing that changed is which model wrote the prose. Scored by validate_native_copy.py — the same gate that runs before Fish reads anything.

Verdict

No model is crowned automatically. The table catches instruction failures and repeated prose; the side-by-side bodies are the actual decision surface. Read each unit blind for hook strength, native voice, mechanism logic, and emotional continuity.

Did they follow Ayden’s current framework?

No. None of the twelve raw outputs are final-framework compliant. GPT followed the dated benchmark brief most closely. Grok produced the best raw native voice. DeepSeek was a clear last.

The benchmark brief itself is stale against the September 2 canon: three supplied hooks are generic instead of naming the 13–18 ADHD-teen-mom lane; it imposes a word cap current doctrine removed; it bans failed-solution prices current doctrine keeps; and it omits the full ADHD Initiation Wall. These are prompt defects, not model defects.

Every final consumer piece also requires a humanizer pass. This test intentionally preserved raw first passes, so it measures writer potential—not publish readiness.

ModelFramework scoreBest atBiggest failure
Grok 4.56.5/10Native voice + avatar fidelityReuses winning lines across units
GPT-5.6 Sol6.3/10Mechanism + brief coverageGrade creep + most convergence
DeepSeek V4 Pro4.5/10Speed + divergenceCopies banned brief language

Blunt read: Grok is the best raw writer. GPT is the best mechanism thinker and exact brief follower. DeepSeek is not competitive for final prose without heavy control and editing. Full flags-only audit is saved with the source artifact.

Scorecard

ModelUnitsCleanFlagsCloned GradeWordsSec4-unit cost
deepseek-v4-pro4/403014.5113154.2
grok-4.54/402044.41002129.4
gpt-5.6-sol4/401256.0958148.3

Cloned counts sentences a model reused across two of its own units, excluding offer-floor lines that are supposed to repeat. It is the metric that matters: it measures the model's prose attractor, which is the one defect a brief cannot fix. Flags are validator violations — largely instruction-following, fixable by editing the brief. Cost is list price for four units at 8k in / 1.5k out.

Method notes

Routes: DeepSeek V4 Pro via OpenRouter; Grok 4.5 via the standalone SuperGrok OAuth entitlement stored by Hermes; GPT-5.6 Sol via Hermes-managed ChatGPT Codex OAuth. All received the same imported system prompt and unit briefs. Every successful raw attempt is retained under raw/; only ≥780-word completed responses enter the comparison.

At a glance

12
usable generations
0
validator-clean units
62
total validator flags
10
cross-unit cloned lines

Side by side