2026.08.08
Cross-Model Validation of prompt-as-code v0.4.0
We had Claude Opus 5 and GPT-5.6 Sol independently review prompt-as-code, then cross-referenced the results. Overclaims found, attribution errors caught, pattern effectiveness confirmed — and the conclusion that scope decisions can only be made by humans.
Read
2026.08.07
The Missing Layer: Design-Time Context Creation
The industry shifted from prompt engineering to context engineering, but the discussion focuses on runtime optimization. The missing layer: how to author the context worth managing in the first place. Observations from document-first development and over 1,000 prompt reviews.
Read
2026.07.04
Marketing vs. Reality — Reasoning, Self-verification, and the Stripe Case
Fable 5 is not the "best general-purpose model." Failure to evaluate source reliability, difficulty maintaining logical consistency, unused tools, defensive responses — structural weaknesses in how it conducts investigations, not just individual answers. Stripe 50M-line case study analysis. Full conversation transcript published.
Read