2026.08.08

What Happened When We Had Multiple AIs Review Our OSS

Cross-Model Validation of prompt-as-code v0.4.0

We had Claude Opus 5 and GPT-5.6 Sol independently review prompt-as-code, then cross-referenced the results. Overclaims found, attribution errors caught, pattern effectiveness confirmed — and the conclusion that scope decisions can only be made by humans.

Read

2026.08.07

What the Context Engineering Discussion Is Missing

The Missing Layer: Design-Time Context Creation

The industry shifted from prompt engineering to context engineering, but the discussion focuses on runtime optimization. The missing layer: how to author the context worth managing in the first place. Observations from document-first development and over 1,000 prompt reviews.

Read

2026.07.04

Claude Fable 5 Hands-on Verification

Marketing vs. Reality — Reasoning, Self-verification, and the Stripe Case

Fable 5 is not the "best general-purpose model." Failure to evaluate source reliability, difficulty maintaining logical consistency, unused tools, defensive responses — structural weaknesses in how it conducts investigations, not just individual answers. Stripe 50M-line case study analysis. Full conversation transcript published.

Read