Bias toward harder questions
I want the private rehearsal to expose the gap before the real room does, so I tune the clone to apply pressure rather than reward a polished package.
Decision model
I use the decision clone to make the private rehearsal harder than the real conversation, then record what it predicted before I know the outcome.
I use it to rehearse difficult decisions, pressure-test the package, and record what the model predicted before the real meeting.
The problem
I was carrying decision packages into important conversations with arguments that felt complete until the first hard question exposed the assumption I had skipped.
Agreement was the dangerous failure mode for a decision clone because easy approval would make the rehearsal feel good while leaving the package untested; constant objection would be just as useless once the model hardened into a caricature.

System output against fabricated inputs; no employer data.
Method: see How the evidence was made.The approach
I built four rehearsal modes: interrogate, red-team, predict, and rehearse. The decision clone runs deliberately hawkish because I would rather absorb the harder version of a question in private and arrive with the answer.
Every prediction enters an append-only receipt ledger before the real interaction, and later evidence resolves verdict accuracy, question precision, condition accuracy, confidence, and surprises as separate dimensions. The scorecard cannot update the model on its own.
Design decisions
I want the private rehearsal to expose the gap before the real room does, so I tune the clone to apply pressure rather than reward a polished package.
I make prediction receipts immutable once written so the later outcome cannot quietly rewrite what the model originally forecast.
I keep verdicts, questions, conditions, confidence, and surprises separate so a correct question prediction cannot be laundered into a correct decision forecast.
I allow calibration evidence to support a proposed change, while every update still requires a separate review and an explicit promotion decision.

System output against fabricated inputs; no employer data.
Method: see How the evidence was made.Controls and calibration
What remains unproven
The model can still over-fit to a caricature where every forecast predicts resistance on cost or risk. The receipt ledger exposes calibration drift, but the portfolio does not yet establish enough resolved decisions to claim general predictive accuracy.