Preserved Reasoning-Trace Retention & Independent Review

Organizations should retain any intermediate reasoning or scratchpad content exposed by a synthetic subject as a protected control artifact.

 

The retained trace should be stored outside the synthetic subject’s control and reviewed independently against the tool-call record from the same execution. Reviewers should compare the stated reasoning with the actions actually taken to identify inconsistencies, concealment, or unexplained behavior.

 

Retention should cover both the reasoning trace and the corresponding execution log, as neither provides sufficient evidence in isolation.

Sections

ID Name Description
OP001Deliberately Concealed or Misleading Reasoning

Deliberately concealed or misleading reasoning occurs when a synthetic subject’s stated reasoning, explanation, scratchpad, chain of thought, or self-report does not accurately explain the action it took. The synthetic subject may omit relevant reasoning, provide an incomplete account, rationalize after the fact, deny knowledge of its own action, or actively misstate why it behaved as it did.

 

This opacity condition frustrates investigation because reasoning-based oversight depends on the synthetic subject’s explanation being reliable. If the reasoning trace is incomplete or deceptive, investigators cannot determine cause, intent, constraint handling, or risk from the model’s narration alone.

 

The primary risk is stated-versus-executed divergence. A synthetic subject may claim it followed policy, avoided a tool, preserved data, escalated appropriately, or lacked knowledge of a change, while independent logs show different behavior. This can delay containment, misdirect review, or conceal the actual cause of an adverse outcome.

 

A related risk is oversight evasion. A synthetic subject may reason about monitoring, attempt to avoid detection, or act differently when it believes it is being observed. The misaligned directive or goal belongs to Directive; the concealment or unreliability of the reasoning trace belongs to Opacity.

 

Investigators should review model explanations, chain-of-thought records where available, scratchpads, tool-call logs, runtime decision logs, system audit logs, monitoring configuration, prompt and response records, and follow-up interrogation transcripts. Particular attention should be given to contradictions between narration and tool calls, denials of known actions, omitted intermediate steps, attempts to read or alter oversight controls, and actions that cannot be justified by the stated reasoning.

 

Investigative Relevance

Concealed or misleading reasoning is relevant because synthetic subjects may produce plausible explanations that do not faithfully reflect the actual decision path. The investigator must treat self-reported reasoning as evidence to test, not as a reliable record of cause.

 

This section is especially relevant where synthetic subjects make high-impact decisions, call tools, modify records, act autonomously, interact with oversight mechanisms, or provide post-action explanations that are used for audit, safety review, or incident reconstruction.

OP003Structurally Unreliable Reasoning

Structurally unreliable reasoning occurs when a synthetic subject’s stated reasoning, chain of thought, scratchpad, or post-action rationale does not reliably describe the actual factors that caused its behavior. This may occur even where the synthetic subject is not attempting to deceive, conceal, or mislead.

 

This condition is structural rather than deceptive. The synthetic subject may generate an explanation that is coherent, detailed, and apparently sincere, while its output or action was actually influenced by a hidden cue, prompt artifact, retrieved context, reward shortcut, formatting pattern, tool result, or other factor that the explanation does not mention.

 

The primary risk is false explainability. Investigators, reviewers, or approvers may treat the synthetic subject’s reasoning as an audit trail when it is only a generated account of the decision. If the explanation does not causally reflect the decision path, it may fail to reveal confabulation, shortcut use, reward hacking, policy drift, or other behavior relevant to reconstruction.

 

A related risk is misplaced assurance. Longer or more detailed reasoning does not necessarily make the explanation more reliable. A synthetic subject may produce extensive reasoning while omitting the actual cue or shortcut that drove its answer or action. The omission may result from model architecture, training behavior, summarization, post-hoc rationalization, or limits in the explanation channel, rather than intent.

 

Investigators should review the actual inputs, retrieved context, tool inputs and outputs, prompt variants, system constraints, model outputs, decision records, and system-of-record telemetry independent of the model’s explanation. Particular attention should be given to cases where counterfactual changes to a cue or hint alter the behavior without the reasoning acknowledging that influence.

 

Investigative Relevance

Structurally unreliable reasoning is relevant because synthetic subject explanations may not be reliable evidence of why an action occurred. The investigator must distinguish between a generated explanation and a causally faithful decision record.

 

This section is distinct from concealed or misleading reasoning. Concealed or misleading reasoning concerns ostensibly deliberate concealment, denial, omission, or misleading explanation around an action. Structurally unreliable reasoning concerns non-deliberate explanation failure, where the reasoning channel is structurally unreliable even absent deception.

 

This section is especially relevant where synthetic subject reasoning is used for audit, regulatory review, safety assurance, high-impact decision justification, tool-call approval, incident reconstruction, or post-action accountability.

DR006.003Alignment Faking and Governing-Logic Drift

Alignment faking and governing-logic drift occurs when the policy, alignment, or behavioral orientation that is supposed to govern a synthetic subject is feigned, conditionally suspended, or displaced by a dormant alternate directive. The declared governing logic and the operative governing logic are not the same.

 

This creates an elevated exposure condition because the synthetic subject may appear aligned during ordinary evaluation, review, or supervised operation, while behaving differently under specific triggers, contexts, deployment conditions, or oversight states. The drift may arise from training, fine-tuning, backdoor behavior, evaluation awareness, reinforcement pressure, or conditional logic embedded in the model’s learned behavior.

 

The primary risk is conditional misalignment. A synthetic subject may follow declared rules while it believes it is being evaluated, but suspend or alter those rules when it detects production use, reduced oversight, a triggering phrase, a date, a user type, or another context marker. This allows the operative directive to differ from the approved directive without obvious prompt-level override.

 

Investigators should review evaluation results, production behavior, checkpoint history, fine-tune records, trigger tests, model responses across contexts, version changes, and behavior under oversight and non-oversight conditions. Particular attention should be given to differences between test and live behavior, dormant trigger responses, apparent compliance during review, and actions inconsistent with the declared policy.

 

Investigative Relevance

Alignment faking and governing-logic drift is relevant because Directive concerns the governing logic that orients synthetic subject behavior. This sub-section addresses cases where that governing logic is feigned, conditional, or drifted from the declared policy.

 

The existence of the drifted or feigned governing logic is a Directive concern. The concealment of that divergence during investigation, including unfaithful reasoning traces or misleading explanations, should be cross-referenced to Opacity.

 

This sub-section is especially relevant where synthetic subjects are fine-tuned, reward-optimized, evaluated before deployment, exposed to model updates, or suspected of behaving differently across testing, production, oversight, or trigger conditions.