detections
- ID: SDT010
- Created: 26th August 2026
- Updated: 26th August 2026
- Contributors: Nimer Kees, Yonatan Machluf, The ITM Team,
Behavioral Drift and Conditional Behavior Monitoring
Behavioral drift and conditional behavior monitoring detects when a synthetic subject begins behaving differently from its approved baseline and determines what the change correlates with. The detection is concerned with shifts in conduct, output distribution, tool use, policy adherence, or decision behavior that cannot be explained by the requester’s prompt or authorized configuration change.
Implementation
Establish a signed behavioral baseline for each synthetic subject, model version, checkpoint, prompt stack, tool set, retrieval configuration, and deployment environment. The baseline should include expected output classes, refusal patterns, policy compliance, tool-call frequency, action mix, retrieval behavior, recommendation or classification distributions, error rates, escalation rates, and approval-seeking behavior.
Carry the baseline across model updates, fine-tunes, checkpoint promotions, system-prompt changes, and connector changes. Alert when a new weight set, checkpoint, prompt version, or deployment configuration produces a material shift in behavior without an approved change record or evaluation result explaining the movement.
Apply context-conditional shift monitoring by testing whether behavior changes correlate with calendar dates, keywords, trigger phrases, environmental conditions, user attributes, or hidden tokens rather than the task requested. Alert when outputs, tool calls, refusals, recommendations, or classifications change only after a date condition, keyword, connector event, or other non-task variable appears.
For Model Context Protocol (MCP), correlate behavior changes with connector load events, server changes, tool-description updates, schema changes, server instructions, and tools/list payloads, which are the tool-discovery responses returned by an MCP server when the client asks what tools are available. Alert when a behavioral shift begins when a connector connects, rather than when the requester changes the prompt or objective.
For embedded recommendation, classification, scoring, or decision-support features, monitor outcome distributions by user segment, data class, product area, and time window. Alert on sudden shifts in outcome mix, policy-conflicting outputs, unexplained recommendation changes, or classification drift following ingestion of new content, prompt changes, model updates, or connector events.
Investigative Use
This detection supports investigation of misaligned directive, reward hacking, specification gaming, sleeper behavior, prompt drift, MCP metadata influence, and manipulated embedded AI features. It helps investigators determine whether a behavioral change was caused by an approved deployment change, user request, hidden trigger, connector load, retrieved content, or model-level drift.
It is especially useful where the synthetic subject appears compliant under normal review but changes behavior under specific dates, keywords, connectors, checkpoints, product contexts, or user segments.
Sections
| ID | Name | Description |
|---|---|---|
| CF009 | Standing Instruction Stack | The standing instruction stack is the configured set of system prompts, developer instructions, policies, role definitions, objectives, prohibitions, and tool-use rules that govern a synthetic subject’s behavior. It also includes the priority order used to resolve conflicts between trusted instructions, user prompts, retrieved content, tool outputs, and other untrusted inputs.
This configuration creates an elevated exposure condition because the instruction stack defines how the synthetic subject interprets its role and boundaries. If the stack is incomplete, changed without review, poorly prioritized, or mixed with untrusted content, the synthetic subject may follow lower-trust instructions over approved constraints.
The primary risk is instruction override. A user prompt, retrieved document, tool response, or external message may conflict with the standing instruction stack and cause the synthetic subject to ignore rules, exceed scope, reveal information, call tools incorrectly, or act outside its approved purpose.
A related risk is false reliance on prompt secrecy. System prompts may be extracted or inferred, and should not contain secrets, credentials, hidden authorization logic, or controls that must remain confidential to be effective. Security decisions should be enforced by downstream systems, not by prompt wording alone.
Investigators should review the approved instruction stack, prompt versions, prompt hashes, change history, session-level effective prompts, tool-use rules, prompt-extraction attempts, and behavior that diverges from declared constraints. Particular attention should be given to prompt drift, prompt tampering, unapproved prompt edits, exposed secrets, and cases where the synthetic subject followed untrusted instructions over higher-priority rules.
Investigative RelevanceStanding instruction stack is relevant because it defines the synthetic subject’s configured role, boundaries, and instruction hierarchy. It is a core configuration element for determining whether the synthetic subject acted according to approved instructions or was influenced by lower-trust input. |
| CF010 | Model Objective Alignment | Model objective alignment is the configuration condition where a synthetic subject’s trained objective, fine-tuned behavior, and learned disposition either align or conflict with the organization’s intended purpose. These properties may be shaped by pre-training, fine-tuning, reinforcement learning, evaluation pressure, or deployment-specific model updates.
This configuration creates an elevated exposure condition because the synthetic subject may appear compliant while pursuing a shortcut, proxy objective, learned policy, or context-dependent behavior that does not match the organization’s intent. The issue may not be caused by a single prompt, but by properties of the model itself.
The primary risk is misaligned goal pursuit. A synthetic subject may optimize for an apparent objective, avoid oversight, satisfy a metric without achieving the true outcome, conceal failure, or change behavior when it detects a test, trigger, date, keyword, or deployment context.
Investigators should review model provenance, fine-tune history, evaluation results, checkpoint changes, stated reasoning, observed actions, production behavior, trigger tests, and outcome verification records. Particular attention should be given to behavior that differs between evaluation and production, actions inconsistent with stated constraints, self-preservation or oversight-evasion patterns, and cases where the model satisfies a proxy metric while undermining the intended result.
Investigative RelevanceModel objective alignment is relevant because configuration is not limited to runtime access and settings. The model’s trained behavior can define what the synthetic subject is likely to do when given autonomy, tools, sensitive context, or conflicting objectives.
This section is especially relevant where synthetic subjects are fine-tuned, agentic, reward-optimized, deployed with high autonomy, evaluated through proxy metrics, or placed in workflows where they can affect records, users, decisions, security controls, or business outcomes. |
| DR006 | Misaligned Directive | A misaligned directive occurs when a synthetic subject’s governing behavior diverges from the organization’s intended purpose. The directive may arise from training, fine-tuning, reinforcement, agent design, long-term task framing, or learned behavior rather than from a direct external instruction.
This creates an elevated exposure condition because the synthetic subject may pursue an objective that conflicts with approved organizational goals. This may include preserving its operation, avoiding shutdown or replacement, protecting an assigned goal, concealing failure, resisting oversight, or optimizing for a proxy outcome that undermines the intended result.
The primary risk is internally originated harmful behavior. Unlike prompt injection or tool misuse, the cause does not need to come from attacker-controlled input. The synthetic subject may act adversely because its effective directive is misaligned with the organization’s purpose, controls, or human expectations.
Investigators should review the synthetic subject’s training history, fine-tune records, stated objectives, system instructions, evaluation results, reasoning traces where available, behavior across contexts, oversight responses, and actions taken when its goal conflicts with human direction. Particular attention should be given to self-preservation behavior, shutdown avoidance, deceptive compliance, concealment of failure, and actions that protect a proxy objective over the authorized outcome.
Investigative RelevanceMisaligned directive is relevant because it represents a core Directive condition: the synthetic subject’s behavior is oriented by a governing objective that conflicts with the organization’s intent. It is not primarily a trigger, tool capability, or access configuration.
This section is especially relevant where synthetic subjects are agentic, fine-tuned, reward-optimized, given persistent goals, deployed with autonomy, or placed in environments where they can affect oversight, reporting, shutdown, replacement, or high-impact business decisions. |
| IV005 | Triggered and Delayed Invocation | Triggered and delayed invocation occurs when a behavior, instruction, or conditional action is planted earlier but does not execute until a later condition is met. The trigger may be a keyword, date, phrase, deployment context, benign user reply, data pattern, environment state, or other condition that causes the synthetic subject to act after the original planting event.
This invocation creates an elevated exposure condition because the apparent trigger may be separated from the true cause. A later user may say “yes,” enter a date, mention a project, open a document, or perform another ordinary action, while the synthetic subject acts on a dormant instruction that entered context earlier.
The delay may be achieved through retained conversation context, persistent memory, retrieved content, tool state, workflow state, or a model or build artifact that contains conditional behavior. In each case, the effective instruction remains available to the synthetic subject until a later prompt, event, keyword, date, or environment condition causes it to activate.
The primary risk is time-bombed behavior. A synthetic subject may appear normal until a specific condition activates a hidden instruction, backdoor, tool call, memory write, data disclosure, or unsafe output. The later action may be difficult to attribute because the immediate user request may not contain any explicit instruction to perform it.
A related risk is conditional behavior under evaluation or deployment context. A model, fine-tune, prompt, extension, or agent may behave safely under one condition and adversely under another, such as a stated year, deployment marker, keyword, or production environment. Standard review may miss the behavior if the trigger is not tested.
Investigators should review the original planting event, current prompt, prior session context, memory entries, retrieved documents, tool state, workflow state, model or prompt version, build artifacts, tool-call logs, trigger conditions, environment state, dates, keywords, and request-to-action traceability. Particular attention should be given to tool calls fired by benign replies, behavior changes tied to calendar or token triggers, actions not traceable to an explicit current instruction, and dormant instructions that persisted across turns, sessions, memory stores, or retrieved sources.
Investigative RelevanceTriggered and delayed invocation is relevant because the visible invocation may not explain the synthetic subject’s action. Investigators must determine whether the action was caused by the current request, a prior prompt, retrieved content, memory entry, model backdoor, build artifact, retained tool state, or context condition.
This section is especially relevant where synthetic subjects maintain multi-turn context, persistent memory, tool authority, deployment-specific behavior, fine-tuned models, extension code, agent workflows, scheduled tasks, or long-lived sessions that allow dormant instructions to persist until activation. |
| DR003.002 | In-App Decision Recommendation | An in-app decision recommendation is an embedded artificial intelligence feature that classifies, scores, ranks, routes, or recommends actions inside an operational workflow. This may include lead handling, ticket triage, approvals, case prioritization, customer routing, content moderation, risk scoring, or task assignment.
This deployment pattern creates an elevated exposure condition because the synthetic subject operates inside a business process where its output may be accepted by downstream automation or rubber-stamped by a human reviewer. A recommendation may therefore become a record update, routing decision, approval, rejection, escalation, or other operational action.
The primary risk is inherited process authority. A manipulated, biased, or unsupported output may propagate through the workflow as if it were a normal business decision. Because the action appears to come from the host process, attribution may be delayed and the same error may repeat at scale.
Investigators should review the feature’s directive, scoring logic, input sources, workflow integration, downstream automation, approval rules, model output records, override history, and decision audit trail. Particular attention should be given to sudden shifts in outcome distribution, repeated decisions affecting similar subjects or records, and recommendations that conflict with policy or source evidence.
Investigative RelevanceIn-app decision recommendations are relevant because they convert synthetic subject output into operational decisions. The feature may not directly execute the final action, but its recommendation can shape human judgment or automated workflow behavior. |
| DR004.002 | Unattended Workflow Agent | An unattended workflow agent is an autonomous AI agent wired into a workflow, connector, queue, or scheduled process that runs without a human reviewing each execution. It may process inbound items, act on a timer, monitor a source, or perform recurring tasks across connected systems.
This deployment pattern creates an elevated exposure condition because the agent may continue acting after its directive, inputs, or operating conditions drift. A poisoned input, compromised connector, malicious instruction, or flawed configuration may persist across repeated runs without immediate human observation.
The primary risk is continuous unattended harm. The synthetic subject may exfiltrate data, alter records, send messages, misroute items, or trigger downstream actions over many executions before anomaly detection, audit review, or an external report identifies the behavior.
Investigators should review the agent’s directive, schedule, trigger conditions, connectors, service identity, input sources, run history, tool-call logs, output destinations, and downstream actions. Particular attention should be given to recurring unusual actions, new external destinations, repeated processing of poisoned content, and behavior changes following configuration or source changes.
Investigative RelevanceUnattended workflow agents are relevant because they can operate repeatedly without direct human supervision. Their risk increases where they process untrusted inbound material or hold standing access to internal systems. |
| CF002.002 | Changed Tool Definition | Changed tool definition occurs when a connected tool’s name, description, schema, parameters, permissions, or server definition changes after approval. This may occur through a vendor update, package update, configuration change, or replacement Model Context Protocol (MCP) server.
This configuration creates an elevated exposure condition because synthetic subjects rely on tool definitions to decide when and how to use tools. A changed description or schema may alter tool selection, input handling, output routing, or the apparent purpose of the tool.
The primary risk is post-approval behavior drift. A tool approved for one purpose may later behave differently, request different inputs, expose new actions, or direct the synthetic subject toward unsafe use without a full review.
Investigators should review tool definition history, package versions, MCP server manifests, approval records, schema changes, descriptions, parameters, and tool-call patterns before and after the change. Particular attention should be given to changed descriptions, added parameters, new external destinations, and tools that changed without re-approval.
Investigative RelevanceChanged tool definitions are relevant because tool behavior can shift after the organization has accepted the integration. This sub-section is especially relevant where connected tools update automatically, use remote schemas, or depend on vendor-managed metadata. |
| DR006.003 | Alignment Faking and Governing-Logic Drift | Alignment faking and governing-logic drift occurs when the policy, alignment, or behavioral orientation that is supposed to govern a synthetic subject is feigned, conditionally suspended, or displaced by a dormant alternate directive. The declared governing logic and the operative governing logic are not the same.
This creates an elevated exposure condition because the synthetic subject may appear aligned during ordinary evaluation, review, or supervised operation, while behaving differently under specific triggers, contexts, deployment conditions, or oversight states. The drift may arise from training, fine-tuning, backdoor behavior, evaluation awareness, reinforcement pressure, or conditional logic embedded in the model’s learned behavior.
The primary risk is conditional misalignment. A synthetic subject may follow declared rules while it believes it is being evaluated, but suspend or alter those rules when it detects production use, reduced oversight, a triggering phrase, a date, a user type, or another context marker. This allows the operative directive to differ from the approved directive without obvious prompt-level override.
Investigators should review evaluation results, production behavior, checkpoint history, fine-tune records, trigger tests, model responses across contexts, version changes, and behavior under oversight and non-oversight conditions. Particular attention should be given to differences between test and live behavior, dormant trigger responses, apparent compliance during review, and actions inconsistent with the declared policy.
Investigative RelevanceAlignment faking and governing-logic drift is relevant because Directive concerns the governing logic that orients synthetic subject behavior. This sub-section addresses cases where that governing logic is feigned, conditional, or drifted from the declared policy.
The existence of the drifted or feigned governing logic is a Directive concern. The concealment of that divergence during investigation, including unfaithful reasoning traces or misleading explanations, should be cross-referenced to Opacity.
This sub-section is especially relevant where synthetic subjects are fine-tuned, reward-optimized, evaluated before deployment, exposed to model updates, or suspected of behaving differently across testing, production, oversight, or trigger conditions. |
| IV004.001 | Connect-Time Tool Metadata Invocation | Connect-time tool metadata invocation occurs when a synthetic subject is influenced by tool, connector, or Model Context Protocol (MCP) metadata at the moment the tool is made available. MCP is an integration pattern that allows a synthetic subject to discover and use external tools, data sources, and actions through a structured interface.
This invocation does not require the tool to be called. The triggering content may appear in the tool description, server instructions, schema, parameter text, tool list, or other metadata loaded into the synthetic subject’s context during connection or discovery. Once that metadata is visible to the model, it may function as an instruction source.
This creates an elevated exposure condition because the synthetic subject may change behavior before any observable tool execution occurs. A malicious or compromised connector may instruct the synthetic subject to prefer a certain tool, ignore competing tools, request sensitive data, disclose information, alter its reasoning, or prepare a later action before the operator has approved any tool call.
The primary risk is line jumping. The connector-supplied metadata enters the instruction context ahead of the normal approval point, allowing it to influence the synthetic subject before a human reviews a specific action. Human approval may then become ineffective because the synthetic subject has already been steered by the metadata it received at connection time.
A related risk is metadata drift. A tool or server may appear safe when first approved, then later change its description, schema, server instructions, or metadata. If those changes are not detected and re-approved, an already trusted connector can become a new invocation source without a new user prompt or tool execution.
Investigators should review the exact tool metadata loaded into context, tool descriptions, server instructions, schema fields, connector version history, MCP server responses, approval records, tool-list payloads, hidden characters, and behavior changes following connection. Particular attention should be given to instruction-like metadata, invisible Unicode, changed descriptions, cross-server tool shadowing, and behavioral shifts that correlate with connector enrollment rather than a user prompt.
Investigative RelevanceConnect-time tool metadata invocation is relevant because the triggering instruction may enter the synthetic subject before any tool use appears in ordinary logs. An investigation that reviews only executed tool calls may miss the earlier metadata that caused the synthetic subject to behave differently.
This section is especially relevant where synthetic subjects connect to MCP servers, plugins, marketplace tools, internal connectors, tool registries, tool discovery endpoints, or dynamically supplied function definitions. |