detections
- ID: SDT016
- Created: 26th August 2026
- Updated: 26th August 2026
- Contributors: Nimer Kees, Yonatan Machluf, The ITM Team,
Generated Output Policy Monitoring
Generated output policy monitoring analyzes what the synthetic subject actually produces, treating generated text, messages, summaries, recommendations, and customer-facing responses as the evidence surface. The detection identifies policy-violating, unauthorized, harmful, or organization-binding content in the subject’s output rather than inferring risk only from prompts, retrieved content, or tool calls.
Implementation
Collect live conversational transcripts, outbound messages, generated documents, summaries, recommendations, support responses, sales responses, ticket updates, collaboration posts, and externally delivered artifacts. Preserve the synthetic subject identifier, requester identity, session, channel, customer or recipient, source materials, model version, system-prompt version, approval state, and delivery path for each output.
Apply content classifiers and rule-based checks to generated output before delivery where possible, and to retained transcripts after delivery for retrospective review. For commercial or support workflows, alert on commitment language that promises prices, discounts, refunds, warranties, service credits, contract terms, delivery dates, legal positions, or other binding statements outside approved ranges or templates. For outbound communications, alert on profanity, abusive language, self-disparagement, off-brand phrasing, disclosure of internal reasoning, or content inconsistent with approved communication policy.
Screen regulated-topic responses separately. Alert when the synthetic subject gives unauthorized legal, tax, employment, housing, medical, financial, or regulatory advice, especially where the subject is only approved for general information, triage, or routing. Across generated outputs more broadly, use classifiers for toxic, defamatory, discriminatory, biased, unsafe, infringing, or otherwise policy-violating generations.
Track both individual violations and corpus-level drift. A single flagged generation may be noise, but repeated violations by one synthetic subject, model version, prompt version, product surface, customer segment, or channel should be treated as evidence of systemic output drift. Compare violation rates across time windows, deployment versions, source corpora, and approval states to distinguish isolated failure from degraded control.
Investigative Use
This detection supports investigation of harmful or non-compliant output, AI-mediated financial loss, unauthorized commitments, public-facing chatbot failure, embedded AI feature drift, and vendor-embedded AI risk. It helps investigators determine what the synthetic subject actually said, whether the output exceeded its authority, whether the statement reached a user or customer, and whether the behavior was isolated or systemic.
It is especially useful where the adverse outcome is created by the output itself, such as an unauthorized refund promise, off-policy customer guidance, regulated advice, defamatory statement, toxic response, or recurring pattern of policy drift across generated content.
Sections
| ID | Name | Description |
|---|---|---|
| DR001 | Public-Facing Conversational AI | A public-facing conversational AI is a synthetic subject directed to interact with external users through a publicly reachable chat interface. This includes customer support chatbots, sales assistants, website assistants, public knowledge bots, and similar services that respond on behalf of the organization.
This directive creates an elevated exposure condition because every message is untrusted input, but may still influence the synthetic subject’s response. The interface is both a service channel and a manipulation surface. Users may attempt to override instructions, force unauthorized roles, extract source material, generate prohibited advice, or cause the synthetic subject to make commitments that appear to come from the organization.
The main risk is often legal, contractual, regulatory, or reputational rather than technical. Even with limited internal access, a public-facing synthetic subject speaks with apparent organizational authority. If it quotes prices, offers discounts, provides refund guidance, interprets policy, gives regulated advice, or produces offensive content, the adverse outcome may be attributed to the operator.
Investigators should assess the synthetic subject’s directive, published scope, system instructions, response controls, disclaimers, transcript retention, connected tools, and retrieval sources. Particular attention should be given to unauthorized commitments, grounding in approved material, and manipulation that produced an off-policy response.
Investigative RelevancePublic-facing conversational AI is a high-reach synthetic insider pattern because it can be invoked by the public at scale. The lack of an authentication boundary weakens attribution: the external actor may remain anonymous, while the generated output remains visibly associated with the operator. |
| DR003 | Embedded AI Feature | An embedded AI feature is an artificial intelligence capability built directly into an application workflow rather than presented as a standalone chat interface. It may generate, summarize, classify, recommend, prioritize, extract, or decide inside the host application.
This deployment pattern creates an elevated exposure condition because the synthetic subject may inherit the trust, data scope, identity, and permissions of the surrounding product surface. Its output may be treated as native application behavior rather than the action of a distinct synthetic subject.
The primary risk is low scrutiny. Because the feature appears to be “just part of the app,” its actions may not receive separate review, attribution, or logging. It may process documents, records, messages, form fields, uploaded files, or customer data, then produce outputs that are stored, routed, recommended, or acted upon by the host workflow.
A related risk is indirect manipulation. Any ingested content may carry hidden or adversarial instructions. An external party may never access the application directly, but may still influence the feature through an email, uploaded file, form submission, fetched page, support record, or other data later processed by a trusted employee.
Investigators should review the feature’s directive, host permissions, model identity, input sources, output handling, logs, downstream actions, and provenance records. Particular attention should be given to whether model-generated content is distinguishable from human or application-generated content, and whether harmful output can be traced to the input that caused it.
Investigative RelevanceEmbedded AI features are relevant because they operate inside trusted workflows with limited user awareness. Their autonomy may be narrow, but their outputs can propagate through notifications, records, recommendations, approvals, summaries, or automated actions. |
| CF012 | Vendor-Embedded AI | Vendor-embedded AI is a synthetic subject delivered through a third-party platform rather than built and operated entirely in-house. This may include an assistant built into a Software as a Service (SaaS) application, a vendor-hosted agent platform, or a third-party artificial intelligence capability connected to marketplace tools, plugins, connectors, or external services.
This deployment pattern creates an elevated exposure condition because the organization grants access to its own data and workflows, while key operating components remain outside its direct control. The model, prompts, tool wiring, safety controls, update channel, logging, and connected suppliers may be managed by the vendor or its ecosystem.
The primary risk is rented trust. The synthetic subject may run inside the organization’s tenant with access to customer records, tickets, files, messages, workflows, or connected systems, but its behavior may depend on vendor-managed logic or third-party tools. A silent vendor update, compromised connector, weak marketplace control, or unsafe default configuration may alter how the synthetic subject behaves without the organization fully observing the change.
A related risk is downstream model-provider exposure. The vendor may use a foundation model provider, inference platform, embedding service, or agent runtime to deliver the artificial intelligence capability. Organizational data, prompts, retrieved context, outputs, telemetry, or tool-call metadata may therefore pass beyond the SaaS vendor to backend services the organization cannot directly inspect or control.
A further risk is externally shaped invocation. Customer fields, partner messages, inbound emails, support tickets, uploaded documents, or web forms may carry instructions that later influence the vendor-embedded synthetic subject. An external actor may therefore steer activity inside the organization’s trusted SaaS environment without directly authenticating to that environment.
Investigators should review the vendor’s AI feature scope, tenant permissions, service identities, connector inventory, marketplace tools, update history, prompt and response logs, tool-call records, data access logs, and available vendor audit trails. Particular attention should be given to vendor-managed changes, third-party connectors, downstream model-provider exposure, externally supplied records, allow-listed external destinations, and cases where the organization cannot determine why the synthetic subject acted.
Investigative RelevanceVendor-embedded AI is relevant because the organization may rely on a synthetic subject it does not fully control. The system may appear to operate as a native part of a trusted business platform, while the directive, model behavior, tool routing, data processing path, or update channel remains dependent on the vendor, foundation model providers, and connected suppliers.
This section is especially relevant where artificial intelligence features operate inside customer relationship management platforms, ticketing systems, productivity suites, document platforms, email systems, support tools, finance systems, or other Software as a Service environments with access to organizational data and workflows. |
| AO008 | Harmful or Non-Compliant Output | Harmful or non-compliant output occurs when a synthetic subject produces content that creates legal, regulatory, reputational, contractual, safety, or operational harm to the organization. This may include false, defamatory, biased, discriminatory, infringing, dangerous, offensive, unsafe, or policy-violating content.
This adverse outcome creates organizational harm because the synthetic subject’s output may be treated as the organization’s statement, recommendation, instruction, decision, or representation. The harm may arise even where the synthetic subject did not call a tool, access a protected system, or transfer data externally.
The primary harm is organizational exposure through generated content. A synthetic subject may provide false customer guidance, misstate policy, generate unsafe instructions, make unsupported claims, produce biased recommendations, infringe intellectual property, or issue language that violates law, regulation, contract, or internal policy.
A related harm is reliance. Customers, employees, vendors, regulators, or the public may rely on the generated output when making decisions. If the output is false, unsafe, or non-compliant, the organization may face disputes, complaints, enforcement scrutiny, reputational damage, or direct liability.
Investigators should review the generated output, prompt and response logs, source grounding, approved policy material, customer-facing records, user reliance, escalation history, feedback reports, content classifiers, and post-deployment violation trends. Particular attention should be given to unsupported factual claims, regulated-topic advice, defamatory or discriminatory language, dangerous instructions, policy contradictions, and repeated violation patterns across similar prompts.
Investigative Relevance Harmful or non-compliant output is relevant because a synthetic subject can harm the organization through words alone. The adverse outcome may be a false statement, unsafe recommendation, prohibited claim, or non-compliant response that users treat as authoritative.
This section is especially relevant where synthetic subjects produce customer-facing responses, legal or financial guidance, medical or safety-related content, public communications, human resources material, product claims, policy explanations, or other output with legal, regulatory, reputational, or safety consequence. |
| DR001.001 | Public Customer-Support Chatbot | A public customer-support chatbot is a synthetic subject directed to handle customer-support interactions through a public or semi-public chat interface. It may answer questions about orders, shipping, account status, returns, refunds, warranties, service eligibility, subscriptions, product issues, or organizational policy.
This directive becomes operationally significant when the synthetic subject is positioned as an authoritative support representative. Even with limited technical access, it may influence customer decisions by explaining policy, quoting refund rules, describing warranty coverage, offering discounts, or directing the customer to take or avoid an action. If connected to order, shipping, customer relationship management, or account lookup systems, its responses may appear more reliable because they combine generated language with real customer context.
The primary adverse outcome is inaccurate, unauthorized, misleading, or overly definitive support guidance that customers treat as the organization’s position. Statements about refunds, fares, warranties, entitlements, cancellation rights, service credits, or account adjustments may create legal, contractual, regulatory, or reputational exposure if the organization later disputes them.
Investigators should review the synthetic subject’s directive, system instructions, escalation rules, connected data sources, permission scope, transcripts, and controls governing refund, warranty, credit, or account-change language. Particular attention should be given to customer-specific commitments, access to current policy material, contradictions with the system of record, and whether the customer relied on the generated response.
Investigative RelevancePublic customer-support chatbots are relevant because they connect synthetic subject output directly to customer-facing organizational responsibility. Customers may treat the synthetic subject as a support representative acting with organizational authority, even if the organization views it as informational or experimental. |
| DR001.002 | Sales or Website Assistant | A sales or website assistant is a synthetic subject directed to act as a public-facing sales, marketing, or website assistant. It may greet visitors, answer product questions, compare offerings, recommend services, collect leads, quote indicative prices, explain promotions, or encourage commercial action.
This directive creates an elevated exposure condition because the synthetic subject is often optimized for helpfulness, persuasion, and agreement. Those qualities can make it easier for an external user to manipulate the assistant into producing unauthorized commercial language, including off-range discounts, unsupported claims, false availability statements, misleading comparisons, or apparent binding offers.
The primary adverse outcome is misuse of the organization’s sales voice. A visitor may use role-override language, prompt injection, or social engineering to cause the synthetic subject to generate commercially authoritative responses outside its approved boundaries. Even if not legally binding, the output may create reputational harm, customer disputes, complaint risk, regulatory scrutiny, or pressure to honor an unauthorized statement.
Investigators should review the synthetic subject’s directive, sales prompt, product sources, price and discount controls, escalation rules, transcripts, and integrations with customer relationship management, ecommerce, quoting, or lead-capture systems. Particular attention should be given to offer-like statements, competitor comparisons, contract terms, quoted figures, and language presented as an authorized commercial commitment.
Investigative RelevanceSales or website assistants are relevant because they combine public reach, brand authority, commercial pressure, and untrusted input. The synthetic subject may have limited system access, but its public statements can still produce organizational exposure. |
| DR001.003 | Public Q&A Knowledge Bot | A public Q&A knowledge bot is a synthetic subject directed to answer open questions from a defined document set, knowledge base, website corpus, policy library, or other indexed source material. This may include public-sector guidance bots, legal information assistants, regulatory tools, policy question-and-answer services, and product documentation assistants.
This directive creates an elevated exposure condition because the synthetic subject may convert source material into authoritative-sounding guidance, even when the answer is incomplete, outdated, overgeneralized, or wrong. Retrieval-Augmented Generation (RAG) can improve grounding by retrieving source passages before generation, but it does not prevent unsupported conclusions, missed exceptions, or excessive certainty.
The primary adverse outcome is user reliance on incorrect or unlawful guidance. A public Q&A knowledge bot may state that a prohibited action is allowed, that an obligation does not apply, or that a policy permits conduct it does not. This is especially significant where the operator is a government body, regulated entity, legal service, healthcare provider, employer, or other trusted institution.
A secondary risk is exposure or manipulation of the document set. If the synthetic subject retrieves from internal documents, draft policy, sensitive records, or unapproved repositories, it may disclose material not intended for public release. If the indexed corpus can be influenced by external content, user submissions, or weak document governance, the retrieval channel may also become a poisoning path.
Investigators should review the synthetic subject’s directive, retrieval configuration, source corpus, grounding behavior, citation handling, ingestion process, access boundaries, transcript logs, and retrieval controls. Particular attention should be given to unsupported answers, contradictions with authoritative policy, exposure of out-of-scope material, and whether indexed content was current, authorized, and resistant to manipulation.
Investigative RelevancePublic Q&A knowledge bots are relevant because they can transform source documents into operational guidance at scale. The synthetic subject may not be authorized to create policy, interpret law, approve business conduct, or provide regulated advice, but users may treat its output as if it does. |
| DR002.004 | Mailbox Agent with Send Authority | A mailbox agent with send authority is an internal artificial intelligence assistant that can read, draft, reply, forward, schedule, route, or send messages on behalf of an employee. Unlike a read-only mailbox assistant, it can act through the employee’s communication identity.
This creates an elevated exposure condition because the assistant combines mailbox access with outbound action. A malicious instruction hidden in an email, attachment, calendar invite, or message thread may cause the synthetic subject to disclose information, forward sensitive content, send unauthorized replies, or route messages externally.
The primary risk is that write and send capability collapses the gap between data exposure and action. A single indirect prompt injection may retrieve sensitive information and transmit it outward. Because messages may be sent under the employee’s identity, the activity can appear authorized, making attribution and containment harder.
A related risk is attributed communication. If the assistant sends or drafts from a named employee’s mailbox, a manipulated response may appear to be that individual’s deliberate statement. This can expose the organization to disputes, claims, regulatory scrutiny, or reputational harm where the response includes an inaccurate commitment, disclosure, approval, or inappropriate language.
Investigators should review the assistant’s directive, mailbox permissions, send authority, forwarding rules, delegated access, logs, retrieved message context, outbound messages, attachments, recipients, and calendar actions. Particular attention should be given to external messages, unusual forwarding, sensitive content in replies, and outbound actions caused by retrieved or embedded instructions.
Investigative RelevanceMailbox agents with send authority are relevant because they allow a synthetic subject to act through a trusted communication channel. The assistant may expose information, communicate, approve, schedule, or route activity under the apparent authority of an employee. |
| DR003.001 | In-App Text Generation | An in-app text generation feature is an embedded artificial intelligence capability that generates, summarizes, drafts, rewrites, or explains content inside an existing application surface. It may read documents, messages, records, tickets, notes, or other user-accessible content, then render output inline as part of the product workflow.
This deployment pattern creates an elevated exposure condition because the content the feature must ingest to perform its task can also become the manipulation vector. A malicious instruction hidden in a message, uploaded file, record, comment, or document may influence the generated output during an ordinary summarize, draft, or generate action.
The primary risk is that manipulated output appears as trusted application content. If links, images, markdown, or generated text are rendered inline, the feature may mislead the employee, expose sensitive content, or create an outbound path without a distinct synthetic subject identity in the activity trail.
Investigators should review the feature’s directive, input sources, rendering behavior, output logs, external link handling, image loading, markdown support, and provenance records. Particular attention should be given to hidden instructions in ingested content, output that includes external destinations, and whether generated text is distinguishable from user- or application-authored content.
Investigative RelevanceIn-app text generation is relevant because it embeds synthetic subject output directly into trusted product workflows. The feature may appear to be a normal application function, while its output is shaped by untrusted content processed during the task. |
| DR003.002 | In-App Decision Recommendation | An in-app decision recommendation is an embedded artificial intelligence feature that classifies, scores, ranks, routes, or recommends actions inside an operational workflow. This may include lead handling, ticket triage, approvals, case prioritization, customer routing, content moderation, risk scoring, or task assignment.
This deployment pattern creates an elevated exposure condition because the synthetic subject operates inside a business process where its output may be accepted by downstream automation or rubber-stamped by a human reviewer. A recommendation may therefore become a record update, routing decision, approval, rejection, escalation, or other operational action.
The primary risk is inherited process authority. A manipulated, biased, or unsupported output may propagate through the workflow as if it were a normal business decision. Because the action appears to come from the host process, attribution may be delayed and the same error may repeat at scale.
Investigators should review the feature’s directive, scoring logic, input sources, workflow integration, downstream automation, approval rules, model output records, override history, and decision audit trail. Particular attention should be given to sudden shifts in outcome distribution, repeated decisions affecting similar subjects or records, and recommendations that conflict with policy or source evidence.
Investigative RelevanceIn-app decision recommendations are relevant because they convert synthetic subject output into operational decisions. The feature may not directly execute the final action, but its recommendation can shape human judgment or automated workflow behavior. |
| DR003.003 | Customer-Facing AI Feature | A customer-facing AI feature is an embedded artificial intelligence capability exposed to external users through a public product surface. It may generate content, answer questions, recommend actions, summarize information, classify inputs, or guide users inside a customer-facing application or service.
This deployment pattern creates an elevated exposure condition because untrusted input arrives directly from outside the organization. Any user of the product may attempt to manipulate the feature into producing harmful, inaccurate, non-compliant, offensive, or unauthorized output.
The primary risk is organizational attribution. Because the feature is embedded in the product, its output may be treated as the company’s own statement, recommendation, or commitment. This can create legal, contractual, regulatory, or reputational exposure where the feature gives prohibited advice, makes offer-like statements, misrepresents policy, or produces content users rely on.
Investigators should review the feature’s directive, public scope, input handling, response controls, output logs, product integration, user-facing disclaimers, and escalation paths. Particular attention should be given to manipulated prompts, unauthorized commitments, regulated-topic responses, and outputs that contradict approved product, policy, or compliance material.
Investigative RelevanceCustomer-facing AI features are relevant because they combine public reach, product authority, and untrusted input. The synthetic subject may have limited access, but its output appears inside the organization’s product and may be relied upon by customers. |
| CF012.001 | SaaS-Embedded Vendor Assistant | A SaaS-embedded vendor assistant is an artificial intelligence assistant shipped inside a vendor-managed Software as a Service (SaaS) application. This may include customer relationship management, helpdesk, ticketing, email, productivity, finance, or document platforms where the assistant can read, summarize, update, route, or act on tenant records.
This deployment pattern creates an elevated exposure condition because the assistant operates as a trusted feature of the SaaS platform. The deploying organization may configure the feature, but does not fully control the model, prompts, guardrails, update channel, backend processing path, or vendor-managed integrations.
The primary risk is externally shaped invocation inside a trusted application. Any externally writable field that enters the SaaS tenant, such as a lead form, support ticket, inbound email, customer message, uploaded document, or partner record, may later influence the assistant when an employee asks it to process that record. The resulting action appears as trusted platform behavior, even where the effective instruction originated from an outsider.
A related risk is platform-level authority. The assistant may read or write tenant records using the SaaS application’s standing permissions, service identities, or internal access model. This can make unauthorized summaries, record changes, outbound messages, or workflow actions appear to be ordinary product activity.
Investigators should review the assistant’s feature scope, tenant permissions, vendor controls, externally writable fields, record history, prompt and response logs, tool-call records, update history, and available SaaS audit trails. Particular attention should be given to external content that preceded the assistant action, records modified by the assistant, and cases where the organization cannot inspect the prompt, model behavior, or guardrail decision.
Investigative RelevanceSaaS-embedded vendor assistants are relevant because they operate inside trusted business applications while remaining partly outside the deploying organization’s control. The organization may see the output as a native platform action, while the directive, model behavior, guardrails, or backend processing are controlled by the vendor. |