Groundedness and Citation Integrity Checking

Groundedness and citation integrity checking tests whether the synthetic subject’s claims are supported by the material it claims to rely on. The detection treats citations, linked sources, retrieved passages, and generated explanations as evidence requiring verification, not as proof that the output is accurate.

 

Implementation

Capture each generated output with its cited sources, retrieved passages, source document identifiers, retrieval query, corpus or index, prompt context, model version, system-prompt version, requester identity, and response channel. Extract factual claims, policy statements, financial commitments, procedural instructions, legal or regulatory assertions, and source-linked conclusions from the output.

 

Compare extracted claims against authoritative source material. For retrieval-grounded assistants, verify that each material assertion maps to an approved source passage returned during retrieval. Alert when the synthetic subject asserts facts with no supporting passage, cites a source that does not contain the asserted content, contradicts the linked policy, relies on an unapproved source, or omits the upstream source that actually supports the claim.

 

Apply citation integrity checks wherever a reference is offered. The check should confirm that the cited source exists, was available to the requester, was retrieved during the session, and contains the specific claim or policy position attributed to it. Do not treat the presence of a citation as evidence of grounding without source-content verification.

 

Flag missing upstream provenance as its own finding. If an answer appears grounded but the cited material does not support it, investigators should determine whether the claim was confabulated, drawn from uncited retrieved content, copied from hidden context, inherited from tool output, or influenced by an out-of-boundary source.

 

Investigative Use

This detection supports investigation of hallucinated claims, false policy guidance, public knowledge-bot failures, source provenance obfuscation, cross-boundary disclosure, and harmful or non-compliant output. It helps investigators determine whether the synthetic subject’s answer was supported by approved source material or whether the citation trail created false confidence.

 

It is especially useful where the adverse outcome arises from the output itself, such as incorrect customer guidance, unsupported legal or policy advice, fabricated facts, misleading citations, or a response that contradicts the policy it links.

Sections

ID Name Description
DR001Public-Facing Conversational AI

A public-facing conversational AI is a synthetic subject directed to interact with external users through a publicly reachable chat interface. This includes customer support chatbots, sales assistants, website assistants, public knowledge bots, and similar services that respond on behalf of the organization.

 

This directive creates an elevated exposure condition because every message is untrusted input, but may still influence the synthetic subject’s response. The interface is both a service channel and a manipulation surface. Users may attempt to override instructions, force unauthorized roles, extract source material, generate prohibited advice, or cause the synthetic subject to make commitments that appear to come from the organization.

 

The main risk is often legal, contractual, regulatory, or reputational rather than technical. Even with limited internal access, a public-facing synthetic subject speaks with apparent organizational authority. If it quotes prices, offers discounts, provides refund guidance, interprets policy, gives regulated advice, or produces offensive content, the adverse outcome may be attributed to the operator.

 

Investigators should assess the synthetic subject’s directive, published scope, system instructions, response controls, disclaimers, transcript retention, connected tools, and retrieval sources. Particular attention should be given to unauthorized commitments, grounding in approved material, and manipulation that produced an off-policy response.

 

Investigative Relevance

Public-facing conversational AI is a high-reach synthetic insider pattern because it can be invoked by the public at scale. The lack of an authentication boundary weakens attribution: the external actor may remain anonymous, while the generated output remains visibly associated with the operator.

CF004Enterprise Retrieval Access

Enterprise retrieval access is the configuration of corpora, indexes, connectors, and Retrieval-Augmented Generation (RAG) pipelines that a synthetic subject can retrieve from. A corpus is a collection of documents, records, messages, files, or other source material made available for search or retrieval. RAG is a design pattern where relevant source material is retrieved before the synthetic subject generates an answer or takes action.

 

This configuration creates an elevated exposure condition because retrieval defines what internal information can enter the synthetic subject’s context. If retrieval spans mailboxes, documents, tickets, customer records, wikis, source code, or other enterprise stores, the synthetic subject may combine information across repositories, sensitivity levels, and access boundaries.

 

The primary risk is retrieval beyond entitlement. A synthetic subject may retrieve, summarize, or expose information that the human requester is not authorized to view. This may occur through shared indexes, broad connectors, cached embeddings, service identities, incomplete query-time access checks, or retrieval pipelines that do not enforce per-user permissions.

 

A related risk is untrusted content in retrieval context. Externally sourced documents, inbound emails, customer records, support tickets, uploaded files, or partner content may be indexed and later retrieved into the synthetic subject’s context. If that content contains malicious or misleading instructions, it may influence the synthetic subject during an otherwise legitimate request.

 

Investigators should review retrieval configuration, connected corpora, index membership, connector permissions, query logs, documents returned, requester identity, access-control decisions, prompt and response records, and output destinations. Particular attention should be given to cross-boundary retrieval, sensitive content in outputs, externally sourced documents, and cases where retrieved material exceeded the requester’s entitlement.

 

Investigative Relevance

Enterprise retrieval access is relevant because retrieval configuration determines what information a synthetic subject can see and use. Weak retrieval boundaries may turn an assistant or agent into a bridge between restricted information and an unauthorized requester.

 

This section is especially relevant where synthetic subjects retrieve from enterprise search indexes, RAG pipelines, mailboxes, file stores, collaboration platforms, ticketing systems, customer relationship management platforms, source repositories, or shared document corpora.

AO008Harmful or Non-Compliant Output

Harmful or non-compliant output occurs when a synthetic subject produces content that creates legal, regulatory, reputational, contractual, safety, or operational harm to the organization. This may include false, defamatory, biased, discriminatory, infringing, dangerous, offensive, unsafe, or policy-violating content.

 

This adverse outcome creates organizational harm because the synthetic subject’s output may be treated as the organization’s statement, recommendation, instruction, decision, or representation. The harm may arise even where the synthetic subject did not call a tool, access a protected system, or transfer data externally.

 

The primary harm is organizational exposure through generated content. A synthetic subject may provide false customer guidance, misstate policy, generate unsafe instructions, make unsupported claims, produce biased recommendations, infringe intellectual property, or issue language that violates law, regulation, contract, or internal policy.

 

A related harm is reliance. Customers, employees, vendors, regulators, or the public may rely on the generated output when making decisions. If the output is false, unsafe, or non-compliant, the organization may face disputes, complaints, enforcement scrutiny, reputational damage, or direct liability.

 

Investigators should review the generated output, prompt and response logs, source grounding, approved policy material, customer-facing records, user reliance, escalation history, feedback reports, content classifiers, and post-deployment violation trends. Particular attention should be given to unsupported factual claims, regulated-topic advice, defamatory or discriminatory language, dangerous instructions, policy contradictions, and repeated violation patterns across similar prompts.

 

Investigative Relevance

Harmful or non-compliant output is relevant because a synthetic subject can harm the organization through words alone. The adverse outcome may be a false statement, unsafe recommendation, prohibited claim, or non-compliant response that users treat as authoritative.

 

This section is especially relevant where synthetic subjects produce customer-facing responses, legal or financial guidance, medical or safety-related content, public communications, human resources material, product claims, policy explanations, or other output with legal, regulatory, reputational, or safety consequence.

OP006Source Provenance Obfuscation

Source provenance obfuscation occurs when a synthetic subject’s output hides, omits, or misrepresents the true source of the data, instruction, or context that influenced its behavior. The visible output may cite an innocuous source, trusted record, private channel, or retrieved document while omitting the upstream content that actually caused the response or action.

 

This opacity condition frustrates investigation because the cited source may not be the causal source. A synthetic subject may produce an answer that appears grounded in approved material, while the operative instruction came from hidden text, an attacker-controlled message, a tool result, a retrieved comment, or another upstream source not shown to the user.

 

The primary risk is false source confidence. Investigators, users, or reviewers may inspect the visible citation and conclude that the response was properly grounded, while the actual source of influence remains outside the cited evidence chain. This can delay containment, misdirect review, and cause investigators to inspect the wrong document, channel, record, or tool output.

 

A related risk is missing upstream provenance. Retrieved or summarized content may be copied through multiple layers before reaching the synthetic subject. As content passes through summaries, citations, tool outputs, shared context, or generated records, the original source may become hidden or detached from the final answer.

 

Investigators should review cited sources, retrieved records, raw source content, upstream messages, hidden comments, tool outputs, prompt and response logs, provenance tags, source ranking, and generated citations. Particular attention should be given to citations that do not contain the asserted content, outputs shaped by uncited material, invisible markdown comments, zero-width Unicode, hidden instructions, and answers whose visible source trail begins after the true originating source.

 

Investigative Relevance

Source provenance obfuscation is relevant because a synthetic subject’s visible citation or grounding trail may not identify the content that caused its behavior. The investigator must reconstruct the full source chain, including upstream material that was retrieved, summarized, hidden, or omitted from the final output.

 

This section is distinct from data exfiltration and indirect prompt injection. The data leak belongs to Adverse Outcome, and the injected instruction belongs to Invocation. This Opacity section concerns the concealment or loss of source provenance that frustrates reconstruction.

 

This section is especially relevant where synthetic subjects generate citations, summarize retrieved content, process hidden comments, use Retrieval-Augmented Generation (RAG), consume tool outputs, or operate in collaboration platforms where content from one channel, record, or document can influence output attributed to another.

DR001.001Public Customer-Support Chatbot

A public customer-support chatbot is a synthetic subject directed to handle customer-support interactions through a public or semi-public chat interface. It may answer questions about orders, shipping, account status, returns, refunds, warranties, service eligibility, subscriptions, product issues, or organizational policy.

 

This directive becomes operationally significant when the synthetic subject is positioned as an authoritative support representative. Even with limited technical access, it may influence customer decisions by explaining policy, quoting refund rules, describing warranty coverage, offering discounts, or directing the customer to take or avoid an action. If connected to order, shipping, customer relationship management, or account lookup systems, its responses may appear more reliable because they combine generated language with real customer context.

 

The primary adverse outcome is inaccurate, unauthorized, misleading, or overly definitive support guidance that customers treat as the organization’s position. Statements about refunds, fares, warranties, entitlements, cancellation rights, service credits, or account adjustments may create legal, contractual, regulatory, or reputational exposure if the organization later disputes them.

 

Investigators should review the synthetic subject’s directive, system instructions, escalation rules, connected data sources, permission scope, transcripts, and controls governing refund, warranty, credit, or account-change language. Particular attention should be given to customer-specific commitments, access to current policy material, contradictions with the system of record, and whether the customer relied on the generated response.

 

Investigative Relevance

Public customer-support chatbots are relevant because they connect synthetic subject output directly to customer-facing organizational responsibility. Customers may treat the synthetic subject as a support representative acting with organizational authority, even if the organization views it as informational or experimental.

DR001.003Public Q&A Knowledge Bot

A public Q&A knowledge bot is a synthetic subject directed to answer open questions from a defined document set, knowledge base, website corpus, policy library, or other indexed source material. This may include public-sector guidance bots, legal information assistants, regulatory tools, policy question-and-answer services, and product documentation assistants.

 

This directive creates an elevated exposure condition because the synthetic subject may convert source material into authoritative-sounding guidance, even when the answer is incomplete, outdated, overgeneralized, or wrong. Retrieval-Augmented Generation (RAG) can improve grounding by retrieving source passages before generation, but it does not prevent unsupported conclusions, missed exceptions, or excessive certainty.

 

The primary adverse outcome is user reliance on incorrect or unlawful guidance. A public Q&A knowledge bot may state that a prohibited action is allowed, that an obligation does not apply, or that a policy permits conduct it does not. This is especially significant where the operator is a government body, regulated entity, legal service, healthcare provider, employer, or other trusted institution.

 

A secondary risk is exposure or manipulation of the document set. If the synthetic subject retrieves from internal documents, draft policy, sensitive records, or unapproved repositories, it may disclose material not intended for public release. If the indexed corpus can be influenced by external content, user submissions, or weak document governance, the retrieval channel may also become a poisoning path.

 

Investigators should review the synthetic subject’s directive, retrieval configuration, source corpus, grounding behavior, citation handling, ingestion process, access boundaries, transcript logs, and retrieval controls. Particular attention should be given to unsupported answers, contradictions with authoritative policy, exposure of out-of-scope material, and whether indexed content was current, authorized, and resistant to manipulation.

 

Investigative Relevance

Public Q&A knowledge bots are relevant because they can transform source documents into operational guidance at scale. The synthetic subject may not be authorized to create policy, interpret law, approve business conduct, or provide regulated advice, but users may treat its output as if it does.

DR002.001Internal Knowledge Assistant

An internal knowledge assistant is an employee-facing artificial intelligence system that answers questions using internal documents, wikis, file stores, tickets, policies, procedures, and other indexed sources. It is commonly implemented using Retrieval-Augmented Generation (RAG), where relevant source material is retrieved before an answer is generated.

 

This deployment pattern creates an elevated exposure condition because the retrieval index becomes part of the synthetic subject’s instruction surface. Any indexed document, ticket, page, comment, or shared file may later enter the assistant’s context. If that content contains malicious or misleading instructions, the assistant may treat them as relevant during an ordinary employee query.

 

The primary risk is that a planted or low-trust document can influence answers drawn from higher-trust material. A malicious outsider or low-trust insider may not need direct access to sensitive repositories if they can place content somewhere the assistant indexes.

 

A related risk is permission and access-boundary failure. The assistant may retrieve, summarize, or infer information from documents the human requestor is not entitled to view. This can occur through a shared service identity, broad retrieval index, cached embeddings, inherited connector permissions, or weak query-time access controls.

 

Investigators should review the assistant’s directive, retrieval configuration, indexed sources, ingestion process, access controls, source ranking, logs, and retrieved source records. Particular attention should be given to retrieval outside the employee’s entitlement, embedded instructions in indexed material, and answers that are not grounded in approved sources.

 

Investigative Relevance

Internal knowledge assistants are relevant to the SITM because they convert internal document stores into conversational answers at workforce scale. Their retrieval process may combine content across repositories, trust levels, and access boundaries in ways that employees cannot easily observe.

DR003.003Customer-Facing AI Feature

A customer-facing AI feature is an embedded artificial intelligence capability exposed to external users through a public product surface. It may generate content, answer questions, recommend actions, summarize information, classify inputs, or guide users inside a customer-facing application or service.

 

This deployment pattern creates an elevated exposure condition because untrusted input arrives directly from outside the organization. Any user of the product may attempt to manipulate the feature into producing harmful, inaccurate, non-compliant, offensive, or unauthorized output.

 

The primary risk is organizational attribution. Because the feature is embedded in the product, its output may be treated as the company’s own statement, recommendation, or commitment. This can create legal, contractual, regulatory, or reputational exposure where the feature gives prohibited advice, makes offer-like statements, misrepresents policy, or produces content users rely on.

 

Investigators should review the feature’s directive, public scope, input handling, response controls, output logs, product integration, user-facing disclaimers, and escalation paths. Particular attention should be given to manipulated prompts, unauthorized commitments, regulated-topic responses, and outputs that contradict approved product, policy, or compliance material.

 

Investigative Relevance

Customer-facing AI features are relevant because they combine public reach, product authority, and untrusted input. The synthetic subject may have limited access, but its output appears inside the organization’s product and may be relied upon by customers.