A citation beside an AI answer is not data lineage. You need the source snapshot, transformations, retrieved context, prompt, available model configuration, raw output and human edits. Without that packet, assurance starts after the evidence trail has already broken.*
An AI-generated risk summary can cite every document and still be impossible to reconstruct. The link may now point to a revised policy. The retrieval layer may have selected different passages. A filter may have excluded the record that changed the answer.
That is the control gap. CPG 235 Managing Data Risk asks regulated entities to understand the flow of data and processing undertaken, which it calls data lineage. It also addresses metadata, end-to-end lifecycle controls, audit trails and the origin and alteration of data. A saved prompt and final answer cover only two points in that chain.
CPG 235 is a prudential practice guide, current from 1 September 2013. APRA states that practice guides describe its view of sound practice but do not themselves create enforceable requirements. APRA's 30 April 2026 letter to industry on AI connects current prudential risk management to AI, including data governance, supplier opacity, lifecycle ownership, assurance and human involvement for high-risk decisions. In its accompanying media release, APRA said it was not proposing additional requirements at that stage. Neither document prescribes a generative AI lineage template. The run-lineage packet below is a proposed control design built from those principles.
What does CPG 235 actually say?
CPG 235 uses a broad definition of data. It includes data that is entered, calculated or derived, and data that is structured or unstructured. Its quality dimensions include accuracy, completeness, consistency, timeliness, availability and fitness for use. It also identifies confidentiality and accountability as potentially relevant dimensions (CPG 235, data and data risk).
For data architecture, APRA points to sources, usages, update mechanisms, owners, authorised users, criticality, sensitivity and quality requirements. It also describes the value of technical information showing data structures, flows, systems, repositories, interfaces and controls across capture, processing, retention, publication and disposal (CPG 235, data architecture and lifecycle).
That language predates large language models, so the next step is an editorial application, not an APRA interpretation. For an AI run, the input is not just the file a user uploaded. It includes any source snapshot, cleaning or redaction, chunking, retrieval result, tool response, instruction and configuration that shaped what the model received. The output lineage continues through human edits, approval and use, which is why tracing the decision matters more than tracing the tokens.
APRA's 27 November 2023 Insight article, reporting a multi-year pilot with selected banks and subsequent prudential reviews, supports the direction without turning it into an AI rule. Its better-practice examples included mapping lineage for critical data elements, documenting and remediating controls across that lineage, and resolving issues across the data lifecycle.
Practitioner context: Data lineage answers where the material came from, what happened to it and where it went. Model explainability asks a different question about why a model behaved as it did. A defensible workflow may need both, but one cannot substitute for the other.
What must a run-lineage packet capture?
Treat each material AI execution as a run with its own identifier. The packet should travel with the work product, rather than sit in a technical log that the control owner cannot access.
Use four linked planes:
- Source plane: Record the source-system name, data owner, exact snapshot or immutable record identifiers, extraction time, business meaning, permitted purpose, classification and quality checks. Preserve retrieved passages inside an approved store, or use governed immutable references and integrity hashes when copying sensitive content would create another exposure. Do not rely on links to documents that may later change.
- Transformation plane: Record filtering, deduplication, redaction, field mapping, calculations, chunking, embedding or indexing steps, retrieval settings and exclusions. Identify the code, rule or configuration version responsible for each material change.
- Execution plane: Capture the user and system instructions available to the organisation, or their governed immutable versions and hashes, plus retrieved context, tool inputs and responses, execution timestamp, service, deployment, region, model alias or version where exposed, and material settings. Mark unavailable provider metadata as unavailable and record the transparency limitation. Do not invent precision the service does not provide.
- Decision plane: Preserve the raw output, automated checks, human edits, sources verified, exceptions raised, reviewer, approver, final version and actual use. Record whether the output informed a person, a paper, a control action or no action at all.

This is deliberately more than an activity log. A timestamp showing that a model was called does not reveal which version of a policy was retrieved, which rows were excluded, what the reviewer changed or whether the output reached a decision maker.
The packet should also inherit the source data's access and retention controls. CPG 235 discusses access based on business need, classification by criticality and sensitivity, retention strategies, change audit trails and desensitisation when data moves to a less trusted environment (CPG 235, retention and auditability). Capturing lineage is not permission to create a second, less controlled copy of sensitive data. A stable governed reference and integrity hash may be safer than duplication where the original material is already retained appropriately.
Use this prompt to prepare a manifest from implementation records before a controlled AI run. The data owner and workflow owner must verify every field and approve the input set before execution.
Risk-tier the packet. A low-impact drafting aid may need a compact manifest. An AI workflow supporting regulatory reporting, customer outcomes or a material control needs deeper source snapshots, reviewer evidence and change records. APRA's 2026 letter expects assurance and monitoring to be proportionate to use-case criticality and includes human involvement for high-risk decisions (APRA AI letter).
Can you rebuild the analysis after the model changes?
Do not define success as reproducing the same sentence. APRA observed that point-in-time assurance can be ill-suited to probabilistic models and that model updates and supplier opacity can limit an entity's understanding of behaviour and outcomes (APRA AI letter). The original model build may no longer be available, even when your own records are complete. Supplier-run change loops make this more likely, not less, as vendor-operated agent platforms put continuous improvement inside the service itself.
Use two reconstruction tests instead.
The input reconstruction test asks whether another authorised reviewer can assemble exactly what the organisation supplied to the model, including transformations, retrieved passages, instructions and the configuration information that was available. The evidence reconstruction test asks whether that reviewer can trace each material statement in the final work product to source evidence, see the raw model output, identify human changes and confirm who decided how the output would be used.
These tests do not require hidden model reasoning. They require observable evidence and accountable decisions.
Fictional worked example: [ENTITYNAME] uses an approved assistant to summarise control-test exceptions for [REPORTINGPERIOD]. The final paper says three exceptions share a root cause. Packet [RUNID] retains the approved extract, transformation rule [RULEVERSION], retrieved records for [CONTROLID], prompt version, available model and deployment details, raw answer, reviewer corrections and the final committee paper.
Six months later, the model has changed. The reviewer cannot demand identical wording, but can rebuild the evidence context and determine whether the three-exception statement was supported when used. If the packet contains only the prompt and final paper, that test fails.
Use this prompt to reconcile a completed run. A human control owner must investigate unsupported claims, approve corrections and decide whether the work product can be relied on.
Do this Monday
- Choose one recurring workflow. Start with a bounded GRC use case whose source data and human approval path are already known.
- Define the minimum packet. Set mandatory fields across the source, transformation, execution and decision planes. Add deeper evidence for higher-risk uses.
- Capture at the point of work. Generate the run ID and snapshot references before execution. Save the raw output before any human edits overwrite it.
- Run a reconstruction drill. Give a completed packet to a reviewer who did not perform the work. Ask that person to rebuild the input context and trace the final claims.
- Close the access gap. Apply the source data's classification, access, retention and disposal rules to the packet. Do not let an assurance artefact become an uncontrolled data store.
- Create a fail condition. Stop reliance on an output when a material source snapshot, transformation, retrieved passage, human approval or final-use record cannot be reconstructed.
Bottom line
A source list is not lineage, and a model log does not establish input provenance. CPG 235 gives practitioners a durable way to think about data flow, metadata, lifecycle controls and auditability, but the run-lineage packet is a proposed generative AI control rather than an APRA-prescribed format. Capture what the model received, what it returned and what the human did next. Under this proposed control standard, if another authorised reviewer cannot rebuild that chain, the output is not ready to carry material GRC work.
This article is general information and education only. It is not legal, compliance, financial or professional advice. Obligations vary by organisation and circumstance. Verify current requirements against the primary sources cited and seek advice specific to your situation.
References
- Australian Prudential Regulation Authority, CPG 235 Managing Data Risk, current from 1 September 2013. https://www.apra.gov.au/practice-guides/cpg-235
- Australian Prudential Regulation Authority, APRA Letter to Industry on Artificial Intelligence (AI), 30 April 2026. https://www.apra.gov.au/news-and-publications/apra-letter-industry-artificial-intelligence-ai
- Australian Prudential Regulation Authority, APRA calls for a step-change in AI-related risk management and governance, 30 April 2026. https://www.apra.gov.au/news-and-publications/apra-calls-step-change-ai-related-risk-management-and-governance
- Australian Prudential Regulation Authority, Quality data as an asset for boards, management, and business, 27 November 2023. https://www.apra.gov.au/news-and-publications/quality-data-asset-boards-management-and-business
TheAICommand. Intelligence, At Your Command.


