5 Chain-of-Custody Breakpoints When AI Touches Police Evidence

Read Time: 13 minutes

Why preservation, version history, and human authorization must be designed before AI reaches the case file

By Chris Ryan and Lili Kazemi


Chris Ryan

Lili Kazemi

An AI-assisted police report is created for an audience that includes people whose job is to test it. A supervisor may review it for completeness. A prosecutor may rely on it to make a charging decision. Defense counsel may examine it for the point where the source, the system, and the officer’s judgment no longer match.

That is a more demanding test than internal sign-off. It is also why the most important chain-of-custody decisions must be made before implementation. Once a source has been deleted, a log has expired, or a version has been overwritten, the agency may not be able to reconstruct what happened by rerunning the model or asking participants to remember the workflow later.

Our prior article, 7 Questions Every Police Chief Should Ask Before AI Enters the Case File, focused on the governance questions agencies should answer before deployment. This article moves to the record itself: after AI has touched a police report, can the agency still show what entered the system, what the system produced, what the officer changed, and who authorized the final record?

What digital chain of custody means

Digital chain of custody is the documented history of how a digital record was collected, preserved, accessed, transformed, transferred, and produced. For an AI-assisted police report, the relevant record may include more than the final narrative. It may include the officer’s dictation or selected source material, the first AI-generated draft, later versions, system and model identifiers, user activity, edits, approvals, and exports to other systems.

A broken chain does not automatically make a record inadmissible. The consequence depends on the jurisdiction, the purpose for which the record is offered, and the nature of the gap. But a weak history can make authentication harder, create disputes over completeness or accuracy, complicate discovery, and give a legitimate basis to question whether the offered record is the same record the system originally produced and the officer actually reviewed.

California has already translated part of this concern into law. California SB 524, now codified at Penal Code section 13663, requires covered agencies using AI to generate police reports to identify the AI program, obtain officer verification, retain the first AI-created draft for as long as the official report, and maintain a minimum audit trail identifying the user and any video or audio used. The statute does not answer every chain-of-custody question. Its minimum audit trail also does not require a full edit-by-edit diff or identify every person who made a later edit. SB 524 therefore establishes a legal floor for version retention and human authorization, while fuller change records remain additional controls agencies may adopt to create a more defensible history.

The evidence rules behind the operational problem

Chain of custody supports evidentiary foundation, but it is not a substitute for the rules of evidence.

  • Authentication. Under Federal Rule of Evidence 901, the proponent must produce enough evidence to support a finding that the item is what the proponent claims. A witness with knowledge may supply part of that foundation, and Rule 901(b)(9) specifically recognizes evidence describing a process or system and showing that it produces an accurate result.
  • Hearsay. Authentication and hearsay are separate questions. Rules 801, 802, and 803 may apply to human statements carried into an AI-assisted report. Rule 801 defines a statement and declarant by reference to a person, so purely machine-generated material may raise foundation and reliability questions rather than classic hearsay. Embedded witness statements still require their own analysis, and the public-records exception has criminal-case limits for matters observed by law-enforcement personnel.
  • Originals and duplicates. Rules 1001, 1002, and 1003 recognize that an accurate readable output of electronically stored information may qualify as an original and that an accurate duplicate may be admissible. An AI-generated summary, regenerated narrative, or partial export is not automatically an accurate counterpart of its source.

The practical point is modest but important: preserving a reliable history can help an agency explain the record, but it does not resolve every admissibility question or replace case-specific legal analysis.

Architecture determines what can be proved later

The architecture of an AI system is part of the evidence story. As Chris Ryan emphasizes, closed-loop processing and client-side encryption address different risks. Closed-loop processing concerns where the input is processed and whether agency information is routed to an outside model provider. Client-side encryption concerns who can view unencrypted agency data. Agencies should ask both questions rather than accepting a generic assurance that information is secure.

Version design matters just as much. In a dictation-to-draft workflow, a useful record may consist of three linked versions: the officer’s original dictation, the AI-generated draft, and the officer’s final edited report. A stable record identifier, timestamps, user identifiers, and an intelligible diff or equivalent change record allow those versions to be examined as one history rather than three disconnected files.

The limits must be equally clear. A system can preserve only what entered the workflow or what it was designed to log. If the officer never dictated a fact, the system cannot preserve that missing fact. A dictation-based system may avoid the privacy and relevance problems created by ambient audio or video, but the record will be only as complete as the information the officer deliberately supplied.

The five breakpoints at a glance

BreakpointWhat can breakSuggested preservation response
1. Source and system lineage separateThe final report cannot be reliably tied to the input, system or model version, user, or processing event.Use a stable record ID; preserve source, first draft, and final; capture timestamps, users, system and version, and material events.
2. Transformation history disappearsAI and human changes cannot be separated, and a later rerun cannot reproduce the original event.Retain the first draft and material versions; preserve the officer edit history; record the operation and timing of changes.
3. Prompt-cycle and reviewer-facing records divergeThe assembled prompt, retrieved context, raw model response, system flags, post-processing history, and displayed draft cannot be matched.Preserve stable request and response IDs plus material prompt, model, flag, validation, and post-processing metadata; document transient or omitted layers.
4. Human review and authorization cannot be reconstructedThe record does not show who reviewed which version, what changed, or who authorized the final report.Capture the exact version reviewed, material edits or acceptance, officer certification, supervisor review, and later corrections.
5. Preservation breaks at a handoff or exportThe device, AI platform, RMS, and prosecutor system retain different versions, metadata, or schedules.Map custodians and schedules; test end-to-end export; coordinate holds; document what a production includes and omits.

Breakpoint 1: The source and the system lineage separate

The first break occurs when the final report survives but the agency cannot reliably connect it to the input, the system that processed it, or the conditions under which it was generated. In a single-channel dictation workflow, source provenance and system provenance are closely related and are better treated as one problem.

Preservation should begin with a stable record ID that connects the input, the first draft, the final version, and any later correction. The record should identify the relevant user, timestamps, application, model or version, and material processing event. Where the source includes audio, video, an image, or another record, the agency should preserve the link to that source and document material omissions.

The goal is not to retain every technical artifact forever. It is to preserve enough reliable information to show what the system received, which system processed it, and how the offered report emerged from that process.

Breakpoint 2: The transformation history disappears

AI can transcribe, translate, structure, summarize, redact, or regenerate language. Each operation can change wording, emphasis, context, or the relationship between a source and a final narrative. If only the polished final report remains, later reviewers may be unable to tell whether a disputed sentence came from the officer, the source material, an AI transformation, or a later human edit.

Version retention is more reliable than attempted reproduction. Models and configurations change through patches, updates, and retraining. Running the same dictation through a system six months later may show what the model does then, not what it did when the original report was generated. A later rerun is therefore not a substitute for the contemporaneous first draft and version history.

Agencies should retain the versions that matter, record the operation and timing of material changes, and preserve an intelligible difference between the AI draft and the officer-authorized report. The prior article expressed the accountability rule simply: AI drafts. The officer verifies. The officer owns.

Breakpoint 3: The prompt-cycle record and the reviewer-facing record diverge

This is the article’s most distinctly AI-specific breakpoint. The officer ordinarily reviews the draft rendered by the application. But that screen is only the last visible layer of a longer prompt cycle, and the data generated at each layer may be stored, logged, or discarded differently.

A simplified request can pass through five technically distinct stages:

  1. Input assembly. The application combines the officer’s dictation with system instructions, a report template, prior-turn context, or retrieved source material. The assembled prompt may therefore contain more than the words the officer entered.
  2. Preprocessing and retrieval. Text may be normalized, segmented, converted into tokens, or supplemented with records returned from a retrieval layer. A token count alone does not preserve prompt content; a token sequence is meaningful only if it can be tied to the tokenizer and reconstructed into human-readable input.
  3. Model execution and controls. A particular model and configuration generate response tokens. Safety classifiers, validation rules, statute checks, tool calls, retry logic, or system flags may affect the result.
  4. Post-processing and display. The application may parse, reformat, redact, merge, or reject parts of the model response before presenting a structured draft to the officer.
  5. Audit capture and export. A separate logging layer determines whether request and response IDs, prompts, token counts, model versions, flags, events, raw outputs, and later exports are retained.

Those stages do not automatically create one continuous record. A system may retain request IDs, token counts, or safety and validation flags without preserving the assembled prompt or exposing those artifacts in the officer-facing audit trail; post-processing may also make the displayed draft differ from the raw model response. Logging may be optional, and some artifacts may be transient. “Intermediate reasoning” therefore requires precision: a rationale, trace, or generated explanation is another output, not necessarily a faithful record of hidden computation, and internal activations should not be treated as an available log unless the architecture actually preserves them.

The goal is not to preserve every token or hidden state. It is to maintain a proportionate event record connecting the source and assembled request, model invocation, material flags or validation events, raw or post-processed output where retained, officer-facing draft, and final authorized report. Depending on the system, that record may include stable record, request, and response IDs; source and template versions; retrieved-source references; model and configuration version; timestamps; material tool calls, retries, or flags; the rendered draft; and known omissions. Officer testimony can establish what appeared on screen and what the officer changed or approved, but not necessarily what the application assembled or the model returned before post-processing. Agencies should determine before deployment what system documentation, logs, and exports exist, who controls them, how long they persist, and whether they can be authenticated and produced in usable form.

Breakpoint 4: Human review and authorization cannot be reconstructed

A notation that a human was involved is not enough. The record should show who reviewed which source and version, what the reviewer changed or accepted, who authorized the final report, and whether a later correction occurred. Review and authorization are related, but they are not interchangeable.

In a dictation-to-draft workflow, the sequence may be straightforward: the officer dictates; the system generates a structured draft; the officer reviews the draft against the officer’s observations and dictation; the officer edits and certifies the final version; a supervisor reviews it; and the report enters the records management system. The system should capture the timestamps, model and version identifier, officer ID, an intelligible diff or equivalent record of material edits, supervisor ID, review time, and officer certification without turning the workflow into an unusable administrative burden.

Human accountability must be visible in the record. The point is not to make the officer responsible for the system’s architecture. It is to distinguish what the system did from the judgment the officer exercised and the approval the agency relied upon.

Breakpoint 5: Preservation fails at a handoff or export

The chain does not end inside the AI application. The earlier workflow map remains useful: officer to device or interface, device to vendor platform, platform to supervisor, supervisor to records management system, and records management system to prosecutor. Each handoff can create a new version, lose metadata, apply a different retention period, or produce an export that looks complete while omitting important history.

Agencies should map the custodian and retention rule for each location, coordinate litigation or preservation holds where applicable, and test the full export path before a real case depends on it. An export should be usable by the records management system and prosecutor case-management system, not merely readable on the vendor’s screen.

Retention parity is a sensible starting point: deleting the first draft or audit history before the final report’s ordinary retention period expires can defeat the purpose of preserving the chain. The exact period must align with applicable law, agency records schedules, contracts, and case-specific holds.

Criminal Justice Information Services (CJIS) is the FBI division whose Security Policy governs how systems protect criminal justice information when it is accessed, processed, stored, transmitted, or destroyed. CJIS is an information-security framework, not a chain-of-custody rule and not a universal police-report retention schedule. Its access, encryption, logging, and data-handling controls may support preservation, but agencies still need separate chain-of-custody and records-retention procedures.

Two field examples show why the record matters

When a missing fact was blamed on the system

In one partner-agency matter described by Ryan, an officer used a dictation-based system to generate an arrest report. During a later case review, a prosecutor concluded that the report did not establish probable cause. The officer attributed the gap to the AI, and a supervisor asked whether the system had omitted information the officer supplied.

The retained history answered the question. The original dictation, AI draft, and officer-edited final report showed that probable cause had not been articulated in the original input. There was no source statement for the AI to drop. The agency could address the event as a training and performance issue rather than a system failure because the record separated what the officer supplied from what the system generated.

When the system raised an issue but the officer made the decision

In a second anonymized example, a platform’s statute checker flagged that an enhanced statute might better fit the circumstances described in a violent-crime arrest report. The officer reviewed the flag against the facts and independently confirmed the change.

According to Ryan, the matter then proceeded under the enhanced statute. The evidentiary point is not to validate that charging decision in the abstract. It is to show that the system surfaced a possible issue, the officer evaluated it, and the final choice remained human and reviewable. A useful audit record should preserve that sequence so a later reader can distinguish a system-generated prompt from the officer’s legal and factual judgment.

These examples illustrate both sides of the same principle. A preserved record can show when the AI caused a problem, but it can also show when it did not.

10 Things Agencies Should Preserve

  1. The original officer input and any linked source material that the workflow actually used.
  2. The first AI-generated draft and any material intermediate version needed to explain a disputed transformation.
  3. The final officer-edited and authorized report.
  4. A stable record identifier connecting the source, draft, edits, final version, and later corrections.
  5. Timestamps, user identifiers, application and model or version information, and material processing events.
  6. A usable record of officer edits, certification, supervisor review, and later correction or approval.
  7. The assembled-request and model-invocation record, including stable request and response identifiers, source and template versions, retrieved-source references, the model and configuration version, timestamps, and material tool calls or retries.
  8. The reviewer-facing output path, including material flags or validation events, raw or post-processed output where retained, the draft rendered to the officer, and known omissions.
  9. The handoff path across the AI platform, records system, prosecutor system, and any other relevant repository.
  10. The applicable retention schedule, legal-hold process, access controls, and integrity protections.

5 Key Takeaways

These five takeaways translate the article’s legal and technical analysis into a practical review framework. Together, they show what agencies need to preserve, separate, document, and test before an AI-assisted report becomes part of the case file.

  1. Preserve the lineage from source input through AI processing to the authorized record, not only the final report.
  2. Treat the prompt cycle and officer review as separate evidence layers, preserve contemporaneous versions, and never treat a later model rerun as proof of the original event; what the system assembled and processed, what the application displayed, and what the officer authorized may not be the same record.
  3. Document human judgment by showing who reviewed what, what changed, and who authorized the result.
  4. Treat authentication, hearsay, and proof of contents as related but distinct evidentiary questions.
  5. Test retention and export across the full workflow before a real case exposes a gap.

Conclusion

The answer is not to preserve everything forever or to turn every officer into a system administrator. It is to design a proportionate, usable history before deployment: source, transformation, system context, human review, authorization, and handoff.

Police reports are built to be relied upon and challenged. An agency should be able to explain not only what the final report says, but how that report became the record. When the architecture and the human workflow preserve that answer, AI can reduce administrative burden without erasing accountability.

Educational note: This article provides general information and editorial analysis, not legal advice for a particular agency, proceeding, or deployment. Federal and state evidentiary rules, records laws, discovery obligations, and agency policies may differ.

Contributor Biographies

Chris Ryan

Chris Ryan is Chief Product Officer at PoliceReports.ai and is a former law-enforcement command staff member with nearly fifteen years of experience across training, tactical operations, undercover operations, leadership, and management. His role spans product development, strategy, sales, and innovation.

Lili Kazemi

Lili Kazemi is General Counsel and AI Policy Leader at Anant Corporation, where she advises on AI governance, risk, compliance, contracts, and policy. She brings more than 20 years of experience across Big Law, Big Four, and federal government roles, with a background in international tax, regulatory strategy, and cross-border legal frameworks. She is completing her certification the London School of Economics and Political Science’s Ethics of AI Masterclass and writes The Human Edge of AI, a LinkedIn newsletter examining AI at the intersection of law, policy, work, and everyday life.

You can now also follow Lili on substack at https://substack.com/@lilikazemi
Subscribe – The Human Edge of AI

About Anant

At Anant, we help forward-thinking teams unlock the power of AI safely, strategically, and at scale. From legal to finance, our experts guide organizations in building workflows that act, automate, and aggregate without losing the human edge. We turn emerging technology into practical, scalable business capability.

Disclaimer


This article is for general informational and educational purposes only and does not constitute legal advice. Law-enforcement agencies should evaluate applicable federal, state, local, tribal, and territorial law; constitutional and evidentiary requirements; prosecutor guidance; public-records and retention rules; labor obligations; CJIS Systems Agency requirements; agency policy; and the specific architecture and contract terms of any proposed deployment.