What a Website Generates That Anyone Would Want
A corporate website produces a continuous record of what it did and what visitors did on it, spread across systems owned by different teams and companies. The set that turns up in disputes is consistent: analytics data at event and user level; server, CDN and firewall access logs; advertising account data, including change history, ad copy versions and disapproval notices; Search Console performance data; CMS revision history and permission changes; tag manager container versions; consent platform configuration and records; form submissions and chat transcripts; and dated versions of the privacy policy.
These records answer the questions that decide website disputes. What did the page say on the day in question. Who requested what, from where, and when. What was the site configured to do, and when did that change. Very little of that can be reconstructed from memory or from the current state of the site.
How Machine-Generated Website Data Gets Admitted
Under Federal Rule of Evidence 901(a), authentication requires "evidence sufficient to support a finding that the item is what the proponent claims it is" — a low bar. For website data the operative illustration is FRE 901(b)(9): "evidence describing a process or system and showing that it produces an accurate result". Analytics reports, access logs and CMS revision histories are not statements by a person; they are the output of an automated system, authenticated by describing the system and showing it produces accurate results. That means establishing what the system is, how data is collected and stored, that the export is a true output of it, and — the part usually missing — the configuration in force at the time: timezone, filters, sampling, bot filtering. Configuration change history is itself evidence.
Two self-authentication rules added effective 1 December 2017 avoid a live foundation witness. FRE 902(13) covers a record generated by an electronic process or system that produces an accurate result, shown by a certification of a qualified person. FRE 902(14) covers data copied from a device, storage medium or file, authenticated by "a process of digital identification" — in practice, hash verification: hashing the original and the copy and showing they match. Both require a certification meeting Rule 902(11) or (12) and the Rule 902(11) notice, which is not optional. Both go to authenticity only; neither answers a hearsay or best evidence objection.
On hearsay, FRE 801(b) defines a declarant as a person, so purely machine-generated output is arguably not a statement at all. Records containing human input, such as CMS entries and form submissions, do contain statements, and the usual route is FRE 803(6), records of a regularly conducted activity, where the opponent bears the burden of showing untrustworthiness.
Default Retention Destroys It Before Anyone Looks
This is the failure mode, and it is silent: every system on the list has a default that deletes.
Google Analytics 4. Property-level data retention for standard properties is two or fourteen months, and fourteen is the maximum; Analytics 360 adds longer options. The setting affects explorations and funnel reports rather than standard aggregated reports, so the trend line may survive while the granular user-level data needed in a dispute is gone. It is not retroactive.
Search Console. Performance data is available for a rolling window of approximately sixteen months, and older data is not recoverable. That is shorter than the time most disputes take to mature, which is why it is the most commonly lost dataset in these matters.
Advertising platforms. Performance history is generally substantial, but change history has its own lookback limit, and paused ad copy, disapproval notices and audience configurations are fragile.
Logs and the CMS. Server, CDN and firewall log rotation is frequently set to days or weeks and is usually invisible to legal and marketing. Many content management systems prune revisions by default.
And no litigation hold reaches a third-party platform automatically: issuing a hold internally does not stop a search, advertising or SaaS provider applying its own retention schedule.
When the Duty to Preserve Starts, and What a Hold Reaches
The common-law duty to preserve attaches when litigation is pending or reasonably anticipated — not when a complaint is served. The triggers are recognisable: an accessibility or tracking demand letter, a cease-and-desist over advertising claims, a regulator's civil investigative demand, a preservation letter, or an internal decision to sue.
The duty extends to relevant electronically stored information in the party's possession, custody or control, and control is the difficult word, because most of the data sits with third parties: analytics, advertising and search platforms, a host or CDN, a CMS vendor, a CRM. Whether a company has control over vendor-held data is contested and turns on its contractual right to obtain it. The practical answer is to export the data rather than litigate the question.
A hold that works has a predictable shape: a written notice identifying the matter, the categories and the date range; distribution not only to custodians but to systems owners — whoever controls analytics retention, Search Console, log rotation, and the ad accounts; suspension of automatic deletion, including backup overwrite cycles and retention windows; acknowledgements; periodic reissue; written release at the end; and documentation of the process, which is the primary evidence of reasonableness if preservation is challenged. Deletion-on-request workflows built for privacy compliance can destroy data under hold, so the carve-out belongs in the tooling before it is needed.
What Rule 37(e) Actually Distinguishes
Federal Rule of Civil Procedure 37(e), adopted in the 2015 amendments, governs failure to preserve electronically stored information in federal civil litigation. It applies only if four predicates are met: the information should have been preserved in the anticipation or conduct of litigation, it is lost, the loss resulted from a failure to take reasonable steps, and it cannot be restored or replaced through additional discovery. That fourth predicate is why exports and backups matter so much — if the data can be obtained from a vendor or reconstituted from a copy, the rule does not apply at all.
| Rule 37(e)(1) | Rule 37(e)(2) | |
|---|---|---|
| Required finding | Prejudice to another party | Intent to deprive another party of the information's use |
| Is negligence enough | Yes; no culpability finding beyond failure to take reasonable steps | No; negligence and gross negligence do not suffice |
| Available measures | Measures no greater than necessary to cure the prejudice | Adverse-inference presumption or instruction, dismissal, or default judgment |
The intent requirement is the gate, and demanding by design; the 2015 amendment was written to reject the position that negligence could support an adverse-inference instruction. Intent can be shown circumstantially — deletion after a hold was in place, disabling logging after a demand letter arrives, failing to disclose a loss. Routine operation of a default retention policy is negligence. So the difference between a curative measure and losing the case usually comes down to what a company did after the trigger, and the worst sequence is to receive a demand letter and then change an analytics or logging setting. Rule 37(e) governs ESI in federal civil litigation; state courts vary.
Capturing It in a Form That Survives
How evidence is captured determines how much argument it takes to use later.
- Export rather than screenshot. A screenshot of a dashboard shows a rendering; an export shows the data, carries the parameters, and can be checked. Screenshots supplement exports rather than replacing them.
- Hash at the time of export and record the value. That is what makes FRE 902(14) available later without argument, and it costs nothing at the time.
- Document how the capture was made, contemporaneously. Who ran it, when, from which account, with which date range, timezone and filters, using which tool. That note is the substance of a 901(b)(9) or 902(13) showing.
- Capture pages as they appeared — full HTTP headers and response content, not only a rendered image — and do it yourself rather than relying on a third-party archive. Archived captures are point-in-time snapshots; there may be none for the date at issue; embedded files may come from different capture dates, so the rendered page is not necessarily what a visitor saw; and logged-in content is not captured.
Set every retention setting to its maximum today, since none work retroactively. Establish scheduled exports into storage the company controls, ideally with object-lock immutability. And keep a written data map of every website-related system: what it holds, who owns it, how long it keeps it, and how to preserve it.
If It Becomes a Dispute
Can we establish what this page displayed on this date, and how? Can we authenticate this analytics export — describe the system, the configuration in force, and show it produces an accurate result? What does the advertising account change history show about who altered what, and when? Was anything lost, and could it have been preserved? Those questions are what an expert witness is engaged to examine, and the answers depend on decisions made before anyone knew there would be a dispute.
Preservation is cheap and reconstruction is expensive, sometimes impossible. Where a demand letter, a regulator's inquiry or a serious dispute has arrived, the sequence that protects a company is to issue the hold in writing, include systems owners as well as custodians, export everything relevant into storage under the company's control, record hash values, document what was done — and change no settings on any system in scope.
Frequently Asked Questions
How do we prove what our website looked like on a particular date?
The strongest evidence is a contemporaneous capture the company made itself: the page's full response content and headers, saved at the time, with a hash value recorded and a note of how the capture was made. That supports self-authentication under FRE 902(14) and a process-and-system showing under FRE 901(b)(9). Third-party archives can be authenticated, but may hold no capture of the date at issue.
How long does Google Analytics keep our data?
For standard GA4 properties the retention setting is two or fourteen months, with fourteen the maximum; Analytics 360 offers longer options. The setting affects explorations and funnel reports rather than standard aggregated reports, so the granular user-level data most useful in a dispute expires while the trend line remains. It is not retroactive.
When does the duty to preserve website data start?
The common-law duty attaches when litigation is pending or reasonably anticipated, which is generally earlier than a complaint being served. A demand letter, a preservation letter, a regulator's civil investigative demand, or an internal decision to sue are all commonly treated as triggers. Because retention windows are short, a hold issued only then often arrives too late.
What happens if we delete data we should have kept?
In federal civil litigation, FRCP 37(e) applies only if the information should have been preserved, is lost through a failure to take reasonable steps, and cannot be restored or replaced from another source. Where it applies, 37(e)(1) permits measures no greater than necessary to cure prejudice. The severe sanctions under 37(e)(2) require intent to deprive; negligence is not enough.