Reputation attacks on businesses
What To Do in the First 48 Hours

How To Preserve Evidence Before It Disappears

The procedure, in order, with the failure waiting at every step and the limit the Internet Archive states about itself

The three rules for the next hour

Before the procedure, three constraints that decide whether running it is worth anything.

Do not contact anyone yet. Not the poster, not the platform, not a letter from a firm. The usual sequence runs: a business finds the content, reacts to it, and discovers days later that the wording has changed or the account has gone.

Do not tidy anything. If the item sits on a profile you control, or a page on your own site is part of the story, leave it exactly as it is until it is captured. Deleting your own side of an exchange removes the context that makes the other side comprehensible, and it gets noticed later.

Do not work from the company account. A page viewed while signed in as the business can render differently from the public version, and on some platforms it tells the other side you have been looking. Sign out first.

What follows is a sequence. Run it in order — the failures at each step are mostly caused by skipping the one before.

Step one: build the list before you capture anything

Nearly every incomplete record I am shown has the same origin. Someone was sent a link, captured that link, and stopped. The item was never alone. Build a list of addresses first, in a plain spreadsheet, and capture afterwards. What belongs on it:

  • The item itself, at its permanent address rather than the address of the feed it appeared in.
  • The account that posted it, which usually disappears before the content does.
  • Everything else that account posted about your company, your executives or your competitors. The pattern lives here, not in the single item.
  • The surrounding context: the thread, the comments, the other reviews on the page.
  • The search results that surface it, captured as results pages. The position of a result is a fact that cannot be recovered later.
  • Copies elsewhere. Aggregators, scrapers, screenshots reposted on other platforms, syndicated versions carrying the same text. Search a distinctive phrase in quotation marks and see where else it lands.

What goes wrong: a list built from memory rather than from a search, and a list built once. Both lose the most useful part of the record for good, because the shape of a campaign across surfaces can only be seen while every piece is still standing.

Step two: capture the page the way a stranger sees it

The state of the browser is part of the exhibit, and an unrecorded state is a hole somebody else fills for you. Work signed out, in an ordinary window, on a normal connection. Record the browser and version, whether a VPN was running, and the machine's time zone. If a page genuinely cannot be seen signed out, capture it signed in and write down that you did and as whom, rather than leaving the difference to be discovered.

Three things break a capture here, silently.

Content that has not loaded. Long pages load as you scroll, comments hide behind a “show more” control, and review text truncates after a few lines. A full-page capture taken before expanding everything produces an image of a page that says less than the page said. Scroll to the bottom, expand it all, then capture.

App screenshots. A capture taken inside a phone app usually contains no address at all, so nothing in the image distinguishes it from any other page with the same layout. Use a browser wherever a browser exists.

Personalization. A results page is assembled for its viewer. Capture results signed out, note the exact query, and treat what you saw as what you saw rather than what everyone sees.

Step three: record what is not in the screenshot

An image shows what a page looked like. A record catches the facts that vanish when the page renders again tomorrow, and this is the step most people skip. For each item, at the time of capture, write down:

  • The date the platform displays, exactly as displayed. Most surfaces show a relative age — “3 weeks ago”. That string means nothing in a folder opened next spring unless the date you read it is beside it. Record both, with the time zone.
  • The handle and the display name, separate things that change independently while the posting history stays put.
  • The numbers on the page: star rating and review count, vote or like counts, a video's view count, the comment count. They move constantly, and nothing recovers yesterday's figure.
  • Where in a video or broadcast the statement occurs, to the second, and what is said around it.
  • Where the machine's clock came from. Your own computer's clock is the weakest timestamp in the collection; the server's own date, visible in the response headers, is an independent one.

What goes wrong: a folder of perfect images with no note of when they were taken or what the page gave as its own date. That is a set of pictures rather than a record of anything.

Step four: put a copy somewhere you do not control

Everything captured so far sits on a machine belonging to the party with an interest in the outcome. An independent copy lodged the same day changes that conversation, and it takes seconds.

The Internet Archive's Save Page Now is the highest return for the least effort in this subject, and it is honest about its limits. From the Archive's help documentation, read 12 August 2026: it “saves the page you enter including the images and CSS,” it “does not save any of the outlinks,” and while it works for most pages, “some sites prohibit crawling, a few have SSL (security) settings that make it break.”

So it takes one address, as the crawler could fetch it, signed out. Content behind a login is not there, and heavily script-rendered pages often come back partial — some review pages archive badly or not at all.

Which produces the failure at this step: submitting the address, seeing a confirmation, and never opening the result. Open it and read it. If what you needed is missing from the archived copy, you have a local capture and no independent one, and it is better to know today.

One further limit, from the Archive's removal guidance: site owners can ask for archives of their site to be excluded, and the Archive says it makes no guarantees beforehand about the outcome of such a request. That cuts both ways. Keep your own copy regardless.

The declaration that sets the terms, and the limit inside it

There is a reason to treat an Archive capture as a serious object and a reason not to overstate it, and both come from one document. The Archive publishes the declaration it will sign when a capture is used in litigation, so the terms on which a capture becomes evidence are public: archive.org/legal/affidavit, read 12 August 2026.

The operative instruction for anyone collecting is the address format, which the declaration gives as http://web.archive.org/web/[Year in yyyy][Month in mm][Day in dd][Time code in hh:mm:ss]/[Archived URL]. Record that whole string every time. The date is carried inside the address, and a shortened link discards the only part identifying which copy was seen.

Then the limitation the Archive states about its own records, which almost nobody quotes:

“The date indicated by an extended URL applies to a preserved instance of a file for a given URL, but not necessarily to any other files linked therein. Thus, in the case of a page constituted by a primary HTML file and other separate files (e.g., files with images, audio, multimedia, design elements, or other embedded content) linked within that primary HTML file, the primary HTML file and the other files will each have their own respective extended URLs and may not have been archived on the same dates.”

— Internet Archive, standard affidavit, archive.org/legal/affidavit, read 12 August 2026

The same document adds that clicking a link inside an archived page returns the archived file “with the closest available date to the initial file containing the hyperlink.”

What that means for how you describe a capture

Read those two sentences together and a habit of speech collapses. What an archived snapshot holds is an assembly of files fetched at their own separate moments, and browsing forward from it walks the viewer silently across time, one nearest-available copy at a time.

Three consequences follow, all procedural.

Describe what the record supports. The defensible form of words identifies the extended address and the timestamp it carries. The indefensible one asserts how a whole page looked on a particular day, which is a statement about files the snapshot may never have held on that day. Nobody notices the difference until a reader who knows this document opens your exhibit, and then it is the only thing being discussed.

Do not navigate inside an archive and treat the result as one capture. If a linked page matters, capture that address separately with its own timestamp.

Capture the archived copy too. Take a local full-page capture of the snapshot itself, showing the extended address, on the day you make it. Archived pages can later be excluded on request, and a link to something withdrawn is not a record.

And the boundary, because it is a genuinely different question: what a declaration establishes about a file is not what a court will admit or what weight it will give. That second question belongs to a defamation attorney.

Live and vanishing content: capture during, or not at all

Everything above assumes a page that will still be there in an hour. Some of this material will not be.

Take a live broadcast: nobody clipped it, the broadcaster kept no recording, and when the stream ends there is nothing left to capture. Nothing on the platform side rebuilds it either; retention was never part of the arrangement. Twitch states that its Community Guidelines apply to all content on the service, including video, chat, whispers and accounts (Twitch Community Guidelines, read 12 August 2026) — the rules reach live content, and reaching it is not retaining it.

So the procedure inverts. Record while it runs, with screen-recording software rather than a phone pointed at a monitor if there is any choice at all. Write down the channel, when the broadcast began and in which time zone, and how far into it the statement comes. Where a clip already exists, grab it and its address in the same minute. Whoever made the clip can delete it, so can the broadcaster, and one of them usually does once attention arrives.

The same urgency covers anything the platform designs to expire, and any chat window attached to a stream — often where the specific allegation actually appears, and rarely retained with the video. What goes wrong here is never technical. It is spending the afternoon deciding whether the statement was serious enough to be worth capturing.

Step five: the log, and why it cannot be written afterwards

The least glamorous step, and the one that decides whether anyone other than the person who made the collection can explain it.

Keep a plain contemporaneous note, written as you go: the date and time of each capture with the time zone, the machine and browser, signed in or signed out, which tool produced each file, and who did the work. Store the files under names that identify them, in a folder nobody touches afterwards, and keep a separate working copy for anything that needs cropping. Nothing in the original folder is ever edited.

Three failures, all common, all avoidable.

Reconstruction. A log written from memory in month four is a document about somebody's memory. This part cannot be back-filled honestly and should not be back-filled at all.

Messaging apps. Screenshots pushed through a group chat come out recompressed, resized and stripped of the file data they arrived with, and what lands in the folder is a copy of a copy nobody can account for. Move files as files.

The single point of knowledge. Whoever ran the collection is the only person who can describe it. If that is a contractor finishing next month, the method has to be written down while they are here.

Step six: do it again, and what to do when it is already gone

Preservation is not an event. The page captured on Monday is a different page next month: reviews are edited, threads grow, responses appear, ratings move, the result carrying it rises or falls.

Set a cadence — weekly while a matter is live, monthly afterwards — and re-capture the same list, keeping every generation. The value of the second capture is comparison: it shows that a review was quietly rewritten, that a thread was seeded from three accounts in one afternoon, or that a response was published and then pulled.

When something vanished before anyone captured it, the honest position is that it may be unrecoverable, and two routes are still worth ten minutes. Look for an archived copy of the address. If none exists, that says nothing about whether the page was ever live; it says the crawler never came. And look for copies, because scrapers, aggregators and reposted screenshots routinely outlive their sources. Reddit states the general shape of this: deleted posts stop being displayed, licensees must stop using them, and it cannot guarantee that third parties have deleted copies made without its permission (Reddit, Public Content Policy, read 12 August 2026).

And what none of it does: preservation removes nothing, makes no platform likelier to act, settles no question about whether anything said was true, and names nobody. It makes every later step possible, and it is the only step that becomes impossible if it is left.

Frequently Asked Questions

Do I need special software to preserve a web page properly?

No. A current browser does most of it: a full-page capture from the developer tools command menu, a saved copy of the page, and the response headers from the network panel. Command-line tools that write a single archive file containing the requests and responses together are better where they are available, and worth setting up if this will run for months. What separates a usable record from a folder of pictures is not the tool. It is capturing the whole list rather than one link, and writing down the surrounding facts at the time.

Can I just email the screenshots to myself?

As a backstop, yes. As the method, no. Images sent through mail and messaging systems are frequently recompressed and stripped of their file data, and what arrives is a copy of a copy carrying a new date. Keep the original files as files, in one untouched folder, and make any cropped or annotated versions separately. Mail yourself the folder as well if you want belt and braces, but do not let a mailbox become the only place the collection exists.

The review only says “3 weeks ago”. What do I write down?

Write down the string exactly as displayed, and the date and time you read it, with the time zone. A relative age is only interpretable against the moment it was read, and in six months the same post will say something different. Where the platform shows an absolute date somewhere else — on the account page, in a tooltip, in the page source — record that too and note where it came from. Two independent time markers are worth far more than one.

How often should I re-capture the same page?

Weekly while a matter is active, monthly once it is stable, and immediately after anything that might prompt a change — a report filed, a response published, press attention. Keep each generation rather than overwriting the last. The comparison between captures is where the useful facts live: text quietly rewritten, a rating moving, a cluster of posts appearing within one afternoon. A single capture shows the page existed. A series shows what happened to it.

Who in the company should actually do the capturing?

One person, consistently, who can later describe what they did. Splitting the work across whoever is free produces inconsistent files, gaps in the log and nobody able to account for the whole. Pick someone who will still be there in six months, write the method down so it survives them anyway, and keep them out of any public exchange with the poster — the person collecting the record should not also be the person arguing in a comment thread.

What should I save from a video or a livestream?

The recording itself if you can get one, plus the address, the channel or account, the title and description, the upload or broadcast time with the time zone, the view and comment counts at capture, and the point in the running time where the statement falls. Capture the comments separately; they move faster than the video. For anything live, record while it runs — a broadcast nobody clipped and the broadcaster did not save does not exist afterwards, and no operator will rebuild it.
Keep reading

The entries behind this guide

Every platform and every response named here has its own entry, with the operator's own policy quoted, the date it said so, and the row that names what will not work.

Top