What Copying Looks Like in Practice
Website copying arrives in four recognisable shapes, and they call for different responses. The first is wholesale duplication: a site mirrored page for page on an offshore host, sometimes with the company name and logo intact. The second is scraped structured data — part numbers, specifications, prices — republished on an aggregator or marketplace. The third is a competitor lifting service-page copy, occasionally verbatim, more often reworded around the same structure. The fourth is republication that was once authorised: a partner or syndication arrangement that ended while the content stayed up.
Those situations differ in harm as much as in law. A full mirror carrying your brand is an impersonation problem before it is a copyright problem, and it warrants a fast, parallel approach to the host and any payment processor involved. An aggregator republishing specification tables may cost nothing at all. Naming which of the four you face is the first useful step, because it determines who to contact and whether contact is worth the effort.
What Is Protected, and What Is Not
Copyright protects expression, not information. 17 U.S.C. 102(b) excludes ideas, procedures, processes, systems and methods of operation from protection, and facts are not authored. A competitor whose product page states the same dimensions, tolerances, certifications and compatibility list is describing the same reality, and that is not infringement however expensive the underlying research was.
What is protected is the writing: the sentences, the ordering, the analogies, the worked examples, the photography and diagrams. Original selection and arrangement can carry protection too, though a claim resting on arrangement alone is narrow. Short phrases, product names and taglines sit largely outside copyright and belong to trademark analysis instead.
Companies that find a competitor covering the same subject in the same order frequently believe they have been copied, when what they have found is two teams describing one specification. Comparing the actual language, paragraph by paragraph, before doing anything else prevents an expensive misreading.
Establishing What You Published, and When
Priority is the fact everything turns on, and it is provable only from records that exist before the dispute does. Four sources carry most of the weight: third-party web archive captures with their capture dates; the site's own revision history and publication timestamps; deployment records and server logs; and search index records, including dated listings and cached copies that sometimes preserve a version nothing else retained.
Gather that material before sending anything. Once a notice lands the copied page often disappears, and with it the comparison the whole claim rests on. Capture the offending page in full, with the URL and date visible, and record how the capture was taken.
Registration posture is worth deciding in advance rather than after discovery. Registration gates a US infringement suit, and the rules conditioning statutory damages and attorney's fees turn on registration before the infringement began or within three months of first publication. For a library published continuously, that is a decision about routine practice, not about one article.
The Notice and Takedown Mechanism
The takedown regime is a bargain with intermediaries, not a remedy against the copier. 17 U.S.C. 512 gives qualifying service providers a safe harbour for material stored at the direction of users, conditioned on designating an agent with the Copyright Office and acting on notices that comply with 512(c)(3). A notice therefore goes to a provider, and what it achieves depends on which one.
- The host can disable or remove the material.
- A search engine can remove the URL from its index. The page stays online for anyone with the link.
- The registrar generally does not police page content; its lever is the domain itself, and it is rarely the right addressee.
- A content delivery network may only pass the complaint to the origin host it fronts.
A compliant notice is generally understood to identify the copyrighted work, identify the infringing material precisely enough to locate it, give contact details, and include statements of good-faith belief that the use is not authorised and, under penalty of perjury, that the information is accurate and the sender is authorised to act. Most providers publish their own form.
Counter-Notices, and the Limits of a Takedown
The target of a notice may file a counter-notification stating, under penalty of perjury, a good-faith belief that the material was removed as a result of mistake or misidentification, and consenting to jurisdiction. The provider may then restore the material after a waiting period unless the complaining party notifies it that a court action has been filed. At that point the dispute leaves the intermediary and becomes litigation, or nothing.
Two limits follow. A takedown is an administrative process operated by a private company: it produces no findings, no damages and no injunction, and it does not stop the same material reappearing on another host the same week. And a notice sent in bad faith carries its own exposure, because the statute provides for liability where a sender knowingly and materially misrepresents that material is infringing. A takedown aimed at criticism, at a lawful comparison, or at a competitor's own writing is the fact pattern that provision exists for.
Deciding Whether It Is Worth Acting On
The reflex assumption is that a copy is damaging the original's search performance, and it is usually asserted rather than tested. Before spending anything, establish the actual harm: whether the copy appears for queries that matter, whether it reaches your customers, whether it presents itself as you, and whether it plausibly diverts revenue.
Three situations tend to justify a fast response regardless of measurable harm: a mirror using your brand in a way that could deceive customers or collect their data; a copy of gated or commercially sensitive material such as pricing, methodology or documentation; and a copier who is also a direct competitor, where the copying may be evidence in a larger commercial dispute.
The rest is a resourcing judgment. Chasing every republication of a blog post costs more attention than the republication does, and companies that try to enforce comprehensively usually stop within a quarter. Enforce selectively, against a written standard, and keep a record of what was sent and to whom.
When a Notice Arrives About Your Own Site
The dangerous reflex is to ignore it, because the decision is not the company's to make. A host acting to protect its own safe harbour may disable the page, or the account, without waiting for a reply; a search engine may drop the URL from its index; and repeat-infringer policies in hosting and platform terms can escalate to termination of service. Silence does not preserve the position, it delegates it.
A workable internal process is short. Route notices to one named owner rather than to whichever inbox received them. Preserve the page, its revision history and the surrounding logs immediately. Identify the asset and where it came from — agency, stock library, Creative Commons source, employee — and pull the licence evidence for it.
Then decide, with counsel, between removing the material, restoring it through a counter-notification, and contacting the sender. A counter-notification carries a consent to jurisdiction and a statement under penalty of perjury, which is why that decision belongs with a lawyer rather than the web team.
AI Crawlers and Content Used for Training
Robots controls are a voluntary convention and nothing more. A robots.txt file is a request that compliant crawlers read and honour: it is not access control, it does not prevent retrieval, and it works only where the operator chooses to comply. Some operators publish separate user-agent tokens for training and for search, so blocking one may affect visibility in a product the company cares about. Controls that actually restrict access sit on the server — authentication, rate limiting and firewall rules — and they also block crawlers a company wants.
On the law, the honest position is that the questions are open. The Copyright Office addressed generative AI training in Part 3 of its report on copyright and artificial intelligence, released 9 May 2025 as a pre-publication version. Copyright Office reports are the views of an administrative agency and do not bind courts. Whether training on copyrighted works infringes, and whether it is fair use, is unresolved, and this page reports no outcomes from pending litigation.
If It Becomes a Dispute
When copying becomes a formal claim, the technical record decides how much can be argued. What did each site display on specific dates, and in what order were the two versions published? Do the texts share copying signatures — the same typographical errors, the same unusual ordering, the same idiosyncratic phrasing? What do server logs, archive captures and index records establish about first publication? And can the captures be authenticated, which means documenting how, when and by whom they were taken.
That reconstruction is what an expert witness is engaged to perform and explain, and it depends on preservation. Web archives are incomplete, a copied page usually vanishes the moment a notice reaches its host, and log retention on most hosting is measured in weeks. Capturing the other side's pages, and freezing your own records, at the point of discovery rather than the point of filing keeps the question answerable.
Frequently Asked Questions
Someone copied our website content. What can we do about it?
The two routes are a takedown notice to an intermediary and a legal claim. A notice to the host can get material removed quickly, and a notice to a search engine can remove a URL from an index, but neither produces damages or a finding.
Before sending anything, capture the copied pages with dates and URLs visible, and assemble your evidence of prior publication. Copies disappear once a notice lands.
Is a DMCA takedown the same as suing someone?
No. A takedown is a request to a service provider, under the safe harbour framework in 17 U.S.C. 512, that it remove or disable access to material. It is administrative, decided by a private company, and produces no damages, no injunction and no ruling on whether anything was infringed.
It also carries risk: the statute provides for liability where a sender knowingly and materially misrepresents that material is infringing.
A competitor's site describes the same product specifications as ours. Is that infringement?
Usually not. Copyright protects expression, and 17 U.S.C. 102(b) excludes ideas, procedures, systems and methods of operation. Facts about a product — dimensions, ratings, certifications — are not authored, and two companies describing one specification produce similar pages.
A claim becomes arguable in the language: paragraphs of prose, worked examples, distinctive explanations, photographs and diagrams reproduced rather than independently written.
We received a DMCA notice about our own site. What happens if we ignore it?
Ignoring it hands the decision to the host. Providers act on notices to preserve their own safe harbour, which can mean the page, or the account, is disabled without discussion, and repeat-notice policies can escalate to termination.
The practical response is to preserve the page and its history, identify where the asset came from and what licence covers it, then take advice before responding.