This page explains a research study being conducted from this domain, what it measures, how it collects data, and how to have your company excluded from it. If you found this page from a user agent string in your server logs, the section on collection conduct is the one you want.
What the study is
The Corporate Website Asset Study measures what can be established about a company's website from the public record: who controls the domain, how the site performs for real users, what third-party services the site sends visitor data to, whether the privacy policy discloses those services, and whether the site meets basic automated accessibility checks.
The study exists because these five things are routinely assumed rather than verified. A company assumes it owns its domain. It assumes its privacy policy describes what its website actually does. It assumes someone checked. In our consulting and expert witness work we have repeatedly found otherwise, and anecdote is not evidence. The study is an attempt to replace a set of professional impressions with measured proportions.
Findings will be published as a free report on this site and will inform a forthcoming book on corporate website marketing.
What it measures
- Domain control. Registration data from RDAP (Registration Data Access Protocol, the structured successor to WHOIS): the registrant organization where disclosed, registry and registrar lock status, DNSSEC delegation, expiry date, and sponsoring registrar. Where the registrant is redacted or withheld, that is recorded as a non-determination, not as a mismatch.
- Core Web Vitals. Field data from the Chrome UX Report, mobile, at the origin level. No page is loaded to obtain this; it is an API query against Google's aggregated dataset.
- Third-party data flows. The third-party services a page contacts when loaded, recorded from the browser's own network log.
- Privacy policy disclosure. Whether each observed third-party organization is named in the site's privacy policy, described by category, or not disclosed at all.
- Automated accessibility failures. Individual axe-core rule violations on the homepage, reported by rule rather than as a composite score, because composite accessibility scores are not meaningful.
Which companies are included
Two groups. The first is the Fortune 500, measured completely rather than sampled. The second is a random sample of US mid-market business-to-business companies with roughly 50 to 999 employees, drawn from a publicly available federal filing dataset so that any reader can reconstruct the sample frame without buying a commercial database. The exact frame definition, filter, sample size, and random seed are published in the protocol before collection begins.
Inclusion is not a judgment about any company. It is the result of a random draw or of appearing on a published list.
Collection conduct
The user agent string. If you are here from your server logs, this is what you saw:
CorporateWebsiteAssetStudy/1.0 (+https://www.corporatewebsitemarketing.com/research/; Hello@Hartzer.com)
Requests come from a US address, at most one per second per host, and honor robots.txt. To block it, disallow CorporateWebsiteAssetStudy in your robots.txt and it will stop on its next pass — or email us and we will remove the domain outright.
These constraints are part of the published protocol, not courtesies we extend at our discretion.
- The crawler identifies itself. It sends a user agent naming the study and linking to this page, so that anyone reading their own logs can find out what it is and who to contact. It does not disguise itself as a browser, a search engine, or anything else.
- robots.txt is respected for all crawling.
- Rate limited to no more than one request per second per host. The full run is deliberately spread over days rather than hours. The traffic should be invisible against ordinary background load.
- Public, unauthenticated pages only. No logins. No form submissions. No purchases. No interaction beyond clicking a cookie consent control where one is present, because consent behavior is one of the things being measured.
- A very small number of pages. The homepage and, at most, a shallow crawl to locate one page containing a form. Not a site-wide crawl.
- No circumvention. Where bot protection blocks the crawler, the site is recorded as blocked and skipped. We do not rotate addresses, change user agents, solve challenges, or otherwise attempt to evade a block. The proportion of sites that blocked us is reported as a finding, because it is one.
- No personal data is collected or stored at any point. The study records the behavior of websites, not of people.
What will and will not be published
Results are published in aggregate only. No company is named in connection with any negative finding.
Reporting true facts derived from public data would be lawful. It would also convert a research report into an accusation, invite disputes with the companies named, and compromise the investigator's position as a neutral expert. We think the report is more useful without that, and considerably more citable.
Individual company results are retained privately so the analysis can be verified, and are not published. Where we find something materially serious and specific to one company, the appropriate response is a private notice to that company, not a line in a report, and that is what we will do.
Requesting exclusion
Any company may be excluded on request, before or after collection, without giving a reason and without any discussion of the findings.
Email Hello@Hartzer.com with the subject line Study exclusion and the domain you are asking to exclude. A request from a company email address at that domain is sufficient; we will not ask you to prove anything further.
On receiving a request we will remove the domain from the collection list, delete any data already collected for it, and confirm in writing. Excluded domains are counted in the published exclusion total, which is reported as a single number with no domains named.
If you would prefer to block the crawler yourself rather than email us, a robots.txt directive is honored and requires nothing from us.
Methodology and pre-registration
The full protocol is written and published before collection begins, including the sample frame, the measurement definitions, the coding rules, the analysis plan, and the limitations. This matters: a study whose method is fixed in advance cannot quietly become a different study once the results are in, and a reader can check that it did not.
The dataset and the analysis scripts will be published alongside the report, excluding the rows that would identify individual companies, so the analysis is reproducible. Limitations are published with the findings rather than buried behind them, and there are several worth stating plainly: registration data redaction limits what can be concluded about domain registrants, automated accessibility testing detects only a minority of real barriers, a browser network log cannot see server-side tracking at all, and every measurement is a snapshot of one collection window.
Who is conducting it
The study is conducted by Bill Hartzer of Hartzer Consulting. Bill has worked in search, domains, and website technology since 1996 and serves as an expert witness in matters involving search engines, digital advertising, web analytics, and domain names. This site is an independent publication and a resource of Hartzer.com.
Questions about the method, requests for the protocol, and corrections are all welcome at Hello@Hartzer.com. Corrections in particular: if you believe a measurement is wrong, we would rather hear it before publication than after.