Corporate Website Marketing logo — B2B website marketing referenceCorporate Website MarketingB2B website marketing since 2003
From the archive

The Personal Search Engine, Revisited

An archive note on searching your own mail and files, and why that idea is finally arriving as an answer engine.

Personal search did not fail, it was absorbed

Personal search — searching your own mail, files and notes rather than the web — was a real product category for a few years in the middle of the 2000s, and it disappeared not because it was wrong but because it became a feature of the operating system. This page has been on this site since around 2004, when that was not obvious.

The subject of the original note was Lookout, a small tool that indexed Microsoft Outlook mail and returned results in a fraction of the time Outlook's own search took. Microsoft acquired it in that era; the exact terms and date are not something this page will assert. The note was short, of the kind people wrote before blogging conventions settled, and it is not reproduced or quoted here. What it recorded was a moment: several companies shipping desktop search tools at once, and an argument that the index of your own working life was the next frontier. That argument was right about the destination and wrong about the shape.

The desktop search moment

For a stretch of the mid-2000s desktop search was a competitive field. Google shipped a desktop product. Microsoft shipped its own and folded search into Windows and into Outlook. Apple built Spotlight into Mac OS X. Yahoo shipped one on licensed technology. Independent tools — Copernic, X1, Lookout and others — competed on index speed and on how well they handled mail.

The reasoning was sound. Web search had been solved well enough to build a large business on, and the corpus nobody had indexed properly was on your own hard drive: years of mail, attachments, drafts, contracts. It was worth more per document than anything on the web and far worse served.

What nobody in that conversation, including this site, took seriously enough was that a search index over your own files is a plumbing feature. It has no network effect, no advertising model, and no reason to be a separate purchase. A feature every operating system needs and nobody will pay separately for is not a market. Within a few years the standalone products were acquired, absorbed or retired.

The successor is an answer, not an index

What survived was not the index but the premise underneath it: that a person with a question should not have to know which corpus holds the answer. In 2004 that meant one search box over mail and files. Now it means an assistant that reads a question, decides which sources to consult, retrieves from a company's own documents and from the open web at once, and returns a composed answer instead of a list of places to look.

Google documents two such features in Search. AI Overviews are described as helping users get to the gist of a complicated topic or question more quickly, and as a jumping-off point to explore links to learn more. AI Mode is described as particularly helpful for queries where further exploration, reasoning, or complex comparisons are needed. Both are grounded in Google's core Search ranking systems, through retrieval-augmented generation and query fan-out — the latter issuing multiple related searches across subtopics and data sources, so the features display a wider set of links than a classic results page.

Fan-out is the descendant of the 2004 idea: personal search wanted to decompose a question across your inbox; this does it across the web.

What changes when a search engine returns an answer

A list of ten links and a composed answer make different demands on a website. Keywords did not stop mattering; the unit of visibility moved from a position on a page to inclusion in an answer assembled partly from queries nobody typed.

The 2004 modelThe current model
Rank a page for a keyword the user typedBe retrievable for subtopics assembled by query fan-out
Ten links, rankedA composed answer, with links as a jumping-off point
Success measured as positionSuccess measured as inclusion and impressions

The translations for a B2B site are concrete. A page that buries its answer under six paragraphs of positioning is harder to retrieve from than one that states the answer and then supports it. A specification held in a PDF is worse than the same specification in text. A topic covered once, thinly, competes badly against a site covering the same territory in depth across pages that link to each other, because fan-out surfaces subtopics and a subtopic needs a destination.

What Google documents, and what vendors are selling

Google is unusually explicit about what is not required, and the statements contradict most of what is being sold. From its guidance on generative AI in Search: you do not need to create new machine readable files, AI text files, markup, or Markdown to appear in generative AI search, and structured data is not required, with no special schema.org markup you need to add. From the AI features documentation: there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary. Chunking content and rewriting content just for AI systems are named as unnecessary. Google's framing is that SEO best practices remain relevant because its AI features are rooted in core ranking systems.

The controls are the existing ones: nosnippet, data-nosnippet, max-snippet, noindex, robots.txt, plus Google-Extended to limit AI training and grounding in some of Google's other systems. Understand the trade-off first: the snippet controls that keep a page out of AI Overviews also remove ordinary search snippets, and Google documents no way to appear normally in Search while being excluded from AI features. There is no free opt-out.

What this site got wrong then, and what the industry gets wrong now

The error in the archived note was category-chasing. Personal search was written up as an emerging field to have an opinion about, which is what a site publishing short notes in 2004 did, and which reliably produced coverage of features mistaken for markets. The useful question — whether the capability had an independent business behind it — was not asked, and the answer was in the economics rather than the technology.

The current version is louder and more expensive. A generative-engine-optimisation industry sells proprietary markup, file formats, content chunking, answer-first templates and visibility scores, on the claim that AI search needs a new discipline. Against that, hold the documentation quoted above. Two claims deserve the same treatment:

  • Traffic-impact figures. No primary source publishes a figure for how much traffic AI Overviews take, or for click-through decline where they appear. Third-party studies use incomparable methods and disagree wildly.
  • AI referral share. Circulating figures range from around one per cent of all web traffic to claims about share within AI referrals, because they measure different things. Any specific percentage is a vendor estimate.

The direction of travel is documented; the magnitude is not.

Measuring visibility you cannot click on

Google states that traffic from AI features is included in Search Console's Performance report under the Web search type, so it has never been invisible — it has been mixed in. In June 2026 Google announced dedicated Search generative AI performance reporting. Coverage of that announcement, secondary rather than Google's own text, describes it as impressions-first, without click data and without separating AI Overviews from AI Mode; check that against the live documentation before building a report on it. For a B2B programme:

  • Expect impressions to lead clicks. Visibility inside an answer can rise while sessions do not. That is information about coverage, not a failure of the page.
  • Instrument assistant referrers. They are identifiable by hostname, and GA4 has an AI Assistants default channel group. Report it; do not forecast on it.
  • Ask buyers where they came from. A self-reported source field is crude, and where influence arrives without a referrer it is one of the few signals available.
  • Stop reporting average position as a headline. Against a composed answer it is close to meaningless.

What to do about it, which is mostly what you should already be doing

There is no new discipline to buy here. What a website needs in order to be retrieved from is what it needed to rank, applied more strictly:

  1. Answer the question in the first two sentences, then support it at length. This serves an impatient buyer and a retrieval system for the same reason.
  2. Give every real question its own page, and cover the subtopics around it. A subtopic with no destination on your site is one your competitor answers.
  3. Put the facts in text. Specifications, compatibility, integrations, standards you comply with, how pricing is structured. A fact in a PDF is harder to retrieve and to cite.
  4. Publish what only you can publish. Original data, tested methods, honest accounts of where your product is the wrong choice.

Keep the ordinary technical requirements met as well — crawlable, indexable, fast enough — because Google documents that its AI features are grounded in core ranking. The 2004 note was about a tool that searched your mail. What it was watching was the start of a shift from finding documents to getting answers, and that shift has now reached the channel most B2B companies depend on for pipeline. The response is unglamorous: fewer pages, written better, saying things specific enough to be worth repeating.

Frequently Asked Questions

Do we need special markup or an llms.txt file to appear in AI Overviews?

No. Google documents that you do not need to create new machine readable files, AI text files, markup, or Markdown to appear in generative AI search, that structured data is not required, and that there are no additional requirements to appear in AI Overviews or AI Mode.

llms.txt is an independent proposal, not a standard, and no major AI vendor's documentation confirms consuming it in production. Publishing one is harmless; paying for it as a visibility service is not.

How much traffic do AI Overviews take away from websites?

No primary source publishes a figure, and this page will not invent one. Google does not report AI Overviews traffic impact, and Search Console's generative AI reporting does not currently expose click data, so the measurement needed is not publicly available.

Third-party studies use different panels, definitions and baselines, and their conclusions are mutually contradictory. AI features clearly change how results are presented; the magnitude of the traffic effect is not established.

Can we appear in Google Search but stay out of AI Overviews?

Not without a cost. The documented controls are the existing ones — nosnippet, data-nosnippet, max-snippet, noindex and robots.txt — and the snippet controls that keep content out of AI Overviews also remove your ordinary search snippets. Google documents no mechanism for appearing normally in Search while excluded from AI features.

Google-Extended is sometimes suggested as the answer. It is documented as limiting AI training and grounding in some of Google's other systems, not as an AI Overviews opt-out.

Is desktop or personal search still a product category?

Not as a separate purchase. Search over your own mail and files is now built into every major operating system and mail client, and the standalone tools of the mid-2000s were acquired, absorbed or retired.

The capability the category reached for has reappeared as assistants that answer from an organisation's own documents and the open web at once — largely bundled into productivity suites rather than sold as search. Internal and external retrieval are converging, and content written for people travels better across both.