The Machine Time Capsule: What Machines Could Read on 100 Small-Business Websites, Sealed August 2026
Every six months, we re-read the same 100 small-business websites exactly the way machines do - titles, headlines, schema, llms.txt, AI-crawler access - and publish the drift. Nobody is recording how the machine-readable web changes over time. Now somebody is.
Baseline: what machines found in August 2026
Of the 78 cohort sites reachable at sealing time:
* Only 2 of 78 publish machine-readable FAQ answers (FAQPage schema) - the single most quotable format for AI assistants.
* 17 of 78 have no H1 headline at all - to a machine, the page has no declared topic. Another 18 have multiple competing H1s.
* 14 of 78 have no meta description - nothing for machines to quote in results.
* 13 of 78 already publish a genuine llms.txt - consistent with our AI Lockout study finding that small businesses adopt faster than reported.
* 5 of 78 fully block at least one major AI crawler in robots.txt.
Why a time capsule
Everyone measures the web as it is today. Nobody measures the same sites again and again, on a fixed schedule, with a fixed method, and publishes the change. Yet the change is the story of this decade: are small businesses adapting to AI-mediated discovery, or being left behind? From now on, that question has a public, longitudinal answer - starting from this page.
What we predict we'll see (on the record, so we can be wrong publicly): llms.txt adoption will keep climbing; FAQ schema will stay rare because it requires real writing, not a plugin switch; and accidental AI-crawler blocks will slowly disappear as plugin defaults change. February 2027 will tell.
Method
The cohort is drawn deterministically from our AI Lockout study sample (sorted, evenly stepped - no cherry-picking), fixed forever: 20 dentists, 20 law firms, 20 solar installers, 30 real estate agencies, 10 plumbers. Each snapshot fetches each site's homepage, robots.txt and llms.txt with a transparently identified research crawler and records only objective, machine-readable facts: title, H1 count, description length, schema types, FAQ count, JS-free word count, AI-bot access, llms.txt presence. No opinions, no scores. Sites are anonymized in our commentary; the raw data lists domains because every fact in it is public and re-checkable in a browser.
Download Snapshot #1 - sealed August 5, 2026 (JSON)
The sealing promise
Snapshot #1 is now frozen. We will not edit it - not to fix our mistakes, not to flatter the trend, not for any reason. If we find an error in our method, we will document it in the next snapshot's notes rather than rewrite history. Re-openings: every February and August, starting February 2027.
Common questions
What happens if a cohort site goes offline?
It stays in the cohort and is recorded as unreachable - that is data too. Small-business websites dying quietly is part of the story we are documenting.
Can my business join the capsule?
No - the cohort is fixed forever so the comparison stays honest. But you can run every check we run, free, at Through Machine Eyes.
Why should anyone trust these numbers?
Don't trust - verify. The raw JSON lists every domain and every recorded fact; each one can be re-checked in a browser in seconds. Our method is published, our crawler identifies itself, and the file is sealed with a public date.
All SynapseIN research: the research hub · claims ledger · proof in 90 seconds.