Bilby / Methodology
How does Bilby check a store?
Bilby opens the store in a real Chromium browser — desktop 1440×900, then mobile 390×844 — from the store's home region, visits home → collection → product → add-to-cart → cart → checkout (stopping at the payment step), records everything the browser sees, runs deterministic checks on the recording, and asks Claude only what needs a human eye. Every finding carries its evidence; every number prints its basis; anything that didn't run is listed.
// THE VISIT
One visit, one recording, five checks
The journey is the one a shopper takes. Desktop first (judged), then mobile (deterministic), one after the other so the two never compete for the CPU whose speed we measure. A step that fails becomes a fact or a finding with a screenshot; a step Bilby can't map on a theme is recorded as “couldn't map” and never scored. The free scan has a hard 90-second wall clock and the report says what didn't complete; nightly visits get six minutes per viewport.
Bilby never logs in, never edits a theme, never installs an app and never places an order. It sees what a shopper sees.
// THE RECORDING
What is captured, and how privacy is handled
Per viewport: every network request via the Chrome DevTools Protocol (URL, type, status, transfer size, timing, initiator, and the POST body for tracking endpoints), console errors, Web Vitals via PerformanceObserver (LCP, CLS, a TBT proxy from long tasks, TTFB, FCP), the main document's headers and certificate, a DOM snapshot (scripts, theme, prices, stock labels, title/meta/schema), a filmstrip at one frame per second (≤ 40 frames at 480 px), and a screenshot per step.
Personal information is blurred before storage: text that looks like an email address or phone number is blurred in Bilby's own render of the page before any screenshot or frame is captured. Recordings are kept 90 days on paid plans and 30 days for free scans.
// DETERMINISTIC FIRST
What is measured, never guessed
Watch (certificate and domain expiry, mixed content, sitemap top-10 status, the cart endpoint's answer to add-to-cart, search results, form validation), Money (which tracking platforms fired which events on which step, from their own collection endpoints), Speed (vitals per page, weight, the heaviest requests by vendor), Change (the night-to-night diff) and the Answer line (robots.txt, llms.txt, schema) are computed from the recording with no model involved. Speed numbers measured from a region far from the store's market, or while Bilby's own browser was starved (TBT above 8 s), are printed as facts with the caveat — never as findings.
// JUDGEMENT
Where Claude is used, and the rules it works under
Claude (the frozen judge model, currently claude-sonnet-5) sees the screenshot of the home, collection and product pages and is asked one question each: does this look broken, stale or confusing to a first-time customer? It works under thirteen rules — never analytics, never speed, never cookie banners, never geo modals, never empty ad slots, never console warnings — and must quote literal evidence from the screenshot or the finding is discarded. A code backstop suppresses the known bot-versus-customer artefacts regardless of wording. An AI finding can only become a P1 when a deterministic signal agrees.
Every check passed a ten-store, zero-false-positive hand-reviewed gate before it shipped; the gate archives are linked from the Bilby 2.0 report.
// SEVERITY AND SCORE
How a score is made
P1 = the buy path is broken (35 points); P2 = a broken feature or a missing purchase-path event (20); P3 = degraded (10); P4 = worth fixing (5). The score is 100 minus the sum, floored at zero, and a P1/P2 always outranks the arithmetic in the verdict line. Findings are de-duplicated across pages by topic.
// REVENUE AT RISK
Dollars, with their basis
Bilby cannot see a store's traffic or sales. It guesses a monthly revenue bracket from store signals (platform, products listed, third-party apps present), labels it “our guess”, and lets the owner correct it with one slider. Each finding's share of monthly revenue is a range anchored to published conversion research (for example ~7% of conversions per extra second — Akamai, The State of Online Retail Performance, 2017 of load beyond 2.5 s; 30–80% of a path's sales when add-to-cart is broken) and the step's position in the buying path. The method is printed under every figure.
// BENCHMARKS
Percentiles from our own visits, with n
Benchmarks come only from Bilby's own dataset of real-browser visits — never scraped indexes. A percentile is shown only when n ≥ 10 and labelled “early dataset — directional” below 100 visits.
// CHANGE AND HISTORY
Night to night
Each visit writes a normalised snapshot — scripts, theme, prices and stock on the visited product, title/meta/schema, a 64-bit perceptual hash of the homepage, robots.txt and llms.txt state, top-page redirects — and the next visit diffs it. Issues have a state machine (open → escalated after 72 h → resolved) so an alert is sent once, on the transition, never every night.
// LIMITS, PRINTED
What Bilby cannot know
Purchase events (no order is placed). Anything behind a login. Whether a change was intentional. A store that walls automated browsers (a 403 or a challenge page) — reported as an allowlist ask, never as a store defect. Rate limits: a 429 to Bilby is skipped and noted. The same rules run against our own store every night, in public, on the scoreboard.
Check pages: watch · money · speed · change · rivals · search · answer · BilbyBot