Public benchmark: raw answers
The phrases we use to test the assistant on our demo shops, what it answered, and how we ran it. Nothing here is a claim about other products.
Method
For each demo shop of a niche we send every shopper phrase from that niche’s list, in each language the list has, plus extra phrases we add by hand. Each phrase is asked once, with no cache and no earlier conversation, through the same pipeline a shopper’s question takes.
We record only what can be measured: whether the assistant showed a product or said it had nothing, the first product shown, the one-sentence reason it gave for that product, and the response time in milliseconds. We do not grade whether an answer was right, because that needs a labelled sample; this page is the raw record so you can judge each row yourself.
Demo shops are public demonstration catalogues, not customer stores. The phrase lists are written by us, so they are not a random sample of shoppers. Response time includes the language-model calls and varies with load.
Raw answers
Not yet run. This page shows results only after the benchmark has been run on our demo shops; until then the table is empty on purpose, and no figures appear here.