skip to content
The Weighted Average

AI Economics for Operators

Onton's Search Benchmark Earns a Pilot, Not a Switch

Onton finds 0.87 extra relevant-result equivalents per query versus Google, but model judges and artifact gaps limit the verdict.

Room with different styles of chairs
Room with different styles of chairs. Photograph by Declan Sun

Onton’s Ontology 1 benchmark produces 0.87 additional relevant-result equivalents per query versus Google on 90 difficult decor searches. That is enough to justify a retailer pilot, but not a production switch: the company designed the challenge set, used no human relevance labels, and publishes neither a merchant API nor enterprise price or SLA.

A strong result on Onton’s home turf

Onton’s canonical benchmark reports mean precision at 10 of 63.0% for Onton, 54.3% for Google Shopping, and 46.9% for Amazon. Across 900 result slots, the 8.7-point gap against Google equals about 78.3 additional judged-relevant result equivalents, or 0.87 per query. Against Amazon, the 16.1-point gap becomes 1.61 per query. These are averages of model scores, not literal human-approved products or purchases.

The win count needs equally careful accounting. Onton won 52 of 90 searches outright, Google won 19, and Amazon won 16. Those add to 87, not 90, because Onton returned fewer than ten results on three queries and missing slots counted as irrelevant under the published protocol. “52 of 90” is accurate; “52 of 87” or “three ties” would not be.

The public dataset card describes Subtext-Decor-90 as curated around aesthetic language, negation, cultural references, and emotional framing. That is a valuable failure cluster—“a chair that feels Bauhaus but not cold” is harder than a SKU lookup—but it is also the terrain where Onton’s explicit domain structure is meant to shine. The set excludes multimodal queries from the main 90 and does not represent an ordinary traffic-weighted catalog mix.

Hard specifications remain the counterexample. Onton’s own failure analysis says Amazon often wins functional-specification searches and attributes some Onton misses to catalog breadth. A decor marketplace with inspiration-led sessions should care; a retailer dominated by brand, size, inventory, and exact attributes should not infer the same lift.

Google is also a moving baseline. Google says its Shopping Graph contains more than 50 billion listings and refreshes over two billion an hour, while AI Mode supports conversational, multimodal refinement. Benchmark capture date, region, account state, personalization, ads, and whether “Google” means classic Shopping or AI Mode can alter a screenshot result. Onton itself calls for longitudinal reruns.

The benchmark used three frontier-model judges and up to 90 × 3 × 10 × 3, or 8,100 judge-result assessments. That is up to 90 nominal assessment slots per query. By comparison, Amazon’s ESCI paper reports roughly 2.6 million human query-product judgments across 130,000 queries, about 20 per query. Onton has 4.5 times the nominal assessment density but only 0.069% as many queries, substituting automated depth for human breadth.

This is the evaluation analogue of today’s Meta Business Agent billing lead: the published unit becomes meaningful only after its semantics are translated. It also connects to the archive’s agent-discovery analysis, where finding the right product or service becomes infrastructure rather than a blue-link feature.

Pilot the failure cluster and repair the benchmark

Judge agreement is the first repair. Onton reports Krippendorff’s alpha of 0.465. All three models rank Onton first on aggregate, but their absolute precision differs materially: Opus scores Onton at 0.759, Gemini at 0.528, and GPT at 0.603. NIST explicitly warns against using LLMs to create TREC-style relevance judgments, favoring trained and monitored human assessors. A retailer should add blinded merchandisers or customers and preserve a human-only holdout.

The artifact is not yet clean enough for a strong reproducibility claim. The Hugging Face repository tree exposes public materials, but the card still contains a placeholder canonical URL, asks maintainers to fill in exact paths and formats, and labels its layout an example. A committed virtual environment adds noise rather than provenance. A pinned commit, final schema, judge prompts, API parameters, capture metadata, bootstrap intervals, and one-command rerun would materially strengthen the result.

Commercial readiness is a separate gate. Onton’s partnership page invites merchants to list products and contact the company, but publishes no search API, feed schema, latency distribution, uptime SLA, data-residency terms, relevance controls, enterprise pricing, or rollback procedure. Consumer subscription pricing cannot stand in for merchant search infrastructure. Migration cost is therefore unknowable from public evidence.

Known alternatives at least make request economics visible. Google AI Commerce Search lists $2.50 per 1,000 requests, while Algolia’s public pricing page includes 10,000 requests and then lists $0.60 per additional 1,000. At one million monthly requests, that is about $2,500 versus roughly $594 for Algolia’s request component—(1,000,000 − 10,000) ÷ 1,000 × $0.60—before records, support, discounts, and implementation. Onton cannot enter that comparison until it publishes or quotes equivalent terms.

The right buyer is a furniture or decor retailer with long natural-language discovery queries, rich imagery, and a measured aesthetic or negation failure cluster. Export 500–1,000 real queries, preserve traffic weights while oversampling that cluster, freeze catalog and availability, and compare incumbent, Onton, and a modern commerce baseline. Randomize engine positions and collect blinded human grades before shadowing production.

An online test should gate on latency, zero-result rate, add-to-cart, conversion, gross margin, return or cancellation, and seller diversity. Keep the incumbent as instant fallback. The verdict changes from pilot to switch only when a preregistered, human-judged evaluation shows a reliable relevance gain and live traffic converts it into commercial lift net of latency and returns—alongside acceptable API, freshness, price, security, and SLA terms.

It changes to no-adopt if the gain disappears on easy or specification queries, humans reverse the model ordering, or conversion fails to move. The archive’s analysis of commercial incentives in AI discovery is the other reason to include margin and seller diversity: relevance alone is not the entire marketplace objective.

Onton has produced an intriguing capability demonstration. 0.87 extra relevant-result equivalents per query is large enough to test. With an alpha of 0.465, no human labels, incomplete artifacts, and absent commercial infrastructure, it is not large enough to trust on faith.

Sources