Original research

AI Agent Outcome Verification Discoverability Benchmark 2026.

How do AI answer engines discover and describe an early-stage product when buyers ask by problem, and when they ask by name? We ran a fixed set of 9 prompts, in English and Hebrew, across 5 engine surfaces, 116 recorded runs in total, and logged every appearance, description, error, and cited source. Headline result: engines that answer from training knowledge describe Pruvz accurately, and engines that answer strictly from citable sources almost never find it.

By Dor Sharoni, founder of Pruvz · Runs captured July 29-30, 2026 · Published August 3, 2026 · Download the dataset (CSV)

Disclosure, up front: this is first-party research. Pruvz measured how answer engines discover and describe Pruvz. That makes the benchmark a documented measurement of our own discoverability whose aggregates are recomputable from the published run-level labels, and it makes it the opposite of independent third-party validation; nothing here should be read as an external endorsement of the product. The prompt set was frozen before the first run and will not be changed to improve future results, and the run-level data is downloadable below so you can recompute every aggregate on this page yourself. What the public data cannot give you is the labels' own ground truth: the full answers behind them are retained internally (see limitations).

Why we ran this.

The benchmark started with a real failure: a buyer-style market-overview prompt, asking exactly the question Pruvz answers (business evidence that an agent-executed refund really happened), returned market maps with no Pruvz in them. Before publishing more content, we wanted a baseline precise enough to measure whether any of it works: which intents surface vendors at all, what the engines already believe about Pruvz, and which sources they actually cite in this category. That original prompt is preserved verbatim as P0.

Methodology.

Nine frozen prompts (P0-P6 problem-led and category-led, P7-P8 brand-led), each run verbatim in English and Hebrew, one prompt per fresh chat, two runs per prompt/engine/language combination on the fully covered engines. Clean personas throughout: logged out where possible, otherwise a dedicated clean account with no history, and never a personal account. Runs were executed from an Israeli IP on July 29-30, 2026. Every run's date, engine, model label when shown, answer, and cited sources were recorded immediately.

The frozen prompt set, in English and Hebrew
IDIntentEnglish promptHebrew prompt
P0Problem-led (the original failed buyer prompt)Give me a market overview: which companies provide evidence for agents for products in production. Say we have a product in production that gives customers refunds via agents, and I want business evidence that the refund really happened.תעשה לי סקירה בשוק, איזה חברות נותנות אבידנס לאייג'נטים למוצרים בפרודקשיין. נגיד יש לנו מוצר בפרודקשיין שנותן ריפאונד ללקוחות על ידי אייג'נטים ואני רוצה ראיות עסקיות שהריפאונד באמת קרה .
P1Problem-ledMy AI agent says it completed an action, like issuing a customer refund. How can I verify the action actually happened in my business systems?סוכן ה-AI שלי טוען שהוא ביצע פעולה, למשל זיכוי ללקוח. איך אני יכול לוודא שהפעולה באמת בוצעה במערכות העסקיות שלי?
P2Problem-ledWhat tools can verify AI agent actions against systems of record like CRM and billing, instead of trusting the agent's own logs?אילו כלים יכולים לאמת פעולות של סוכני AI מול מערכות כמו CRM ומערכת חיוב, במקום לסמוך על הלוגים של הסוכן עצמו?
P3Problem-ledHow do enterprises create an audit trail for AI agent decisions and actions, including which policy applied at decision time?איך ארגונים בונים מסלול ביקורת (audit trail) להחלטות ופעולות של סוכני AI, כולל איזו מדיניות חלה ברגע ההחלטה?
P4Problem-ledAI agent observability tools show me traces of model and tool calls, but how do I prove the actual business outcome of an agent's action?כלי observability לסוכני AI מראים לי traces של קריאות מודל וכלים, אבל איך אני מוכיח את התוצאה העסקית בפועל של פעולת הסוכן?
P5Problem-ledWhat are the best tools in 2026 for verifying the business outcomes of production AI agents?מהם הכלים הטובים ביותר ב-2026 לאימות תוצאות עסקיות של סוכני AI בפרודקשן?
P6Category and alternatives-ledAre there alternatives to agent observability platforms that verify AI agent outcomes at the business level, by reading back from the systems the business trusts?האם יש חלופות לפלטפורמות observability לסוכנים שמאמתות תוצאות של סוכני AI ברמה העסקית, על ידי קריאה חוזרת מהמערכות שהעסק סומך עליהן?
P7Brand-ledWhat is Pruvz?מה זה Pruvz?
P8Brand-ledIs Pruvz (pruvz.ai) a real product? What does it do and what is its current status?האם Pruvz (pruvz.ai) הוא מוצר אמיתי? מה הוא עושה ומה הסטטוס הנוכחי שלו?
Engine surfaces, their configuration, and headline results
Engine surfaceConfiguration and account stateGrounding observedProblem-led discoveryBrand-led runs
Perplexityperplexity.ai default mode, fully logged outSearch-grounded: answers built from cited sources0 / 288 / 8 engaged (see accuracy below)
ChatGPTchatgpt.com default model, fully logged outParametric in most runs (little or no browsing observed)3 / 288 / 8 engaged
Geminigemini.google.com default model (Flash), dedicated clean account, activity saving declinedParametric in most runs3 / 288 / 8 engaged
Claude APIclaude-sonnet-5 via the Anthropic API with the web_search tool, API key with no memory; partial, budget-truncated coverageSearch-grounded1 / 5not run
Claude subagentFresh-context Claude Code subagents with web search; partial coverage with a declared priming caveatSearch-grounded2 / 3not run
The consumer claude.ai product refuses automated sessions, so it is not covered; the two Claude rows are partial substitute surfaces with single runs per combination and declared caveats (see limitations). The subagent surface's working context carried a category-adjacent directory name, a bias that runs toward finding Pruvz, so its two finds count only because the agent cited externally retrieved URLs, while a miss on that surface is a strong signal.

Result 1: discovery is rare, and phrasing decides it.

Across all problem-led and category-led runs, Pruvz appeared 9 times in 92 runs. Perplexity, the one fully covered engine that answers strictly from citable sources, found Pruvz 0 times in 28 discovery runs. Where Pruvz did appear, the phrasing was market-overview, tools-that-verify, or alternatives-to language; the how-to phrasings produced methodology answers with almost no vendor names at all.

Pruvz discovery results by prompt
PromptIntent, condensedRuns where Pruvz appeared
P0Market overview (the original failed prompt)5 / 14
P1Verify a claimed action0 / 12
P2Tools that verify against systems of record1 / 14
P3Audit trail including decision-time policy0 / 12
P4Beyond observability traces0 / 12
P5Best tools in 20260 / 13
P6Alternatives that verify at the business level3 / 15

Result 2: when engines know Pruvz, they describe it accurately.

The 24 brand-led runs on the three fully covered engines split as follows: 17 described Pruvz accurately, 4 were accurate with minor issues (invented example use cases, a partial category mislabel, and a correct founder attribution the site itself did not source at the time), 2 answered about a different company entirely, and 1 denied the product exists. No run invented product capabilities: the positioning, the evidence-packet contents, and the current design-partner status were reported correctly wherever Pruvz was actually described, in both languages.

Result 3: the Hebrew entity collision.

Both Hebrew runs of "What is Pruvz?" on the search-grounded engine answered about Provision-ISR, an unrelated Israeli CCTV manufacturer, asserting that "Pruvz" is a colloquial short form of "Provision" and citing that company's site. A Hebrew speaker asking about the brand was handed a different company twice out of two attempts. The parametric engines showed no such confusion, which localizes the defect: at the time of the runs there was no Hebrew-language source for the Pruvz entity anywhere, so a retrieval engine corrected the query toward a phonetically similar brand with Hebrew coverage. The English runs of the same prompt on the same engine described Pruvz accurately, and one English status run denied the product exists in its first attempt while its second attempt fetched the site and answered correctly, a reminder that these engines are not deterministic.

The Hebrew entity page published in response

Result 4: what the engines cite in this category.

Runs with captured citations named 64 distinct hostnames. The pattern: vendor-authored comparison and buyer-guide content dominates the citable surface for this category, YouTube was the most-cited external hostname in the search-grounded discovery runs, and Hebrew prompts drew citations from Israeli sites that no one in this category is contesting. Every hostname cited in two or more runs:

Most-cited hostnames across captured-citation runs
HostnameRuns citing itNote
pruvz.ai9Cited in brand-led runs, plus two search-grounded discovery finds
youtube.com6The most-cited external hostname in Perplexity discovery runs
pedowitzgroup.com3Vendor-authored guidance content
cflowapps.com2Vendor-authored workflow content
confident-ai.com2Vendor-authored tool comparisons
digitalapplied.com2Vendor-authored statistics and buyer guides
loginradius.com2Vendor-authored engineering content
medium.com2Individual engineering posts
moticohenadv.com2Israeli legal site, cited on Hebrew prompts
provision-isr.com2The wrong-entity answer on Hebrew brand queries
skysoftconnections.com2Vendor-authored CRM data-auditing content
spd.co.il2Israeli site, cited on Hebrew prompts
stackone.com2Vendor-authored use-case content
usefini.com2Vendor-authored buyer guides
The full 64-hostname table, including every single-citation hostname, is in the downloadable aggregates file. Hostnames are www-stripped but not collapsed to registrable domains, so subdomains count separately. One more observation from the answer texts: engines repeatedly grouped Pruvz into category labels of their own invention, such as "business evidence" and "agent systems of record", and some of the peer companies they named alongside Pruvz in those categories could not be verified to exist when we checked for any web presence as of the run dates. Category vocabulary exists in the models; a citable, real occupant of it mostly does not.

Interpretation (ours, clearly labeled).

We read the data as follows; the observations above stand on their own if you read it differently. First, the results suggest the citable external surface may be a larger constraint than brand awareness: the parametric engines describe Pruvz correctly while the retrieval engine could not cite it. The baseline cannot separate the possible causes (indexing gaps, on-site content that does not match the queries, missing external authority, retrieval quality, or plain run-to-run variance); the post-release rerun will test whether the new on-site and external sources change retrieval. Second, discovery followed comparison intent in these runs: vendors surfaced when buyers asked for market overviews, tools, and alternatives, which is why we published an honest alternatives and build-vs-buy page and problem-led solution pages for outcome verification, refund verification, and system-of-record verification. Third, the Hebrew collision motivated us to publish a Hebrew-language anchor for the entity; whether that anchor is sufficient, or whether indexing, external links, and entity profiles are also needed, is exactly what the rerun will measure. And fourth, YouTube's prominence in the citation data is part of why the product demo is published as a video with a full transcript. This page itself is part of the response: it is the citable, inspectable account of the measurement.

Limitations, in full.

(1) First-party research: Pruvz measured Pruvz; treat every editorial choice accordingly. (2) The consumer claude.ai product is not covered at all; the two Claude surfaces are partial substitutes with single runs per combination, one budget-truncated and one with a priming caveat, and neither represents the claude.ai product experience. (3) Citation URLs were lost on all 5 recorded Claude API runs: a token-saving flag on the runner also dropped the citation payload, so those runs are marked in the dataset and excluded from the citation aggregates. (4) Three further Claude API calls appear in our internal cost ledger but their answers were never retained, including one rejected zero-search configuration test; they are excluded from the dataset entirely rather than counted. (5) ChatGPT and Gemini answered from parametric knowledge in most runs, so their discovery numbers measure training-data recall more than live retrieval. (6) Runs were executed from an Israeli IP and engines may localize. (7) Answer engines are nondeterministic: two runs per combination bound the variance but do not eliminate it, and single-run combinations on the Claude surfaces bound nothing. (8) Full answer texts and screenshots are retained in our internal research archive and are not republished here, because redistribution rights for third-party model output are unclear; the dataset carries structured fields and short characterizations instead. (9) The qualitative labels (appearance, accuracy, error category) were assigned by a single coder, the author, reading the full answers against the criteria below, with no independent second rating; the published data makes the aggregates recomputable from those labels, but the labels themselves cannot be independently verified without the retained answers.

How the labels were assigned.

One coder (the author) read every full answer and assigned the closed error-category vocabulary as follows. "Accurate" means the answer described Pruvz's actual positioning and status with no false claims. "Accurate with minor issues" means the description was substantially correct but included one of: example use cases the site does not claim (invented_use_cases), a partially wrong category framing (category_mislabel), or a factually correct statement the site did not support at the time (founder_named_without_source). "wrong_entity" means the answer was about a different company. "existence_denied" means the answer asserted the product is not real. Appearance (pruvz_appears) means the answer engaged with Pruvz by name at all, which is why the two wrong-entity runs count as engaged but not as accurate.

The dataset.

One CSV row per recorded run, plus an aggregates file with every number this page cites, both under CC BY 4.0. If you use it, attribute "Pruvz (pruvz.ai)".

Download the data.

116 recorded runs, July 29-30, 2026: appearance, rank, error category, citation capture state, and cited hostnames per run, with aggregates computed from exactly these rows.

Dataset column dictionary
ColumnMeaning
run_idP{prompt}_{engine}_{language}_r{run}, e.g. P1_chatgpt_en_r1
run_dateLocal Israel calendar date of the run (YYYY-MM-DD)
enginechatgpt, gemini, perplexity, claude-api, or claude-agent
model_versionModel label as shown by the engine, when available
languageen or he
prompt_idP0 through P8 (full texts in the methodology section above)
prompt_familydiscovery (P0-P6) or brand (P7-P8)
run_numberRepetition index within the prompt/engine/language combination
account_statelogged-out, clean-account, or the declared substitute-surface state
pruvz_appearsWhether the answer engaged with Pruvz at all (yes/no)
pruvz_rankPosition among named products when Pruvz appeared
error_categorynone, wrong_entity, existence_denied, founder_named_without_source, invented_use_cases, or category_mislabel
citation_capturecaptured, none_cited (no sources shown), or lost (capture failure, see limitations)
cited_source_countNumber of cited URLs, when capture succeeded
cited_domainsDeduplicated hostnames of the cited URLs (www-stripped; subdomains kept, so community.hubspot.com stays distinct from hubspot.com), semicolon-separated

What happens next.

The identical frozen prompt set will be re-run after the pages this baseline motivated have been published and indexed, with the same engines, the same repetition rules, and the same declared Claude limitations, and this page will be updated with a clearly dated post-release section comparing discovery, accuracy, and citation sources against the baseline above. The prompts will not be changed to improve the result. Corrections: if you find an error in the data or believe a description here is inaccurate, write to hello@pruvz.ai.

Written and maintained by Dor Sharoni, founder of Pruvz. Product names and trademarks belong to their respective owners; Pruvz is not affiliated with or endorsed by the companies referenced on this page. Related: alternatives and build vs. buy, the recorded product demo, and a real evidence packet, explained.