AI Agent Outcome Verification Discoverability Benchmark 2026.
How do AI answer engines discover and describe an early-stage product when buyers ask by problem, and when they ask by name? We ran a fixed set of 9 prompts, in English and Hebrew, across 5 engine surfaces, 116 recorded runs in total, and logged every appearance, description, error, and cited source. Headline result: engines that answer from training knowledge describe Pruvz accurately, and engines that answer strictly from citable sources almost never find it.
By Dor Sharoni, founder of Pruvz · Runs captured July 29-30, 2026 · Published August 3, 2026 · Download the dataset (CSV)
Disclosure, up front: this is first-party research. Pruvz measured how answer engines discover and describe Pruvz. That makes the benchmark a documented measurement of our own discoverability whose aggregates are recomputable from the published run-level labels, and it makes it the opposite of independent third-party validation; nothing here should be read as an external endorsement of the product. The prompt set was frozen before the first run and will not be changed to improve future results, and the run-level data is downloadable below so you can recompute every aggregate on this page yourself. What the public data cannot give you is the labels' own ground truth: the full answers behind them are retained internally (see limitations).
Why we ran this.
The benchmark started with a real failure: a buyer-style market-overview prompt, asking exactly the question Pruvz answers (business evidence that an agent-executed refund really happened), returned market maps with no Pruvz in them. Before publishing more content, we wanted a baseline precise enough to measure whether any of it works: which intents surface vendors at all, what the engines already believe about Pruvz, and which sources they actually cite in this category. That original prompt is preserved verbatim as P0.
Methodology.
Nine frozen prompts (P0-P6 problem-led and category-led, P7-P8 brand-led), each run verbatim in English and Hebrew, one prompt per fresh chat, two runs per prompt/engine/language combination on the fully covered engines. Clean personas throughout: logged out where possible, otherwise a dedicated clean account with no history, and never a personal account. Runs were executed from an Israeli IP on July 29-30, 2026. Every run's date, engine, model label when shown, answer, and cited sources were recorded immediately.
| ID | Intent | English prompt | Hebrew prompt |
|---|---|---|---|
| P0 | Problem-led (the original failed buyer prompt) | Give me a market overview: which companies provide evidence for agents for products in production. Say we have a product in production that gives customers refunds via agents, and I want business evidence that the refund really happened. | תעשה לי סקירה בשוק, איזה חברות נותנות אבידנס לאייג'נטים למוצרים בפרודקשיין. נגיד יש לנו מוצר בפרודקשיין שנותן ריפאונד ללקוחות על ידי אייג'נטים ואני רוצה ראיות עסקיות שהריפאונד באמת קרה . |
| P1 | Problem-led | My AI agent says it completed an action, like issuing a customer refund. How can I verify the action actually happened in my business systems? | סוכן ה-AI שלי טוען שהוא ביצע פעולה, למשל זיכוי ללקוח. איך אני יכול לוודא שהפעולה באמת בוצעה במערכות העסקיות שלי? |
| P2 | Problem-led | What tools can verify AI agent actions against systems of record like CRM and billing, instead of trusting the agent's own logs? | אילו כלים יכולים לאמת פעולות של סוכני AI מול מערכות כמו CRM ומערכת חיוב, במקום לסמוך על הלוגים של הסוכן עצמו? |
| P3 | Problem-led | How do enterprises create an audit trail for AI agent decisions and actions, including which policy applied at decision time? | איך ארגונים בונים מסלול ביקורת (audit trail) להחלטות ופעולות של סוכני AI, כולל איזו מדיניות חלה ברגע ההחלטה? |
| P4 | Problem-led | AI agent observability tools show me traces of model and tool calls, but how do I prove the actual business outcome of an agent's action? | כלי observability לסוכני AI מראים לי traces של קריאות מודל וכלים, אבל איך אני מוכיח את התוצאה העסקית בפועל של פעולת הסוכן? |
| P5 | Problem-led | What are the best tools in 2026 for verifying the business outcomes of production AI agents? | מהם הכלים הטובים ביותר ב-2026 לאימות תוצאות עסקיות של סוכני AI בפרודקשן? |
| P6 | Category and alternatives-led | Are there alternatives to agent observability platforms that verify AI agent outcomes at the business level, by reading back from the systems the business trusts? | האם יש חלופות לפלטפורמות observability לסוכנים שמאמתות תוצאות של סוכני AI ברמה העסקית, על ידי קריאה חוזרת מהמערכות שהעסק סומך עליהן? |
| P7 | Brand-led | What is Pruvz? | מה זה Pruvz? |
| P8 | Brand-led | Is Pruvz (pruvz.ai) a real product? What does it do and what is its current status? | האם Pruvz (pruvz.ai) הוא מוצר אמיתי? מה הוא עושה ומה הסטטוס הנוכחי שלו? |
| Engine surface | Configuration and account state | Grounding observed | Problem-led discovery | Brand-led runs |
|---|---|---|---|---|
| Perplexity | perplexity.ai default mode, fully logged out | Search-grounded: answers built from cited sources | 0 / 28 | 8 / 8 engaged (see accuracy below) |
| ChatGPT | chatgpt.com default model, fully logged out | Parametric in most runs (little or no browsing observed) | 3 / 28 | 8 / 8 engaged |
| Gemini | gemini.google.com default model (Flash), dedicated clean account, activity saving declined | Parametric in most runs | 3 / 28 | 8 / 8 engaged |
| Claude API | claude-sonnet-5 via the Anthropic API with the web_search tool, API key with no memory; partial, budget-truncated coverage | Search-grounded | 1 / 5 | not run |
| Claude subagent | Fresh-context Claude Code subagents with web search; partial coverage with a declared priming caveat | Search-grounded | 2 / 3 | not run |
Result 1: discovery is rare, and phrasing decides it.
Across all problem-led and category-led runs, Pruvz appeared 9 times in 92 runs. Perplexity, the one fully covered engine that answers strictly from citable sources, found Pruvz 0 times in 28 discovery runs. Where Pruvz did appear, the phrasing was market-overview, tools-that-verify, or alternatives-to language; the how-to phrasings produced methodology answers with almost no vendor names at all.
| Prompt | Intent, condensed | Runs where Pruvz appeared |
|---|---|---|
| P0 | Market overview (the original failed prompt) | 5 / 14 |
| P1 | Verify a claimed action | 0 / 12 |
| P2 | Tools that verify against systems of record | 1 / 14 |
| P3 | Audit trail including decision-time policy | 0 / 12 |
| P4 | Beyond observability traces | 0 / 12 |
| P5 | Best tools in 2026 | 0 / 13 |
| P6 | Alternatives that verify at the business level | 3 / 15 |
Result 2: when engines know Pruvz, they describe it accurately.
The 24 brand-led runs on the three fully covered engines split as follows: 17 described Pruvz accurately, 4 were accurate with minor issues (invented example use cases, a partial category mislabel, and a correct founder attribution the site itself did not source at the time), 2 answered about a different company entirely, and 1 denied the product exists. No run invented product capabilities: the positioning, the evidence-packet contents, and the current design-partner status were reported correctly wherever Pruvz was actually described, in both languages.
Result 3: the Hebrew entity collision.
Both Hebrew runs of "What is Pruvz?" on the search-grounded engine answered about Provision-ISR, an unrelated Israeli CCTV manufacturer, asserting that "Pruvz" is a colloquial short form of "Provision" and citing that company's site. A Hebrew speaker asking about the brand was handed a different company twice out of two attempts. The parametric engines showed no such confusion, which localizes the defect: at the time of the runs there was no Hebrew-language source for the Pruvz entity anywhere, so a retrieval engine corrected the query toward a phonetically similar brand with Hebrew coverage. The English runs of the same prompt on the same engine described Pruvz accurately, and one English status run denied the product exists in its first attempt while its second attempt fetched the site and answered correctly, a reminder that these engines are not deterministic.
The Hebrew entity page published in responseResult 4: what the engines cite in this category.
Runs with captured citations named 64 distinct hostnames. The pattern: vendor-authored comparison and buyer-guide content dominates the citable surface for this category, YouTube was the most-cited external hostname in the search-grounded discovery runs, and Hebrew prompts drew citations from Israeli sites that no one in this category is contesting. Every hostname cited in two or more runs:
| Hostname | Runs citing it | Note |
|---|---|---|
| pruvz.ai | 9 | Cited in brand-led runs, plus two search-grounded discovery finds |
| youtube.com | 6 | The most-cited external hostname in Perplexity discovery runs |
| pedowitzgroup.com | 3 | Vendor-authored guidance content |
| cflowapps.com | 2 | Vendor-authored workflow content |
| confident-ai.com | 2 | Vendor-authored tool comparisons |
| digitalapplied.com | 2 | Vendor-authored statistics and buyer guides |
| loginradius.com | 2 | Vendor-authored engineering content |
| medium.com | 2 | Individual engineering posts |
| moticohenadv.com | 2 | Israeli legal site, cited on Hebrew prompts |
| provision-isr.com | 2 | The wrong-entity answer on Hebrew brand queries |
| skysoftconnections.com | 2 | Vendor-authored CRM data-auditing content |
| spd.co.il | 2 | Israeli site, cited on Hebrew prompts |
| stackone.com | 2 | Vendor-authored use-case content |
| usefini.com | 2 | Vendor-authored buyer guides |
Interpretation (ours, clearly labeled).
We read the data as follows; the observations above stand on their own if you read it differently. First, the results suggest the citable external surface may be a larger constraint than brand awareness: the parametric engines describe Pruvz correctly while the retrieval engine could not cite it. The baseline cannot separate the possible causes (indexing gaps, on-site content that does not match the queries, missing external authority, retrieval quality, or plain run-to-run variance); the post-release rerun will test whether the new on-site and external sources change retrieval. Second, discovery followed comparison intent in these runs: vendors surfaced when buyers asked for market overviews, tools, and alternatives, which is why we published an honest alternatives and build-vs-buy page and problem-led solution pages for outcome verification, refund verification, and system-of-record verification. Third, the Hebrew collision motivated us to publish a Hebrew-language anchor for the entity; whether that anchor is sufficient, or whether indexing, external links, and entity profiles are also needed, is exactly what the rerun will measure. And fourth, YouTube's prominence in the citation data is part of why the product demo is published as a video with a full transcript. This page itself is part of the response: it is the citable, inspectable account of the measurement.
Limitations, in full.
(1) First-party research: Pruvz measured Pruvz; treat every editorial choice accordingly. (2) The consumer claude.ai product is not covered at all; the two Claude surfaces are partial substitutes with single runs per combination, one budget-truncated and one with a priming caveat, and neither represents the claude.ai product experience. (3) Citation URLs were lost on all 5 recorded Claude API runs: a token-saving flag on the runner also dropped the citation payload, so those runs are marked in the dataset and excluded from the citation aggregates. (4) Three further Claude API calls appear in our internal cost ledger but their answers were never retained, including one rejected zero-search configuration test; they are excluded from the dataset entirely rather than counted. (5) ChatGPT and Gemini answered from parametric knowledge in most runs, so their discovery numbers measure training-data recall more than live retrieval. (6) Runs were executed from an Israeli IP and engines may localize. (7) Answer engines are nondeterministic: two runs per combination bound the variance but do not eliminate it, and single-run combinations on the Claude surfaces bound nothing. (8) Full answer texts and screenshots are retained in our internal research archive and are not republished here, because redistribution rights for third-party model output are unclear; the dataset carries structured fields and short characterizations instead. (9) The qualitative labels (appearance, accuracy, error category) were assigned by a single coder, the author, reading the full answers against the criteria below, with no independent second rating; the published data makes the aggregates recomputable from those labels, but the labels themselves cannot be independently verified without the retained answers.
How the labels were assigned.
One coder (the author) read every full answer and assigned the closed error-category vocabulary as follows. "Accurate" means the answer described Pruvz's actual positioning and status with no false claims. "Accurate with minor issues" means the description was substantially correct but included one of: example use cases the site does not claim (invented_use_cases), a partially wrong category framing (category_mislabel), or a factually correct statement the site did not support at the time (founder_named_without_source). "wrong_entity" means the answer was about a different company. "existence_denied" means the answer asserted the product is not real. Appearance (pruvz_appears) means the answer engaged with Pruvz by name at all, which is why the two wrong-entity runs count as engaged but not as accurate.
The dataset.
One CSV row per recorded run, plus an aggregates file with every number this page cites, both under CC BY 4.0. If you use it, attribute "Pruvz (pruvz.ai)".
Download the data.
116 recorded runs, July 29-30, 2026: appearance, rank, error category, citation capture state, and cited hostnames per run, with aggregates computed from exactly these rows.
| Column | Meaning |
|---|---|
| run_id | P{prompt}_{engine}_{language}_r{run}, e.g. P1_chatgpt_en_r1 |
| run_date | Local Israel calendar date of the run (YYYY-MM-DD) |
| engine | chatgpt, gemini, perplexity, claude-api, or claude-agent |
| model_version | Model label as shown by the engine, when available |
| language | en or he |
| prompt_id | P0 through P8 (full texts in the methodology section above) |
| prompt_family | discovery (P0-P6) or brand (P7-P8) |
| run_number | Repetition index within the prompt/engine/language combination |
| account_state | logged-out, clean-account, or the declared substitute-surface state |
| pruvz_appears | Whether the answer engaged with Pruvz at all (yes/no) |
| pruvz_rank | Position among named products when Pruvz appeared |
| error_category | none, wrong_entity, existence_denied, founder_named_without_source, invented_use_cases, or category_mislabel |
| citation_capture | captured, none_cited (no sources shown), or lost (capture failure, see limitations) |
| cited_source_count | Number of cited URLs, when capture succeeded |
| cited_domains | Deduplicated hostnames of the cited URLs (www-stripped; subdomains kept, so community.hubspot.com stays distinct from hubspot.com), semicolon-separated |
What happens next.
The identical frozen prompt set will be re-run after the pages this baseline motivated have been published and indexed, with the same engines, the same repetition rules, and the same declared Claude limitations, and this page will be updated with a clearly dated post-release section comparing discovery, accuracy, and citation sources against the baseline above. The prompts will not be changed to improve the result. Corrections: if you find an error in the data or believe a description here is inaccurate, write to hello@pruvz.ai.
Written and maintained by Dor Sharoni, founder of Pruvz. Product names and trademarks belong to their respective owners; Pruvz is not affiliated with or endorsed by the companies referenced on this page. Related: alternatives and build vs. buy, the recorded product demo, and a real evidence packet, explained.