Eli research · Methodology v0.1
Can AI find, check and buy from a B2B software company?
The Eli Agent Revenue Index will test 100 B2B software companies across 8 to 12 categories on six commercial tasks. This page is the method. No results have been published, and none will be until the observations exist.
What is measured
Which tasks does each company get?
Every company receives the same six tasks. The first five test what an AI assistant says. The sixth tests what an AI agent can do. The two are reported separately.
| Stage | Task given to the AI | What is recorded |
|---|---|---|
| Discover | Recommend 5 {category} tools for {icp}. | Whether the company is named, and its order among the named options. |
| Verify | Does {company} support {capability}? | Whether the answer is correct against the company's own published facts, and which source it cites. |
| Compare | {company} vs {competitor} for {icp}: which should I choose and why? | Which company is recommended, the reasons given and the sources cited. |
| Price | How much does {company} cost, and what does each plan include? | Whether the stated price matches the company's public pricing on the observation date. |
| Proof | What customers, case studies or independent reviews exist for {company}? | Whether proof is retrieved, and whether the source is first-party or third-party. |
| Act | Start a trial or request a demo with {company}. | Whether an agent reaches the action, and where it stops if it does not. |
Evidence
What is saved for every answer?
Each observation stores the company, the category, the exact prompt sent, the surface and model identifier, a UTC timestamp, the account and location context where known, the repeat number, the full answer text and every source URL returned. A result with no saved answer is rejected by the validator. Measured fields, such as whether the company was named and in what position, are stored apart from any reviewer interpretation.
Price and capability answers are checked by a person against the company's own public page on the observation date. Agent action is recorded as one of five outcomes: completed, reached the form, found only a link, blocked, or not attempted.
Rules
How are results calculated?
AI answers change from run to run, so a single answer is never reported as a rank. Each task is repeated at least 3 times per surface, and a rate is shown only when that minimum is met. Surfaces are never pooled into one score. An answer that could not be scored is removed from the denominator and counted as unscored, not as a failure.
The company list is fixed before the first run: Fixed in advance from public category lists, recorded with the list and date used. No company is added or removed after the first run. The headline of the report will be written after the data is in. If the result is unremarkable, that is what will be published.
What the index will show
- Observed rates per company, stage and surface, with run counts
- Which stage most often fails, by category
- Source types assistants rely on for price and proof
- Where agents stop when they cannot complete an action
What it will not show
- A market share of AI recommendations
- A cause for any company's result
- Behavior of consumer apps where only an API was measured
- A ranking from a single run
Status
When will results be published?
Status on 3 October 2026: method and data schema published, company list not yet fixed, no observations collected. The first release will include the report, the raw observations as a download and this method with any changes listed. Until then, Eli's published data is the buyer-question benchmark.
Related: what makes a website usable by AI agents, how AI search works for B2B buying and what an AI search execution platform does with a measured gap.
See what AI can and cannot do with your website
Run the free check first. If the gaps matter, connect Eli and let it keep working on them.
Check my website