Independent software research / Operators spanning multiple industries
The yardstick for AI: pick the closest industry, get a calibrated audit.
We test every B2B AI vendor on the same rubric, but the scoring weights and the benchmarks change by industry. A retail operator and a pharma operator should not be graded against the same baseline. If your business spans multiple industries, or doesn't fit any one of them cleanly, pick the closest match below. You get that industry's full audit: its sub-segment routing, the compliance frameworks procurement actually checks, its vendor pool, and its peer benchmarks. Senior decision-makers choose on evidence, not vendor demos.
Take the free 4-minute AI readiness auditHow we test, score, and publish
Yardstick Research is an independent software research and consulting agency for B2B AI tools. We test the tools ourselves, score them on outcomes that matter, and publish the results. Methodology in plain sight, so any board or executive committee can check our work. The audit covers sixteen industries today. Once you pick the closest match, the rubric stays the same but the scoring is calibrated to that industry's reality: a manufacturer's Data Readiness benchmark sits at 40 percent of the maximum, a FinTech operator's at 75 percent, a pharma operator's at 70 percent. Here's how that actually happens:
-
01
We evaluate every vendor on each list using public information and free-tier hands-on.
Our researchers evaluate each vendor using a defensible mix of inputs: vendor documentation and pricing pages, free-tier or trial-seat hands-on where the vendor offers one, video walkthroughs, third-party reviews (G2, Capterra, Gartner Peer Insights), published customer case studies, practitioner discussion, and recent funding and news coverage. Where we can sign up and exercise the product directly, we do, and grade the output against a sample workflow drawn from the industry you select. We do not pay for paid tiers, and we do not run a held-out workload through every tool. Both are cost-prohibitive at the scale these guides cover.
Every claim in a tear-sheet is labelled MEASURED (free-tier hands-on observation, or output graded against a sample workflow), ESTIMATED (cost-per-seat efficiency derived from the vendor's pricing page and feature limits), or CITED (vendor-published or third-party benchmark, with the source linked).
-
02
We score on outcomes buyers care about, with weights we publish.
Vendor decks sell features. Operators actually buy outcomes. Every audit, regardless of industry, scores five dimensions: Strategy & Use Cases, Data Readiness, Tool Stack, Team & Workflow, and Budget & Procurement. What changes by industry is the benchmark a mature operator should hit. Mature SaaS operators sit near 75 / 70 / 70 / 65 / 75 percent; FinTech near 70 / 75 / 60 / 70 / 80; healthcare and pharma near 65 / 70 / 50 / 60 / 70; manufacturing near 50 / 40 / 40 / 50 / 55; non-profits near 50 / 45 / 40 / 50 / 45. That spread is the whole point: a single cross-industry score would tell a manufacturer they're behind and a pharma operator they're ahead, when both might be exactly where their peers are. The dimensions and benchmarks are public so your board can defend the pick, and so vendors can't quietly negotiate them.
-
03
We publish. Vendors check facts. Affiliate links are disclosed.
Every vendor receives their scored tear-sheet seven days before publication and can flag factual errors (wrong pricing tier, misquoted feature, integration listed as native that's actually via a third party). Rankings can't be appealed; only factual corrections are accepted. Where the guide links to a vendor's product, that link may earn us a commission. Disclosed on every page where the link appears. Vendors do not pay for inclusion, placement, or ranking.
Take the audit
See your score
Get your results
Free. Pick the closest industry
AI Readiness Audit. Mixed Industry edition
Select Your Industry