Independent software research / Investor-owned utilities, municipal and public power, electric cooperatives, competitive energy retailers, grid and generation operators

The yardstick for AI in energy and utilities - calibrated for your utility type and your metering and customer systems.

We test every B2B AI vendor on the same rubric, and we score for what utility teams actually need: better customer engagement and billing, grid and distributed-energy-resource (DER) orchestration, load forecasting and meter-data analytics, asset performance and grid reliability, and the compliance posture that decides whether a vendor is even buyable - NERC CIP, SOC 2, and a clean separation between your customer systems and your grid-control (operational-technology) systems. Whether you run an investor-owned utility, a municipal or public power utility, an electric cooperative, a competitive energy retailer, or a grid or generation operation, the audit routes by your utility type and your existing metering and customer systems, and it treats the grid-safety boundary as a hard filter rather than a footnote. Utility COOs, CIOs, and VPs of Operations choose on evidence, not vendor demos.

Take the free 4-minute energy and utilities readiness audit

Score, gaps, and three utility-fit tool recommendations benchmarked against investor-owned, municipal, cooperative, retailer, and grid-operator peers. No email required to see your score.

How we test, score, and publish

Yardstick Research is an independent software research and consulting agency for B2B AI tools. We test the tools ourselves, score them on outcomes that matter, and publish the results. Methodology in plain sight, so any board, commission, or oversight committee can check our work. For investor-owned utilities, municipal and public power, electric cooperatives, energy retailers, and grid and generation operators, we weight Data Readiness and Team & Workflow heavily because in utilities the integration spine (the customer-information system, the meter-data layer, and the boundary between customer systems and grid-control systems) decides what AI you can actually deploy, and the workforce that owns the workflow - much of it union-represented and safety-critical - decides whether it sticks. Here's how that actually happens:

  1. 01

    We evaluate every energy and utilities vendor on this list using public information and free-tier hands-on.

    Our researchers evaluate each vendor on the list using a defensible mix of inputs: vendor documentation and pricing pages, free-tier or trial-seat hands-on where the vendor offers one, video walkthroughs, third-party reviews (G2, Capterra, Gartner Peer Insights), published utility case studies, practitioner discussion (Utility Dive, T&D World, GTM / Wood Mackenzie, industry conferences), and recent funding and news coverage. Where we can sign up and exercise the product directly, we do, and grade the output against a sample workflow: in the utility case, a billing-and-usage inquiry through a tool like Kraken Technologies or SEW.AI, a demand-response or DER-dispatch cycle through EnergyHub or Voltus, a meter-data disaggregation pass through Bidgely or Arcadia, or an asset-performance query through Power Factors. We do not pay for paid tiers and we do not run a held-out benchmark through every tool. Both are cost-prohibitive at the scale this guide covers.

    Every claim in a tear-sheet is labelled MEASURED (free-tier hands-on observation, or output graded against a sample workflow), ESTIMATED (cost-per-seat efficiency derived from the vendor's pricing page and feature limits), or CITED (vendor-published or third-party benchmark, with the source linked).

  2. 02

    We score on outcomes buyers care about, with weights we publish.

    Vendor decks sell features. Utility teams actually buy outcomes: customers who self-serve instead of calling, bills that are right the first time, distributed energy resources orchestrated without a reliability event, load forecasts a rate case can defend, generation and grid assets that fail predictably rather than unexpectedly, and a stack that clears the security and reliability review the first time. We score seven dimensions: Asset performance and grid reliability; Cost economics and time-to-value; Utility customer engagement and billing; Ease of data integration and accuracy; Energy data analytics and forecasting; Grid and DER orchestration; and Regulatory, security, and grid compliance. The dimensions and benchmarks are public so your board or commission can defend the pick, and so vendors can't quietly negotiate them. The audit also captures your operating baselines (utility type, meters or accounts served, metering and customer systems, and compliance gates) and fans return-on-investment scenarios out per baseline.

  3. 03

    We publish. Vendors check facts. Affiliate links are disclosed.

    Every vendor receives their scored tear-sheet seven days before publication and can flag factual errors (wrong pricing tier, misquoted feature, a certification listed that the vendor does not actually hold, integration listed as native that's actually via a third party). Rankings can't be appealed; only factual corrections are accepted. Where the guide links to a vendor's product, that link may earn us a commission. Disclosed on every page where the link appears. Vendors do not pay for inclusion, placement, or ranking.

Take the audit

See your score

Get your results

Free. Calibrated for energy and utilities

AI Readiness Audit. Energy & Utilities edition

Select Your Industry