10 min

How an AI System Found and Scored 42 Automation Opportunities

Automation Discovery Dataset Scoring Validation OpenClaw

On 2 May 2026, an early automation-discovery workflow produced 42 opportunity briefs: seven candidates in each of six domains. The number is real. The original description—“production-ready”—was not. Every item in the surviving registry is still marked identified, not validated, designed, built, or live.

This article publishes the dataset and shows what its scores can and cannot tell us.

The dataset at a glance

DomainItemsAverage ICETop itemTop ICE
German healthcare745.9AM-00475
Real estate and smart buildings750.4AM-00875
Carrier bidding and e-commerce748.9AM-01580
Precision irrigation754.9AM-02280
Greenhouse climate757.1AM-02980
Wine-cellar inventory764.6AM-036100

Across all 42, average ICE was 53.6. The portfolio contained 11 document, nine decision, eight monitoring, seven communication, and seven knowledge opportunities.

How the early workflow generated candidates

Each domain run deliberately produced one small portfolio rather than one “best” answer. The workflow examined five process layers—document, communication, decision, monitoring, and knowledge—then created seven self-contained briefs with global AM identifiers.

Each brief received Impact, Confidence, and Ease factors from one to five:

ICE = Impact × Confidence × Ease
range = 1–125

The arithmetic is deterministic. The factors were model judgments. That distinction prevents a score such as 80 from masquerading as measured ROI.

The audit also found one integrity defect in the historical registry: AM-002 records factors 5 × 3 × 1 but an ICE total of 30; the product is 15. The downloadable CSV preserves the recorded value instead of silently rewriting history. The other 41 rows are arithmetically consistent. This defect is one reason the open-source successor moved score validation into code.

What the scores reveal

The distribution favored bounded, legible workflows. AM-036—unified barrel identity and a digital cellar map—scored 100 because the proposed impact, evidence confidence, and implementation ease were all high. More ambitious integrations scored lower when ease or confidence fell.

The layer distribution also exposes a sampling effect: the workflow was instructed to cover five predefined layers. It therefore cannot prove that organizations naturally contain those proportions.

What the scores do not establish

  • Demand: no buyer interviews or willingness-to-pay evidence are present.
  • Feasibility: source-system access, data quality, security, and integration cost remain untested.
  • Safety: healthcare, compliance, and physical-control candidates need domain-specific assurance.
  • Novelty: a distinct title does not prove the underlying workflow is non-duplicative.
  • ROI: ICE ranks hypotheses; it does not calculate cash flow, adoption, or operating cost.

Selection bias in the experiment

The six domains were chosen, not randomly sampled. Every run was forced to return seven items. The same prompt structure and scoring rubric shaped all outputs. The corpus therefore supports comparison within this experiment, not a claim about the prevalence of automation across the economy.

The “one day” duration measures generation throughput. It does not include domain-expert validation, solution design, procurement, implementation, or change management.

A validation funnel for any selected item

  1. Interview the workflow owner and at least two operators.
  2. Observe real cases; measure volume, cycle time, rework, exceptions, and failure cost.
  3. Map systems, data rights, decision authority, and human-oversight requirements.
  4. Re-score Confidence and Ease using evidence, never the original prose.
  5. Prototype the narrowest reversible step and compare it with the existing baseline.
  6. Advance the status only when the evidence for that lifecycle transition is retained.

Download the complete 42-item inventory

The CSV contains every original ID, title, domain, layer, factor, ICE score, and lifecycle status. It is a historical dataset, not an endorsement of the ideas.

Download the 42-item dataset