AI ROI

Useful Intelligence per Dollar: A Scorecard for AI Investment Returns

OpenAI proposes a framework measuring AI ROI through useful work accomplished, cost per successful task, dependability, and scaling economics.

Useful Intelligence per Dollar: A Scorecard for AI Investment Returns — article cover
On this page10 SECTIONS
  1. The Shift from Adoption to Work Accomplished
  2. The Four Measures of Useful Intelligence per Dollar
  3. Measure 1: Useful Work Done
  4. Measure 2: Cost per Successful Task
  5. Measure 3: Dependability
  6. Measure 4: Value at Scale
  7. Practical Implications for Product Builders
  8. Limitations and Trade-offs
  9. Conclusion
  10. Sources

The Shift from Adoption to Work Accomplished

For years, enterprise software success was measured by adoption metrics: seats purchased, active users, and license renewals. OpenAI argues that in the AI era, these numbers miss the point. The real measure should be work accomplished—the actual tasks AI completes that create value.

This article breaks down OpenAI’s proposed framework, “Useful Intelligence per Dollar,” as outlined in their 2026 A scorecard for the AI age. We examine the four key questions it answers and what product builders and AI learners can take away.

The Four Measures of Useful Intelligence per Dollar

OpenAI defines the ultimate scorecard as answering four questions:

  1. Is AI completing work that matters?
  2. What does each successful task cost?
  3. Can people depend on the result?
  4. Does each AI dollar produce more value as usage grows?

Each measure targets a different aspect of AI return on investment, shifting focus from token counts to business outcomes.

Measure 1: Useful Work Done

Start with a specific workflow. Define what “done” means and track outcomes in the system where work happens. Examples from OpenAI:

  • Support team: a customer issue resolved.
  • Engineering team: a code change passing tests.
  • Legal team: a contract reviewed accurately and on time.
  • Finance team: for forecast review, tasks like finding data, moving it into spreadsheets, identifying changes, reconciling tabs, rebuilding slides—all work that AI can automate, freeing people for judgment calls.

OpenAI cites ChatGPT Work as an example of taking on such processes, allowing teams to focus on higher-value questions like “What changed? Why? What should we do next?”

Measure 2: Cost per Successful Task

Raw cost per token is misleading. A cheaper model may require more attempts, more time, or more human review. A more capable model might complete the task in one pass. The true metric is the full cost of producing a successful outcome.

OpenAI provides a simple formula:

  • Total full cost (including employee time, human review, retries, rework)
  • Count tasks that met the quality bar
  • Divide total cost by successful tasks

As an illustration, OpenAI’s GPT‑5.6 family (released in July 2026) offers three tiers: Luna (fast, affordable), Terra (balanced), Sol (flagship). On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning set a new state of the art while using 54% fewer output tokens than another leading model. On DeepSWE v1.1, Sol achieved 72.7% pass rate above Claude Fable 5’s 69.9%, at 36.2% lower estimated API cost (as reported by OpenAI).

These numbers highlight that a more capable model can lower cost per successful task even if its per-token price is higher.

Measure 3: Dependability

Dependability directly affects economics. When results are accurate and consistent, people spend less time reviewing and correcting. OpenAI suggests tracking three outcomes:

  • Ready to use: result met quality bar as delivered.
  • Needs correction: required another attempt or human edits.
  • Needs escalation: a person had to step in and finish the work.

These metrics reflect real-world reliability better than model accuracy scores. Dependability also demands clear boundaries: organizations must define what data the system can access, what systems it can modify, and when human approval is needed.

OpenAI notes that capability earns first use, but dependability makes AI part of how work gets done. ChatGPT Work builds on ChatGPT Enterprise’s security and compliance foundation to enable deeper integration while maintaining oversight.

Measure 4: Value at Scale

The final question is whether economics improve as usage grows. Companies should track the same workflow over time: tasks meeting quality bar, total cost, and cost per successful task. If completed work grows faster than total cost while quality holds, each AI dollar is producing more value.

Compute sits at the center. Better models, more efficient inference, purpose-built hardware, and smarter routing all improve return on compute. OpenAI emphasizes that improvements compound: better infrastructure enables research, which produces more capable and efficient models, which improve products and drive adoption, funding the next cycle.

Practical Implications for Product Builders

From OpenAI’s framework, product teams can take several concrete actions:

  1. Define one core workflow that delivers clear business value (e.g., customer reply summarization, code review, report generation).
  2. Establish a success definition with stakeholders: what output counts as “done and usable”?
  3. Compute full task cost including human time, review, retries—not just API spend.
  4. Track dependability distribution (ready vs. needs correction vs. escalation) and work to reduce rework.
  5. Monitor scaling economics monthly or quarterly: is cost per successful task declining as volume grows?

For those learning AI tools, the framework reframes evaluation: ignore token prices alone and instead benchmark total effort needed to get acceptable results.

Limitations and Trade-offs

OpenAI’s framework comes from a vendor perspective. It naturally favors capable models like GPT‑5.6. The benchmarks cited (e.g., Artificial Analysis Coding Agent Index, DeepSWE v1.1) are specific to coding tasks; results may not generalize. CFOs and product builders should cross-validate with their own tasks, quality thresholds, and cost data.

Additionally, the “useful work” definition requires careful scoping. Not all valuable AI output is easily countable (e.g., creative inspiration, incidental learning). The framework works best for repeatable, measurable workflows.

Dependability tracking assumes the organization has systems to capture ready/correct/escalate status—not always trivial to implement.

Conclusion

OpenAI’s “Useful Intelligence per Dollar” shifts AI investment measurement from vague adoption stats to concrete work outputs and cost efficiency. For product builders, it serves both as an evaluation tool and a design guide: choose models, design workflows, and set quality standards with the goal of producing more reliable work at lower cost. The real value of AI lies not in tokens consumed but in meaningful work completed.

All claims and figures in this article are sourced from OpenAI’s original scorecard document (2026) unless otherwise noted.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL