The Klade Ratio — Methodology v0.1 (draft)
This document is written so that a skeptical CFO, auditor, or economist can attack it and find the limitations already stated. That is the point.
1. Definitions
- Workflow: a repeatable unit of operational work whose output is recorded in a system of record (e.g., claims adjudication, diligence memo drafting, LP reporting, invoice reconciliation). The workflow is the only unit of analysis. Klade does not score prompts, sessions, or individuals.
- Counted output: a discrete output unit logged by a system of record independent of Klade and independent of the AI tooling (claims processed, memos filed, deals screened, tickets closed, billable hours by task code).
- Fully-loaded AI cost (denominator): metered inference spend (tokens × rate, per model tier) + seat licenses attributable to the workflow + directly attributable AI tooling. Sourced from provider admin APIs (Anthropic Enterprise Analytics, OpenAI per-user consumption) and invoices. Metadata only — Klade never accesses prompt or completion content.
2. Qualification test
A workflow is in scope if and only if its output is counted in a system of record. If not, Klade classifies the workflow as unvalued and reports its spend without a ratio. Klade will decline engagements where no workflows qualify. Software engineering workflows are excluded categorically (measured adequately by an existing owned category).
3. The three tiers
Reported separately. Never blended into a single composite score.
Tier 1 — Hard waste (measured)
Observed, non-inferential dollars:
- Tier mismatch: frontier-model tokens spent on tasks with a defined cheaper-tier policy equivalent (classification, extraction, routine drafting), priced at the rate delta
- Redundant inference: near-duplicate calls, retry storms, recurring scheduled jobs producing unconsumed output
- Runaway agents: agentic loops exceeding step/cost budgets without terminal output
- Idle capacity: paid seats with zero or de-minimis usage over the observation window
Tier 1 findings are arithmetic on observed telemetry. No attribution model is involved. They lead every report.
Tier 2 — Efficiency vs. benchmark (compared)
Cost per counted output unit, compared against (a) the customer’s own pre-AI or trailing baseline, and/or (b) the cross-company benchmark for the same workflow class. Because both sides of the comparison are ratios, no dollar value is ever assigned to the output itself. Reported as a percentile / multiple vs. baseline with the comparison set size disclosed.
Tier 3 — Value attribution (only where priced)
Only where the system of record prices the output (billable hours by task code, revenue-bearing claims, fee-bearing filings):
Klade Ratio (workflow) = attributable priced output over window ÷ fully-loaded AI cost over window
Attribution rule: priced output is attributed via (in preference order): (1) matched non-AI comparison group within the same org; (2) the customer’s own pre-adoption baseline, trend-adjusted; (3) benchmark counterfactual, flagged as lowest confidence. Self-reported time savings are inadmissible.
Ranges, not points. Every Tier 3 ratio is reported as an interval (e.g., 3.2–4.6×) whose bounds are the attribution methods above applied conservatively vs. neutrally, with all assumptions enumerated in an appendix. A single-point score is not offered.
4. Data sources & handling
- Provider admin APIs via customer-created, read-only, scoped keys (usage/cost endpoints only)
- Output exports from customer systems of record (counts, timestamps, task codes — content redacted at source where applicable)
- Structured intake interviews (week 1): used to map workflows to spend streams and output systems, never as measurement input
- Engagement runs inside Klade’s SOC 2-scoped environment; customer data destroyed at engagement close per contract; no client-side capture of any kind
5. What the observation window must satisfy
Minimum 60 days of telemetry overlap between spend data and output data; workflows below a minimum output count (n < 50 units in window) are reported Tier 1/Tier 2 only, flagged low-n.
6. Stated limitations (standing section)
- Output counts measure throughput, not quality. Klade does not measure quality and does not claim to.
- Attribution in Tier 3 is a counterfactual estimate; the range brackets it, it does not eliminate it.
- Copilot/Microsoft estates expose aggregate-only telemetry by vendor policy; such estates receive Tier 1 (partial) and Tier 2 treatment only, disclosed as such.
- The benchmark’s power grows with rows; early benchmarks disclose n and should be weighted accordingly.
- Uncounted knowledge work may be the most valuable AI use in an organization. Klade reports its spend and refuses to guess its value. Buyers who need that number guessed should hire someone else.
7. Individual-level measurement (policy)
Workflow/team level is the default and the product. Per-employee detail exists only as a separately contracted SKU inside enterprise engagements, contractually barred from compensation or termination use. No consumer self-upload scoring is offered: self-supplied session exports are unverifiable and gameable by the measured tool itself. A future verified-credential tier, issued only from enterprise telemetry during engagements, is a roadmap item, not a product.
8. Versioning
The methodology is versioned; every audit report cites the version used. Changes are logged in CHANGELOG order here. Public release target: after first design-partner audits validate the mechanics.