Evidence that matters beyond time saved
Two lawyers use legal AI to prepare first drafts of responses in recurring debt-recovery claims. One finishes faster than usual. The other takes longer because the assigned matter raises a limitation defence, a territorial-jurisdiction issue and objections concerning the available evidence. Averaging their minutes would create a neat number, but not reliable purchasing evidence. Time saved matters only after the firm establishes that the work being compared is genuinely comparable. Otherwise, the metric rewards easy task selection and penalises the tool for complexity it did not create.
A pilot without controlled variables mixes together at least three effects: the technology, the matter and the person performing the assignment. An uncontested claim supported by a short payment history is not equivalent to a dispute involving partial payments, an assignment of debt or uncertain proof. Nor is an associate’s hour operationally identical to a partner’s hour or a trainee’s hour. A before-and-after comparison becomes credible only when the pilot records these differences and prevents one unusually simple or difficult matter from determining the conclusion.
| Comparable assignments | Match matters by document volume, issue profile, deadline pressure and lawyer seniority before comparing performance. |
|---|---|
| Usable outputs | Measure drafts that advance without restarting, together with rework, missing issues and throughput. |
| Predefined thresholds | Set success criteria before reviewing results and apply them within each complexity tier. |
Define the comparison unit and matched baseline
Start by defining the unit of comparison. In this example, it could be one complete first draft of a response addressing a fixed set of procedural and substantive issues and ready to enter the next work stage. Define what “complete” means before anyone starts timing: the draft identifies the claim and proposed defences, addresses jurisdiction and limitation where applicable, connects material propositions to the evidence available for the exercise, and flags missing information. Without this completion standard, a bare outline and a usable draft may both be counted as finished work.
Build a matched baseline rather than comparing the pilot with a vague recollection of how long drafting usually takes. For each pilot assignment, identify prior or parallel work with a similar document volume, issue count, defence profile and deadline pressure. Exact twins will rarely exist, so record material differences rather than pretending they do not matter. A compact intake sheet can capture the number of source items, whether limitation or jurisdiction is disputed, whether the debt amount is contested, the condition of the evidential record and the seniority of the lawyer.
Stratify complexity and decompose time
Next, stratify the matters by complexity. A basic tier might cover an uncontested debt, a short factual sequence and a complete evidential set. An intermediate tier might include disputes about payment allocation, interest or quantum. A high-complexity tier might involve limitation, jurisdiction, assignment of the debt, several connected legal relationships or evidential objections. Compare baseline and pilot results within each tier before combining them. This shows whether the tool performs consistently on representative routine work while becoming less predictable in atypical matters.
Time remains useful, but it should be decomposed. Record active time to the first draft, additional rework time and the time contributed by anyone who takes over the task. Distinguish partner, associate and trainee time, because the same number of minutes has different consequences for capacity and work allocation. Record familiarisation time as well. Early users may be slower while learning the environment; silently excluding that cost would flatter the pilot, while treating it as a permanent task cost would understate performance after adoption.
Measure usable outputs and the distribution
Add the usable-output rate: the proportion of drafts that reach the next work stage without being discarded and restarted. This does not classify an output as final or automatically correct; Wisanna outputs are not automatically correct or final. It asks a narrower operational question: did the first draft provide enough structure and issue coverage to continue? Pair that rate with rework intensity, issues later found to be missing, and throughput—the number of comparable assignments reaching the defined completion standard during a fixed period.
Read the distribution, not just the average. The median describes the typical result when an exceptionally easy or difficult matter would distort the mean. The range shows how variable performance was, while separate failure cases expose the boundary of a promising use case. A favourable median for basic claims may coexist with wide variation in matters involving limitation and jurisdiction. That is actionable evidence: the firm might expand routine testing while keeping the complex tier under a narrower protocol.
Wisanna is a private and secure legal-AI workspace built for lawyers. Its public product surfaces include AI Chat, a Microsoft Word add-in and Wisanna Draft for editable legal documents. A serious pilot should test those surfaces through repeated, comparable assignments in a working legal environment. A curated demonstration can show what a product surface looks like, but it cannot establish a usable-output rate, quantify variation across complexity tiers or reveal how time effects differ by lawyer seniority.
Set the decision rule for Wisanna
Set the decision rule before seeing the results. A firm might require an improvement in median active time for basic matters, no deterioration in the usable-output rate, and no disproportionate increase in rework. It might also require a minimum sample in each complexity tier so that one strong example cannot carry the decision. The threshold need not imply an immediate firm-wide rollout: it can authorise a larger test, continuation within one use case or termination where the evidence is consistently weak.
The strongest conclusion is therefore not “the pilot saved an average of twelve minutes.” It is a segmented statement: for which debt-recovery assignments, performed by which roles and under what level of complexity did Wisanna help work reach the next stage, with what typical time effect and what variation? If the threshold is reached only for basic matters, broaden testing there. If results diverge sharply by seniority or issue profile, redesign the pilot around that finding. A balanced scorecard turns a novelty presentation into an evidence-based operating decision.
Evaluate Wisanna with representative legal work
Test AI Chat, the Microsoft Word add-in and Wisanna Draft through repeated assignments, matched baselines and predefined decision thresholds.
See Wisanna's lawyer-controlled workflow