Skip to content

The product

One CRM. Every next step.

Contacts, pipeline and delivery in one workspace, with AI that prepares the work for your approval.

See how Meibo works
AI & AUTOMATION / THE PRACTICAL GUIDE

Measure the outcome.
Improve the work.

A fast answer is useful only when it helps complete the right job. Build a measurement system that accounts for the result, the review and the work needed to recover failures.

Get the free worksheet
m.MEASUREMENT / A DEFINED OUTCOME
THE JOBDid the automation improve the work?
  1. 01Count accepted outcomes
  2. 02Include review and repair
  3. 03Compare equivalent work
THE INTENDED OUTCOMEA decision backed by evidence
Completion · quality · time · operating costILLUSTRATIVE PROCESS / NOT A LIVE RUN
UNDERSTAND THE WORK.CONNECT THE DETAILS.MAKE THE NEXT STEP CLEAR.
In this guide
THE SHORT VERSION

Make the next decision clearer.

  • Define a logical business task separately from its runs, retries and actions.
  • Track first-pass quality and eventual completion as different measures.
  • Include review, correction and operating effort in time comparisons.
  • Keep synthetic evaluation results separate from live customer outcomes.

Define the unit of useful work

A useful CRM automation measure starts with an observable business task. That might be an enquiry routed to the correct owner, a meeting brief supported by relevant evidence or a follow-up draft ready for review. Define the accepted outcome before counting successes.

A task can require several runs or retries, and a run can propose several actions. Keep those units separate. Retrying one failed enquiry three times does not create three new customer enquiries; approving three changes within one requested handoff does not necessarily represent three completed business jobs.

Anthropic’s evaluation guidance distinguishes tasks, repeated trials, grading and the resulting environment state [1]. For a CRM, the practical check is whether the required record, draft or decision actually exists and meets the acceptance criteria, not merely whether the agent said the task was complete.

A PRACTICAL FRAMEWORK

Six measures, with explicit denominators

MeasureDefinitionKeep separate from
First-pass acceptanceAccepted initial results ÷ eligible tasksEventual completion after correction
CompletionCompleted tasks by cutoff ÷ eligible tasksRuns or tool calls
Correction rateTasks requiring correction ÷ eligible tasksRejected requests outside scope
Human effortReview + correction + administration minutesModel execution and waiting time
Time to completionElapsed time from agreed start to final outcomeTime to the first response
Cost per completionDefined operating cost ÷ completed tasksCost per model request
Define the population, observation period and evidence source before comparing these measures.

Establish a comparable baseline

Observe the current manual process on a representative set of tasks. Record active handling time, waiting time, mistakes requiring repair and the final accepted outcome. Include straightforward and awkward cases. A pilot consisting only of easy requests can misrepresent the work the team normally receives.

Define the population and exclusions. For example, include website enquiries accepted during a particular week and exclude known test submissions. Keep spam-handling performance separate if it uses a different acceptance rule. Record the timezone and observation window so someone else can reproduce the count.

Compare like with like after automation. If the new process handles only simple enquiries while the manual baseline includes complex negotiations, the apparent improvement is not a fair estimate of the change. Separate the cohorts or rerun a matched evaluation.

Track first-pass quality and eventual completion

First-pass acceptance measures tasks that meet the criteria without correction on the initial result. Eventual completion measures tasks that reach the required outcome by the agreed cutoff, including recovery. Reporting both exposes a process that ultimately works but consumes too much repair effort.

Write the denominator beside every percentage. In an illustrative set of 100 eligible tasks, 80 pass on their first result and 20 require intervention. If all 20 are later corrected successfully, first-pass acceptance is 80% and eventual completion is 100%. Those numbers answer different questions; neither should be labelled simply “AI accuracy”.

For factual quality, use specific checks: correct relationship, supported claims, no invented deadline, intended action and preserved permissions. A proposal can contain several accurate facts while failing the one criterion that makes it unsafe or unusable. Keep critical failures visible rather than averaging them away.

ILLUSTRATIVE COHORT / 100 ELIGIBLE TASKS

First-pass quality and final completion.

80Accepted first time
+
20Recovered after correction
=
100Completed by cutoff
First-pass acceptance 80%Eventual completion 100%
All 20 initial failures are successfully recovered in this fictional example. If any remain unresolved at cutoff, completion is lower.

Calculate time after review and recovery

Count human effort on both sides. In the illustrative 100-task batch, the manual baseline is eight minutes per task, or 800 minutes. The assisted process uses two minutes of review per task, another 120 minutes correcting the 20 initial failures, and 60 minutes of batch administration. Total assisted human effort is 380 minutes.

The difference is 420 minutes, or seven hours of released capacity for the same 100 completed tasks. This is not automatically seven hours of cash savings. The team needs to use that capacity productively, and the comparison needs to include the actual review and recovery work rather than just time spent generating outputs.

Measure elapsed time separately. A worker might complete research in a minute but wait a day for approval. Report time to a review-ready result and time to the accepted final outcome. If customer response is the objective, the second measure may matter more than model latency.

ILLUSTRATIVE TIME COMPARISON / SAME 100 COMPLETED TASKS

Include the work around the automation.

Manual baseline 800 min

100 tasks × 8 minutes

Assisted human effort 380 min

200 review + 120 correction + 60 administration

Released capacity 420 minutes / 7 hours
Illustrative arithmetic, not a measured customer saving. Waiting time, setup and financial assumptions are separate.

Measure cost per accepted outcome

Define the cost boundary. Model charges are one input; integration services, hosting, monitoring and human review may also belong in the comparison. Separate recurring operating cost from one-off setup effort. Record currency, period and whether a number is measured, estimated or allocated.

Continuing the fictional example, suppose the batch costs £40 in incremental software usage and completes all 100 tasks. Software usage cost per completed task is £0.40. If 80 completed tasks were the final outcome instead, the same £40 would be £0.50 per completed task. Use the observed completion count, not the number of requests sent to the model.

If you assign human time an illustrative internal rate of £30 per hour, 380 minutes represents £190 of assisted labour. Adding £40 gives £230 of operating cost against £400 for the 800-minute manual baseline. The £170 difference is a modelled capacity value before setup and other excluded costs, not a customer result or guaranteed saving.

Use repeatable tests and production evidence together

Keep a set of known-answer cases covering the jobs the automation must perform. Include the correct relationship, a misleading message, missing information and a case that should produce no action. Write the checks before observing the output so the rubric does not move to accommodate a plausible answer.

Run repeated trials where generation can vary, and report the number of tasks and attempts. A small set of synthetic cases helps detect regressions but does not establish a population-wide accuracy rate. A structured check for a source identifier also does not prove that every sentence accurately represents that source.

Meibo’s current implementation includes fictional evaluation cases for meeting briefs, opportunity discovery and follow-up preparation. These check specific evidence and action requirements. Keep those results separate from real-workspace acceptance and member feedback, where source coverage, usefulness and exceptions can differ.

Build an operational review that leads to decisions

Review unresolved tasks, oldest waiting work and the reason for corrections. Group failures by identity, missing context, unsupported claims, permissions, integration errors and unsuitable actions. This turns “the automation feels unreliable” into a set of fixable causes.

Use medians and a tail measure for latency when the sample is large enough to support them, and retain the sample size. An average can hide a few stalled jobs. Compare changes with consistent task definitions and flag when a shift in workload makes a simple before-and-after comparison misleading.

Decide what would make you expand, revise or pause a workflow. A strong result might combine dependable completion, manageable review and clear recovery. A recurring critical permission failure warrants attention even if time savings look attractive. The scorecard should support those decisions without reducing everything to one flattering number.

Use the scorecard as a living definition

The downloadable resource records each measure, its population, formula and evidence source. Adapt it to one workflow first. Record who owns each number and when it was last checked. If the completion definition changes, preserve the old definition so historical results remain interpretable.

Start with counts you can explain: eligible tasks, first-pass accepted tasks, recovered tasks, unresolved tasks, human minutes and attributable operating cost. Add more granular measures when they support an actual decision. Do not infer sales uplift from a faster draft without evidence connecting the two.

The goal is a clear account of useful work: what completed, what needed a person, how much effort it consumed and what should improve next. That account makes expansion easier to justify and helps the team trust the automation for the right reasons.

FREE RESOURCE / NO SIGNUP REQUIRED

CRM automation measurement scorecard.

Define task populations, acceptance, recovery, human effort and cost. Includes explicit formulas and the illustrative 100-task batch for reference.

Download CSVOpens in spreadsheet software. Planning worksheet, not a direct CRM import file. All example rows are illustrative.

Common questions.

What is the most useful CRM automation metric?+

Start with accepted business outcomes for a defined task population. Pair completion with first-pass quality, review effort and cost so a high completion count does not hide extensive repair work.

How is first-pass acceptance different from completion?+

First-pass acceptance counts tasks meeting the criteria without correction on the initial result. Eventual completion includes tasks recovered by the agreed cutoff. Keep the same eligible-task denominator explicit.

How should retries be counted?+

Count them as attempts or operating work attached to the original logical task. Do not inflate the business-task count because one task required several runs.

Can I call a synthetic test pass rate AI accuracy?+

Only describe what the test actually measures, with its dataset and rubric. A small fictional suite does not establish general accuracy on customer data or prove the correctness of every generated claim.

How do I measure time saved fairly?+

Compare equivalent completed work and include human review, correction and administration. Keep elapsed waiting time separate from active human effort.

Does released time equal financial savings?+

Not automatically. It can represent useful capacity. Cash savings depend on actual cost changes; modelled labour value should be labelled and should state setup and operating exclusions.

Sources & methodology.

Meibo’s recommended framework, with primary documentation for the specific product facts cited above. This is AI-assisted editorial content. Examples, diagrams and calculations are illustrative; they are not customer results or independent research findings.

  1. Anthropic — Demystifying evals for AI agents

    Primary reference for separating a task, an attempted run, evaluation checks and the resulting state. Numerical examples below are illustrative, not research findings.

Sources checked 8 October 2026. Read the editorial policy or suggest a correction.

William Mattey

Founder of Wall & Fifth, builder of Meibo and editorial contact for the learning library.

About the editorial lead
THE COMPLETE IMPLEMENTATION PLAYBOOKBring the decisions together in a practical rollout plan.