“People tried it” is an encouraging update at the beginning of an AI rollout. Some people found the tool, got access, and were curious enough to open it. If that same update appears in the next three leadership meetings, though, it is time to ask what happened after they logged in.

Did they complete useful work? Did they return? Did they quietly go back to the old process after an impressive first attempt? A usage dashboard can look busy without answering any of those questions. You need to follow the work far enough to see whether it changed.

Start with one recurring task. Define what you want to improve, then measure its use and its results separately. A process can be popular without helping much, and a useful process can remain hard for people to find or fit into their day.

Follow one task through five stages

Suppose a sales team introduces AI assistance for account preparation before customer meetings. The job is to gather relevant account information, check it, correct any errors, and prepare for a specific conversation. Generating a summary is one step along the way.

Five measures help explain what happens:

  1. Access: How many intended users can use the approved tool? Count the people who need it, and find out who still lacks access or setup help.
  2. Trial: How many completed at least one relevant task? Opening the application and finishing meeting preparation are different events.
  3. Repeat use: How many returned for another task during the period you are measuring?
  4. Workflow adoption: How often did people use the agreed process, including review, when a suitable task came up?
  5. Outcome: Did the work become faster, more useful, or more reliable, including the effort to check and correct it?

Account for how often someone has a reason to use the process. A person preparing for one major meeting each month should not look like a failed adopter because someone else has five meetings a day. Measure the opportunities to use it as well as the people who do.

Look at what happens to the old process, too. If the AI summary becomes the starting point for checked preparation, the work has changed. If people generate it and then do their usual research from scratch, they may be doing both versions. The activity count will not show that extra work unless you ask.

Know what you are comparing

Before introducing the new process, examine a small, representative sample of meeting preparation. Record the time spent researching, checking information, and writing the final notes. Include corrections and follow-up work caused by gaps in the preparation.

Choose a few quality criteria you can apply to both methods. Are important claims supported by the account material? Are the customer's stated priorities represented accurately? Are missing facts identified? Does the preparation help the seller choose useful questions for this particular meeting? A tidy document can still contain the wrong account history, so judge it by what the seller needs to do with it.

Use the same standards for both methods. Where practical, have reviewers assess samples without knowing which method produced them. Compare similar tasks, too: a routine check-in with an existing account and a first meeting with a complex new buyer require different preparation. Keep those differences visible, along with changes in employee experience, source material, and workload. Otherwise, the average may tell you more about the mix of meetings than the effect of the tool.

Include the checking and repair

Suppose the old process takes 30 minutes for a preparation task. The AI-assisted version takes eight minutes to produce a draft, twelve minutes to verify it, and five minutes to repair it. You are comparing 30 minutes with 25 minutes. Reporting an eight-minute result makes the reviewer's work disappear.

Five minutes saved could still be worthwhile. To decide, look at the quality of the preparation and the ongoing work of maintaining the process. Faster preparation that brings unsupported claims into customer conversations is a poor trade. Preparation that takes the same time but produces more useful questions may be worth keeping. Decide which improvements matter before the results arrive.

Business outcomes need care as well. Conversion rates and sales cycles can change with pricing, staffing, the customers in the pipeline, and other parts of the sales process. Track them, but do not attribute the whole change to AI just because the dates line up. First establish what changed in preparation, then look for evidence of how that affected the conversations and results that followed.

Use the numbers to make a decision

A short weekly review can cover how many suitable tasks came up, how many used the assisted process, whether people returned, and whether review was completed. Add total effort and the quality issues found in a sample. Keep the definitions and reporting period beside the numbers so everyone knows what they are looking at.

Ask why people chose the old process when the new one was available. They may lack usable source material, have a task the tool handles poorly, or find checking the output slower than starting again. They may simply struggle to find the tool when they need it. Those answers point to different fixes. A reminder to use more AI does little to resolve them.

Set a date to decide whether the pilot should continue, change, or stop. Use the measures you agreed on to make that decision: where is the work better, what still needs repair, and which uses add more effort than they save? The login count can tell you who arrived. The rest of the review should tell you whether the process earned a place in their working day.