A growing number of companies now track how often employees use AI tools, and some have started building that number into performance reviews. HR Dive covered the trend this summer, and the reaction among leaders I talk to splits right down the middle. Half think it is overdue. The other half think it is surveillance with a dashboard.
I have spent the last few years putting AI into real workflows inside real businesses, recruiting, field service, operations, back office. My view is that tracking usage is not wrong. Rewarding it is.
Before I argue against it, the other side deserves a fair hearing, because the people pushing usage metrics are not foolish.
Companies are spending real money on AI licenses, and leaders have a duty to know whether anyone is using them. Adoption inside most organizations is uneven. A few enthusiasts run ahead, and a lot of people never open the tool. If you never measure it, you never see that gap.
There is also a culture argument. Saying "we expect you to learn this" means more when it shows up in how people are evaluated. Gartner's Kate Jensen told HR Dive that including AI use in reviews can be a great first step toward building it into broader workforce conversations. I agree with that framing. It is a first step.
Here is the problem. The moment usage becomes a scored metric, people optimize for usage.
I have watched this happen with every activity metric I have ever seen tied to evaluations. Calls made. Emails sent. Tickets touched. Hours logged. People are smart, and they will give you exactly the number you ask for. They will not necessarily give you the outcome you wanted.
With AI, the cost of that is higher than it looks. If I am graded on how often I use the tool, I will route more work through it, including work it does poorly, and I will spend less time checking what comes back. The same HR Dive piece quotes Jensen saying usage alone "doesn't encourage the right behaviors," and others warn about a flood of low-quality AI output that looks like productivity. That matches what I have seen. Volume goes up. Judgment does not.
There is a quieter cost too. Your best people are often the ones most willing to say, "I tried it on this, and it is not good enough yet." A usage score punishes them for the exact discernment you need.
If you want AI to change how work gets done, measure the work.
Pick specific workflows, not general usage. Name three or four processes where AI should help, such as first drafts of proposals, candidate screening notes, or service call summaries. Track cycle time, error rate, and rework on those processes before and after.
Assign an owner for each workflow. Someone needs to be accountable for whether the AI version is actually better, and empowered to change it or stop it.
Ask people to show their work. In one-on-ones, have team members walk through how they used AI on a real task, what they kept, and what they threw out. You will learn more from ten minutes of that than from a month of login counts.
Reward good calls in both directions. Recognize the person who found a smart new use and the person who correctly said the tool was not ready for a task. Both are exercising the judgment you are paying for.
Keep the skills that AI is supposed to support. If a role requires someone to write clearly, analyze a deal, or diagnose a problem, keep evaluating those skills directly. Tools change. Judgment is the asset.
Across my companies, the difference between AI that sticks and AI that stalls has rarely been the tool. It has been whether a leader got close enough to the work to see what was actually happening. Usage dashboards feel like control. They give you something to show the board. But they let a leader stay at a distance from the very thing they are trying to change.
So track usage if it helps you spot who needs training. Just do not confuse the number with the result. The question that matters in a review is not "how often did you use AI?" It is "what got better, and how do you know?"
Connect with Steve: linkedin.com/in/stevepurban