A customer-facing robot is easy to judge by impression, because everyone forms one. Impressions shift with whatever was observed most recently, which makes them a poor basis for the next investment decision.
Four groups of metrics
Volume. Interactions per day, questions answered, guided visits or deliveries completed. The most basic metric, and it also reveals whether the robot is well placed.
Quality. Proportion of correct answers, proportion where it declined, handovers to staff. This group needs sampling and manual reading; it cannot be fully automated.
Availability. Actual operating hours against intended hours, stoppages and reasons. A robot sitting on charge half the day loses half its value with nobody noticing.
Impact. Staff time saved, visitors served outside staffed hours, customer feedback. The hardest to measure and the most important.
All four should be recorded from day one. Reconstructing them after three months is not possible because the underlying data no longer exists.
Collecting numbers cheaply
Take what the robot provides. Most products log conversations and interaction counts. Ask the supplier how to export this — and ask before buying, because some products do not allow it.
Read a conversation sample weekly. Twenty random conversations, marked correct or wrong. Twenty minutes, and it produces the most trustworthy quality figure available.
One column in the shift log. Staff record daily whether the robot ran the full day and whether anything went wrong. Simpler and more effective than any elaborate system.
Ask the customer one question. After an interaction, whether they were helped. Response rates are low but the trend still means something.
Ask staff monthly. What helps, what annoys. Their answers are the earliest indicator of whether the project is rising or declining.
Reducing to one usable number
To compare against alternatives, reduce everything to cost per unit of value.
Total annual cost. Purchase price divided over useful life, plus content, maintenance, power, and the owner's time valued in money.
Divide by annual interactions. That gives cost per interaction.
Compare with alternatives. Cost per interaction if staff handled it. Cost with a kiosk. Cost of doing nothing — that is, the visitor going unserved.
The result frequently surprises in both directions. Sometimes the robot is much cheaper than it feels because interaction volume is higher than assumed. Sometimes the reverse, because the robot is idle most of the time.
More important than the number itself: it points at what to do next. High cost per interaction caused by low volume is a placement or job-selection problem, not a robot problem.
The three-month review
A structured meeting with numbers on the table.
Compare against the criteria written beforehand. If none were written, that is the lesson for the next project.
Four questions. What is the robot doing better than the old way. What worse. What obstructs it most. What would be lost if it were removed.
Hear staff first. They spend eight hours a day beside it and know things the numbers do not show.
Three possible conclusions. Continue and expand. Continue but change how it is used — usually a different position or a different job. Or stop and record what was learned.
The third is legitimate. Stopping an ineffective project after three months is far better than sustaining it for two years to avoid admitting it. The learning retains its value either way.
From the first project to the second
The largest value of a first project is usually not the saving but the knowledge.
Knowing what your customers ask. The conversation log is a customer survey you did not pay for. Many businesses discover questions they had never considered.
Knowing how your team handles technology. Who engages, who avoids, and which style of communication works.
Knowing how far your estimates were off. Actual cost against budget, actual time against plan. Apply that factor to the next project.
Knowing how the supplier works. After three months it is clear whether support is genuine or was only good during the sale.
Having an internal reference. A second project is far easier to approve when a first one exists with clear numbers — even numbers that are merely adequate.
Frequently asked questions
Which metrics should be tracked?
Volume of interactions, answer quality, actual availability, and impact such as staff time saved and visitors served outside staffed hours. All four should be recorded from day one.
What is the most reliable way to measure answer quality?
Reading twenty random conversations each week and marking them correct or wrong. It takes about twenty minutes and produces a more trustworthy figure than any automated measure.
What does a high cost per interaction indicate?
Usually low interaction volume, which points at placement or job selection rather than at the robot. The fix is moving it or changing its task, not replacing the machine.
What is the biggest value of a first project?
Knowledge rather than savings — a free record of what customers actually ask, an understanding of how the team handles technology, and a measure of how far the organisation's estimates were off.
More in Service robots in hospitality and retail and Robots in education.