MintMarbles Start a project

Practical AI: measuring automation in hours saved, not demos shipped

AI · July 2026 · 3 min read

It has never been easier to build an impressive AI demo. It is still hard to build automation that saves a team real time every week. The difference is almost always in how success is measured.

Demos are cheap, hours are not

A chatbot that answers a few questions well can be built in an afternoon. It looks great in a meeting. Then it meets real customers, real documents and real edge cases, and the team quietly goes back to doing the work by hand.

We have stopped measuring AI projects by what they can do in a demo. We measure them by the hours they give back and the errors they remove. If we cannot estimate those numbers before we start, we do not start.

The question is never "can AI do this?". It is "how many hours a week will this give back, and how will we know?".

Start with the work, not the model

Every automation project begins with a simple exercise: list the repetitive tasks a team does, how often they happen and how long each one takes. Then we look for the ones that are frequent, rule-based enough to check and painful enough that people will welcome help.

Good candidates usually look like this:

  • Answering the same customer questions about products, pricing or policies.
  • Reading documents such as invoices, forms or applications and typing their contents into another system.
  • Sorting and routing incoming enquiries to the right person.
  • Drafting first versions of routine replies, summaries or reports for a person to approve.

Keep the scope narrow

The most useful assistants we have built do one job well. BELSense, the assistant we built for Beacon Energy, only answers questions about solar panels, inverters, batteries, net metering and the company's own services. It politely declines everything else. That narrow scope makes it more accurate, easier to test and far easier for the team to trust.

A general assistant that can talk about anything is impressive for a day. A focused one that answers the fifty most common customer questions correctly, at any hour, is useful for years.

Measure before and after

Before we build, we record a baseline: how many of these tasks happen each week, how long they take and how often they go wrong. After launch, we measure the same things. We report three numbers:

  1. Hours saved per week, based on tasks the automation completed or shortened.
  2. Error rate, compared with the manual process, including mistakes the automation made.
  3. Hand-off rate, meaning how often a person still had to step in.

A falling hand-off rate is the clearest sign that an automation is earning trust. A rising one is an early warning that the scope or the data needs attention.

Design the hand-off

No automation is right every time, so the moment it hands over to a person matters as much as the moment it succeeds. Good hand-offs include the full context, explain why the automation stopped and never leave a customer waiting without a reply. When people can see and correct what the system did, they are far more willing to let it do more.

A simple checklist

  • Can you name the task, how often it happens and how long it takes today?
  • Is there a way to check whether the output is right?
  • Is the scope narrow enough to test thoroughly?
  • Is there a clear path to a person when the automation is unsure?
  • Will you measure hours saved and errors after launch, not just usage?

If the answer to all five is yes, the project is worth building. If not, it is probably a demo.

Start a project

Have an idea?
Let’s build it together.

Tell us what you are building. A senior engineer replies within one business day with next steps, not a sales deck.