“Can we speed it up?” is a perfectly fair stakeholder question. “Can we skip user testing?” is where I start reaching for a cup of tea. In complex enterprise products, a fast untested decision is often just expensive rework arriving early.
There is a better option: use AI-assisted tools to lower the cost of making a realistic prototype, then test the important assumptions before production development begins. In my experience, the value is not AI for its own sake. It is a tighter learning loop.
Why this mattered at ShipServ
In maritime procurement, users work through complex, high-stakes flows: requisitions, approvals, vendor bidding, compliance checks and invoices. People may be working aboard a vessel with intermittent connectivity or managing a fleet from an office. A confusing journey can mean a missed approval, an incorrect part or an invoice problem—not merely an annoyed click.
When we explored automated data extraction and matching for e-invoicing at ShipServ, the solution touched existing monetised functionality. We needed to see how people understood the possible approaches before committing a team to build one. I describe the broader context in my ShipServ UX strategy case study.
Use AI prototyping to make the question testable
Tools such as Lovable and Cursor can help a team turn a concept into an app-like prototype quickly. That can be especially useful when a static wireframe is not enough to test the interaction: conditional fields, calculations, navigation logic or state changes need to feel real enough for a participant to react naturally.
The workflow I recommend is simple:
Write the hypothesis and the task a participant should be able to complete.
Build only enough interaction to test that hypothesis.
Review the prototype for misleading polish, inaccessible shortcuts and invented data before showing it to anyone.
Test with representative users, observe the behaviour, then revise.
The prototype is not production code in fancy dress. Treat it as a research instrument. If it teaches you that your idea is wrong, it has done an excellent job.
Where Maze fits
Maze can test prototypes and live websites, giving teams task-level evidence such as completion, paths, misclicks, survey responses and—where tracking is enabled—richer behavioural data including heatmaps. Its current live-site results guide is the best source for what each result type means and what your setup captures.
In practical terms, I would use Maze to answer questions like:
Can users complete the task without help?
Where do they hesitate, take an unexpected route or abandon the journey?
What language or control creates the wrong expectation?
How do the behavioural signals compare with what users say afterwards?
A session recording or heatmap is not a verdict. It is a clue. Five people clicking the wrong label might reveal a labelling problem, a prototype limitation or an unrepresentative sample. The follow-up conversation is where the evidence becomes insight.
How we applied this to an e-invoicing redesign
The most useful example from my own work was an e-invoicing redesign at ShipServ. This was not a speculative “what if an AI made invoices nicer?” exercise. The invoicing experience sat in the middle of a live, monetised workflow, where accuracy, confidence and auditability mattered just as much as speed. An invoice that looks beautifully streamlined but cannot be reconciled, queried or approved is not good UX; it is a very expensive piece of theatre.
The existing journey involved several connected jobs: receiving an invoice, reading the supplied information, matching it to the underlying purchase order and goods receipt, resolving mismatches, routing it for approval, and keeping a clear record of what happened. Different people entered the flow with different priorities. An accounts-payable user needed to deal with exceptions quickly. A buyer wanted confidence that the charge matched what had been ordered. An approver wanted enough context to make a decision without opening five tabs and consulting a spreadsheet from 2017.
Start with the workflow, not the shiny feature
Before making a prototype, I broke the work into the decisions users actually had to make. That prevented us from treating automated extraction as a magic black box and helped us focus on the experience around it.
Invoice intake: Can the user see what has arrived, what is ready to process and what needs attention?
Extraction review: Can they quickly understand what the system has read from the invoice, and which fields need human confirmation?
Matching: Can they see the relationship between invoice lines, purchase orders and receipts without losing the thread?
Exception handling: When values, quantities or suppliers do not match, is the problem clear and is the next action obvious?
Approval: Does an approver have enough evidence to approve, reject or send back an invoice with confidence?
Audit trail: Can the team later understand what changed, who changed it and why?
That list became the backbone of the prototype and the research plan. It sounds obvious, but it is surprisingly easy to build a clever demo that covers the happy path and quietly leaves the real work—exceptions, hand-offs and accountability—somewhere off-camera.
Using AI to get to something testable
We used AI-assisted prototyping to create interactive versions of the core screens and transitions quickly enough to compare approaches before implementation. The aim was not to generate a production-ready interface from a sentence. It was to get past static screens and give people a believable workflow to react to.
For example, we could explore how an extracted invoice might appear beside the purchase order it was being matched against; whether exceptions should be grouped into a queue or shown in context; and how much detail an approver needed at each point. We used realistic but safely fictional data so participants could reason about line items, quantities, currencies and status without exposing operational information.
The prototype also let us make the system’s confidence visible. That was important. An extraction tool should not present an uncertain value with the same visual certainty as a confirmed one. We explored clear statuses, highlighted fields that required review, and made it easy to compare the source information with the suggested value. The design principle was simple: automation should reduce clerical effort without hiding the evidence a person needs to trust it.
Testing the moments where trust can break
We did not ask participants whether they “liked” the redesign. We gave them tasks. Find an invoice that needs attention. Check why a line cannot be matched. Resolve a discrepancy. Decide whether an invoice is ready for approval. Explain what you would do next if the supplier value did not match the purchase order.
Those tasks revealed the points that mattered most:
Whether users could distinguish a system suggestion from a confirmed value.
Whether the reason for an exception was understandable at a glance.
Whether people could move from an exception to the relevant supporting information without losing their place.
Whether status labels reflected the language users used in their daily work.
Whether the interface gave approvers enough context to act without making them perform a forensic investigation.
This is where the combination of an interactive prototype and Maze was useful. The prototype made the decision points concrete; the test results showed where people hesitated, took an unexpected route or failed to complete the intended task. We then paired those signals with the feedback participants gave us. A misclick is a clue, not a diagnosis. The explanation from the person who made it is often where the real design work begins.
What changed because we tested early
The early rounds helped us refine the information hierarchy before development. We made exceptions more explicit, improved the way related records were surfaced, and gave the review state greater prominence so users did not mistake an extracted value for a verified one. We also kept the focus on a clear next action: resolve, query, approve or send back. In an operational workflow, that clarity is more valuable than a dashboard that merely looks calm.
It also gave the product team a better conversation with stakeholders. Instead of debating abstract opinions about an AI-enabled invoice flow, we could show a working concept and discuss evidence from representative tasks. That did not eliminate judgement calls—nor should it—but it moved the discussion closer to the realities of the people doing the work.
The principle I would carry forward
For invoicing, and for any workflow where automation affects money, compliance or trust, the most important design question is not “How much can AI do?” It is “Where does a person need clarity, control and proof?” Use AI to accelerate exploration and reduce repetitive effort. Then design the human checkpoints with the same care as the automated steps.
That was the real value of the exercise. We were able to learn about the core invoicing workflow earlier, turn assumptions into testable interactions and make better-informed choices before the cost of changing direction rose sharply. The AI helped us move faster. The research helped ensure we were moving in the right direction.



