Procurement happens on one platform. Payment happens on another. Amazon Business, NetSuite, SAP, email, spreadsheets. The invoice, the purchase order, and the receipt all live in different places, and someone has to hold the whole picture in their head. Witness is that single intelligent layer. It pulls the three documents together, does the matching, and flags what's off. The human stays the final decision-maker, always.
Witness is a concept product for accounts payable teams who process invoices against purchase orders and receipts, the classic "3-way match," using Amazon Business as the purchasing platform and an ERP like NetSuite or SAP S/4HANA on the other end. Instead of switching between four or five systems to reconcile one invoice, everything relevant lives in one workspace, and AI does the first pass of the matching while a person makes the call.
Give finance teams an AI that does 3-way matching for them.
People didn't want AI to do the matching for them. They wanted it to do the tedious cross-checking and then show its work, because the accuracy of a payment decision is their responsibility, not the tool's.
Product Designer. I built the design system in Figma, generated the design.md spec, and built the interactive prototype end to end, from information architecture through a working multi-page HTML/CSS/JS build.
Figma for the design system and screens, Claude for translating that system into a working prototype. I directed every decision; the tools moved faster.
A full multi-page interactive build on realistic, simulated accounts-payable data. Three moments from it, click to play.
This interactive prototype includes the following workflows.
What this design gets right today, and what it's built to deliver at scale.
Not "can AI do the match." It's "how do you let AI do the match without the human losing the thread of why."
Making the AI feel active, like named agents at work, not a quiet background feature.
Role-based views for AP specialist, controller, and approver: the next thing I'm designing, not a finished feature yet.
The full story: the research, the pivot, what I built, and how AI became part of the build process itself.
Before sketching a single screen, I wanted to understand where a 3-way match actually breaks down for the people who do it every day, and how much of that was a fragmentation problem versus a trust problem.
Three findings kept surfacing across the survey and interviews: people were constantly switching systems to reconcile one invoice, they had no centralized view of payment status, and they didn't trust AI enough to hand off a financial decision without checking it. The next section breaks down exactly what that meant for design.
"[AI] makes mistakes. It makes mistakes on every part. But as I said, it's my job to make sure that AI does the work correctly." Interview participant
That single line reframed the whole project for me. The job wasn't to build an AI that never makes mistakes. It was to build a system where a mistake is easy to catch, and where the human never has to wonder what the AI actually did.
A journey map from requisition to payment confirmed where confidence drops hardest: invoice generation, 3-way matching, and payment approval. The exact moments where a wrong call carries a dollar amount. That's where I focused the design.
None of these needed a prettier screen. They needed a different structure: one workspace, transparent reasoning, and a way for people to control exactly how much AI does.
The same invoice lives in four places at once, and reconciling it means opening all four, every time.
Accuracy is the top concern in financial workflows. People won't hand off a decision they can't verify.
Manual entry errors and vendor discrepancies create repeated verification work and email back-and-forth.
No single place shows payment status, blockers, or compliance state, so people track things twice.
An AI that runs 3-way matching and outputs a pass/fail reads well on paper. It also asks a finance professional to trust a black box with money leaving the company.
People wanted AI to do the tedious cross-referencing, then hand back a clear, checkable answer, evidence attached, so a human could sign off with confidence, not blind faith.
That distinction shaped every design principle that followed: the system should centralize the fragmented workflow, keep every AI action transparent and explainable, and always leave the final authority with the person, not the model.
AI assists through recommendations and automation. Final authority always stays with the person.
Pull everything about one invoice into a single, coherent view instead of five open tabs.
Every AI action is predictable and explainable. The person can see why, not just what.
Data is handled reliably, and every recommendation is dependable enough to act on.
Witness is designed for the Accounts Payable Specialist or Finance Professional who processes invoices day in and day out. The dashboard, the chat agent, and the duplicate catch are all built for that one job.
Works the invoice queue. Reviews matches, resolves flagged mismatches, and clears the routine cases AI already handled.
Every screen below is the actual working prototype, running inline, not a screenshot. Built on realistic, simulated accounts-payable data. Scroll inside any frame to explore it.
Invoices sort into "must handle today," "waiting on others," and "handled by AI." The question isn't what exists, it's what needs a person right now.
Testing showed people assumed AI-handled meant fully done. I added a visible review count so no one is surprised later.
The Witness agent answers in plain language, right inside the workspace. Ask about an invoice, a vendor, or a payment status, and it retrieves the answer instead of sending you to go find it.
Ask about INV-1027 and the agent pulls that invoice's actual history, its PO, its match status, and its flags, live from the workspace it's sitting in.
The same chat panel follows the user across every page, so context never resets when they move from the dashboard to an invoice.
Approve, dispute, reroute, or escalate. Surfaced first, so the next step is visible without scrolling.
Witness flags INV-1027 against Praxis Logistics as a likely duplicate of a previously paid invoice tied to PO-7718. Confidence score and financial exposure stated plainly: a $890 double-payment risk.
The source invoice image sits before the AI's reasoning, so the primary document is checked first, not last.
People wanted moderate to high automation, but disagreed on how much supervision they needed. Autoflow makes that adjustable per step.
Testers were split on per-vendor control. I made vendor-specific autoflow opt-in, not a forced step.
Trust links directly to explainability. The log shows what happened, why, and when, for every invoice.
If AI got a match wrong, reverting it is one click from the log, not a support ticket.
Four screens build understanding before any configuration. Four more make it real: connect tools, set preferences, choose autonomy.
Assist: surface what matters, human decides everything. Draft for review: AI prepares, human approves. Handle it: AI runs it end to end and reports back.
"I surface what matters. You review and decide everything." Lowest automation, highest oversight.
Best for: new vendors, high-dollar invoices, or the first weeks after setup, anywhere trust hasn't been earned yet.
"I prepare the action. You approve before anything happens." The middle ground most people land on.
Best for: routine vendors with an established pattern, where a quick approval beats a full review.
"I take care of it end to end and keep you informed." Highest automation, for low-risk, high-confidence cases only.
Best for: a repeat invoice from a trusted vendor that has matched cleanly before, with nothing unusual flagged.
I wanted a prototype real enough to usability-test properly, not a click-through of static frames. That meant building an actual working front end, and using AI to do it faster without giving up design control.
The full design system, plus key screens like the Dashboard designed in detail, so the AI had a concrete design language to follow, not just a token list.
Step 1Generated a design.md file from the Figma system: tokens, colors, typography, and spacing rules documented as a spec.
Step 2Built the prototype from that spec, following the reference screens, including a few design variations for select screens before I picked and refined the final direction.
Step 3Every AI output reviewed and hand-edited for correctness before it shipped; version control kept design decisions traceable.
Step 4Stack, accurately: vanilla HTML, CSS, and JavaScript. No framework, no build step. Nine full pages: dashboard, all invoices, invoice detail, autoflow, autoflow editor, AI activity, vendors, settings, and onboarding. It's never processed a real invoice, every number in it is simulated.
Warm and confident rather than sterile, with enough restraint that a duplicate-invoice warning still reads as urgent rather than decorative.
The chat panel shared across every page, so the agent's context follows the user rather than resetting per screen.
A shared data-wt-date attribute keeps every "3 days ago" and due-date label accurate and consistent across all nine pages.
An 8-ray sunburst rotated 45°, used as the product's logo across onboarding, sign-in, and the workspace.
Design system built and documented in Figma; extracted into code via Claude Code and hand-maintained across the prototype.
I ran moderated think-aloud sessions with domain experts in procurement and accounts payable, walking through two scenarios: an invoice AI handled overnight, and one requiring mandatory human review. Here's what changed as a direct result.
| What I Noticed | What I Changed | Why |
|---|---|---|
| Duplicate & exception alerts were easy to miss | Made key alerts visually louder: stronger color, larger text | Trust |
| Invoice detail page felt dense, next step buried | Moved secondary info into tabs, and surfaced the recommended action at the top | Clarity |
| "Handled by AI" read as fully done, no review needed | Added a visible count of items still requiring review | Trust |
| Testers split on whether vendor-specific rules should be mandatory | Made vendor-specific autoflow an optional setting | Control |
Findings from 3 domain-expert usability sessions on a working prototype using simulated data.
Early explorations reduced everything to one trust score. I moved away from it. A single number hides which part of a match is uncertain, and financial decisions need the specific mismatch, not a vague percentage.
It was tempting to frame "Handle It" as the best option. I didn't. Higher automation isn't higher quality. It's a tradeoff the controller should make deliberately, per vendor and per risk, not a ladder to climb.
Every screen right now shows AI being right. The highest-value thing I haven't built yet is the flow for when a person overrides a flagged match: what that correction looks like, and how it teaches the system.
An AP specialist and a controller shouldn't see the same screen. It's sketched, not shipped, the honest next step, not a finished claim.
What I actually learned building this: writing a prompt for a financial screen is the same discipline as writing a design brief. Vague input gets vague output, and the quality of what Claude produced tracked directly with how precisely I'd already thought through the interaction. AI compressed the build. It didn't replace the judgment about what should exist in the first place, that decision, on every screen, stayed mine.