Personal Finance
Personal Finance
Personal Finance
|
|
|
Dec 2023 – Feb 2024
Dec 2023 – Feb 2024
Dec 2023 – Feb 2024
Designing an AI copilot that catches what you'd miss.
Designing an AI copilot that catches what you'd miss.
Designing an AI copilot that catches what you'd miss.
Detection types
Detection types
Detection types
3
3
3
Duplicate, category,
missing receipt
Duplicate, category,
missing receipt
Duplicate, category,
missing receipt
Confidence floor
Confidence floor
Confidence floor
85%
85%
85%
Before any change is
suggested
Before any change is
suggested
Before any change is
suggested
Resolved without you
Resolved without you
Resolved without you
0%
0%
0%
Nothing moves on
its own
Nothing moves on
its own
Nothing moves on
its own
ROLE
ROLE
ROLE
Senior Product Designer
Senior Product Designer
Senior Product
Designer
TOOL
TOOL
TOOL
Figma + FigJam
After Effects
Notion + Jira
Framer
Claude Code
Figma + FigJam
After Effects
Notion + Jira
Framer
Claude Code
Figma + FigJam
After Effects
Notion + Jira
Framer
Claude Code
WHAT I'M DOING
WHAT I'M DOING
WHAT I'M DOING
Research, IA,
wireframes to hi-fi,
motion, prototype,
handoff & QA
Research, IA,
wireframes to hi-fi,
motion, prototype,
handoff & QA
Research, IA,
wireframes to hi-fi,
motion, prototype,
handoff & QA
PLATFORM
PLATFORM
PLATFORM
Mobile ( Andriod /
IOS)
Mobile ( Andriod /
IOS)
Mobile ( Andriod /
IOS)
TEAM
TEAM
TEAM
PM + 3 Engineers
PM + 3 Engineers
Founder +
2 Eng
01 — Context & Overview
01 — Context & Overview
01 — Context & Overview
What is Potion?
What is Potion?
Potion is a personal-finance copilot for people who keep their own books. Freelancers, sole traders, and anyone juggling money across a few accounts and cards. It watches every transaction and surfaces the quiet errors that usually only show up months later.
The hard part was never the detection. It was the handoff. How does an AI hand a suspected error to a person without either taking over the decision or getting ignored? I designed the answer to that, mobile-first, and shipped it 0 to 1.
Potion is a personal-finance copilot for people who keep their own books. Freelancers, sole traders, and anyone juggling money across a few accounts and cards. It watches every transaction and surfaces the quiet errors that usually only show up months later.
The hard part was never the detection. It was the handoff. How does an AI hand a suspected error to a person without either taking over the decision or getting ignored? I designed the answer to that, mobile-first, and shipped it 0 to 1.


02 — Problem
02 — Problem
02 — Problem
People find their own money errors by accident.
People find their own money errors by accident.
When you do your own books, mistakes do not announce themselves. They sit quietly in the ledger and compound until tax time, when finding them is far more expensive than catching them would have been.
When you do your own books, mistakes do not announce themselves. They sit quietly in the ledger and compound until tax time, when finding them is far more expensive than catching them would have been.


03 — Research & Insights
03 — Research & Insights
03 — Research & Insights
The person who has to vouch for it.
The person who has to vouch for it.
I spoke with freelancers and sole traders who do their own books, and watched how they reconcile at month end. The gap was never the math. It was attention, and the trust that follows from it. One person came to stand in for what they all shared, and one insight reframed the whole problem.
I spoke with freelancers and sole traders who do their own books, and watched how they reconcile at month end. The gap was never the math. It was attention, and the trust that follows from it. One person came to stand in for what they all shared, and one insight reframed the whole problem.




04 — Ideation & Exploration
04 — Ideation & Exploration
04 — Ideation & Exploration
Mapping the flow before designing the screens
Mapping the flow before designing the screens
Early on, the flows split into two philosophies. One let the AI quietly fix anything it was confident about. The other treated every flag as a decision that belonged to the person. I built both to feel the tradeoff before committing.
Early on, the flows split into two philosophies. One let the AI quietly fix anything it was confident about. The other treated every flag as a decision that belonged to the person. I built both to feel the tradeoff before committing.
Early on, the flows split into two philosophies. One let the AI quietly fix anything it was confident about. The other treated every flag as a decision that belonged to the person. I built both to feel the tradeoff before committing.
Explored
Explored
Explored
Auto-resolve high-confidence matches
Auto-resolve high-confidence matches
When the AI was very confident, an obvious duplicate or a clear category fit, it could resolve the issue silently and surface only the ones it was unsure about. Fewer interruptions, a cleaner feed.
When the AI was very confident, an obvious duplicate or a clear category fit, it could resolve the issue silently and surface only the ones it was unsure about. Fewer interruptions, a cleaner feed.


The tradeoff that killed it
The tradeoff that killed it
Confidence is not correctness. A person's trust in the tool depends on seeing every correction, not just the uncertain ones. Auto-resolving even the obvious cases meant a silent edit to the books that nobody signed off on.
Confidence is not correctness. A person's trust in the tool depends on seeing every correction, not just the uncertain ones. Auto-resolving even the obvious cases meant a silent edit to the books that nobody signed off on.
Confidence is not correctness. A person's trust in the tool depends on seeing every correction, not just the uncertain ones. Auto-resolving even the obvious cases meant a silent edit to the books that nobody signed off on.
SHIPPED
SHIPPED
SHIPPED
Every flag stops for a decision
Every flag stops for a decision
A duplicate charge or category mismatch always surfaces with the reasoning attached and the likely causes already laid out. A person picks the resolution. Nothing gets fixed without someone seeing it happen.
A duplicate charge or category mismatch always surfaces with the reasoning attached and the likely causes already laid out. A person picks the resolution. Nothing gets fixed without someone seeing it happen.

Why it won
Why it won
The pattern held even for easy cases. AI does the reasoning; a person makes the call. That consistency is what lets people trust it with the numbers that actually mattered, instead of wondering what it had quietly changed.
The pattern held even for easy cases. AI does the reasoning; a person makes the call. That consistency is what lets people trust it with the numbers that actually mattered, instead of wondering what it had quietly changed.
05 — Design Execution
05 — Design Execution
05 — Design Execution
AI flags it. You decide.
AI flags it. You decide.
Every flag resolves inside the same card. The transaction stays in view, the reasoning is spelled out, and the fix is a short list of causes the AI already thought through.
Every flag resolves inside the same card. The transaction stays in view, the reasoning is spelled out, and the fix is a short list of causes the AI already thought through.
One conversation handles every flag
One conversation handles every flag
One chat handles everything flagged. Potion explains each and offers the fixes as replies. Nothing changes until you answer.
One chat handles everything flagged. Potion explains each and offers the fixes as replies. Nothing changes until you answer.
One chat handles everything flagged. Potion explains each and offers the fixes as replies. Nothing changes until you answer.
Reasoning and confidence, then a one-tap choice
Reasoning and confidence, then a one-tap choice
The reasoning and a confidence score, then a one-tap category fix. Filed under Shopping, the pattern points to Software.
The reasoning and a confidence score, then a one-tap category fix. Filed under Shopping, the pattern points to Software.
The reasoning and a confidence score, then a one-tap category fix. Filed under Shopping, the pattern points to Software.
Both charges, side by side
Both charges, side by side
Both charges shown side by side. You say what happened, and Potion resolves it and tracks any refund.
Both charges shown side by side. You say what happened, and Potion resolves it and tracks any refund.
Both charges shown side by side. You say what happened, and Potion resolves it and tracks any refund.
A short list of real causes, never a blank investigation
A short list of real causes, never a blank investigation
A payment with no receipt attached. Potion offers to find it in your email, mark it as paid in cash, or upload one. Match it and it attaches to the transaction, so the books stay complete.
A payment with no receipt attached. Potion offers to find it in your email, mark it as paid in cash, or upload one. Match it and it attaches to the transaction, so the books stay complete.
06 — Iteration & Testing
06 — Iteration & Testing
Three decisions that made the AI feel trustworthy, not intrusive.
Three decisions that made the AI feel trustworthy, not intrusive.
Testing pushed the design toward restraint. Each round, the version that earned more trust was the one that did less on its own and showed more of its thinking.
Testing pushed the design toward restraint. Each round, the version that earned more trust was the one that did less on its own and showed more of its thinking.
Testing pushed the design toward restraint. Each round, the version that earned more trust was the one that did less on its own and showed more of its thinking.

07 — Impact & Results
07 — Impact & Results
Catching what a person would have missed.
Catching what a person would have missed.
The detection layer caught duplicate charges and category mismatches that manual review had been missing, the exact kind of quiet error that compounds across a growing set of accounts.
The detection layer caught duplicate charges and category mismatches that manual review had been missing, the exact kind of quiet error that compounds across a growing set of accounts.
The detection layer caught duplicate charges and category mismatches that manual review had been missing, the exact kind of quiet error that compounds across a growing set of accounts.

08 — Reflection
08 — Reflection
What I took from it.
What I took from it.
The hard design problem in an AI product is rarely the automation. It is the handoff. Who does the work, who owns the outcome, and whether the reasoning is visible enough to make the review real instead of a rubber stamp.
This is the same shape across very different domains. Any high-stakes, high-volume manual job where the AI can draft but should not decide wants the same structure. Solving it here made it portable.
Confidence is not correctness. Designing for the moment a system is sure and still wrong is what earns trust.
The hard design problem in an AI product is rarely the automation. It is the handoff. Who does the work, who owns the outcome, and whether the reasoning is visible enough to make the review real instead of a rubber stamp.
This is the same shape across very different domains. Any high-stakes, high-volume manual job where the AI can draft but should not decide wants the same structure. Solving it here made it portable.
Confidence is not correctness. Designing for the moment a system is sure and still wrong is what earns trust.
The hard design problem in an AI product is rarely the automation. It is the handoff. Who does the work, who owns the outcome, and whether the reasoning is visible enough to make the review real instead of a rubber stamp.
This is the same shape across very different domains. Any high-stakes, high-volume manual job where the AI can draft but should not decide wants the same structure. Solving it here made it portable.
Confidence is not correctness. Designing for the moment a system is sure and still wrong is what earns trust.
What I'd improve
What I'd improve
Watch for review fatigue. Stopping for every flag is right for trust, but at high volume it can become noise. I'd test a weekly digest that batches similar low-stakes flags without ever resolving them automatically.
Introduce a trust budget. After a run of confirmed corrections of the same type, let a person opt in to grouping them, keeping the decision theirs while cutting repetition.
Instrument the skip. "Skip for now" is a signal I under-measured. Tracking what gets skipped and why would sharpen which flags are worth raising at all.
Watch for review fatigue. Stopping for every flag is right for trust, but at high volume it can become noise. I'd test a weekly digest that batches similar low-stakes flags without ever resolving them automatically.
Introduce a trust budget. After a run of confirmed corrections of the same type, let a person opt in to grouping them, keeping the decision theirs while cutting repetition.
Instrument the skip. "Skip for now" is a signal I under-measured. Tracking what gets skipped and why would sharpen which flags are worth raising at all.
Watch for review fatigue. Stopping for every flag is right for trust, but at high volume it can become noise. I'd test a weekly digest that batches similar low-stakes flags without ever resolving them automatically.
Introduce a trust budget. After a run of confirmed corrections of the same type, let a person opt in to grouping them, keeping the decision theirs while cutting repetition.
Instrument the skip. "Skip for now" is a signal I under-measured. Tracking what gets skipped and why would sharpen which flags are worth raising at all.