A Maison credit card on a desk beside a laptop, pen and keys

Rethinking Product Discovery

Role
Sole researcher and designer

Industry
Finance

Duration
2 months

Reading Time: 12 minutes

Overview

A 51-participant study comparing form and conversational shopping that reshaped the product direction around how people actually evaluate and trust recommendations.

Challenge

Two ways to discover the same product: a form or a chatbot. I designed the same six-step shopping journey in both formats under an invented brand, Maison, removing the influence of an existing issuer’s reputation from the study.

Both experiences used the same recommendation and personalisation logic, isolating the interaction model as the primary variable. The question was whether changing that model would meaningfully affect how people perceived the experience, behaved within it, and ultimately made a decision.

Test Round #1Test Round #2
ResearchFramed the study: the same six-step shopping journey in two interaction models, form and chatbot, under an invented brand, so no issuer's reputation could colour the result.Study design
IdeationBoth experiences designed on identical recommendation and match logic, so the interaction model stayed the only variable. Round two came back through here carrying the first round's findings.Both rounds
PrototypeBuilt both flows as working prototypes rather than screens. Round two rebuilt them so both sides carried the same cards, figures and match logic.Rebuilt for round 2
Test51 unmoderated think-aloud sessions on participants' own phones: 26 in round one to find what broke, 25 in round two against hypotheses written before anyone was recruited.26 + 25 sessions
SynthesizeEvery transcript coded to timestamped moments, so each finding traced to a participant rather than an impression. The study changed the product direction instead of naming a winner.51 transcripts

Approach

I led the study end to end: framing the questions, designing and building both prototypes, running 51 unmoderated sessions on participants' own phones, coding transcripts, and synthesising the findings. I consulted with a senior and staff researcher throughout to challenge assumptions and surface blind spots.

Round one began as an evaluative study and became generative. Participants responded positively to both experiences, but what they said did not always align with what they did, and that gap set up round two.

I rebuilt both experiences and shifted from asking about preferences to observing behaviour: what people selected, where they focused, what they sought out, and what they needed before committing. Round two's measures were defined before anyone was recruited.

Context

Choosing a credit card is an infrequent, high-consequence decision. People must compare unfamiliar terms and competing benefits without the expertise to know which differences matter for their own financial priorities.

Where the evidence sits
Hover or tap a point
The most common reason people gave for trusting the recommendation: the questions it asked made the result feel personal. About 8 of 24 said this, more than any other reason.Round 2 · most-cited trust driverThe single most requested change, in both arms. People want to say rewards matter more than a low rate, not tick both as equally important.Rounds 1 and 2 · top requestAlmost everyone, in both arms, reported trusting the recommendation. If the study had stopped at asking, this is the answer it would have shipped.Round 2 · wrap-up surveyHigh and steady in the form. Split in the chatbot, where sevens sat next to twos and threes from the few people who actually checked something.Rounds 1 and 2 · confidence ratingIt led with a single match, and 8 of 12 never realised more cards existed. Round two showed three instead, and then everyone looked at two or more.Round 1 problem · fixed in round 2The same problem in a smaller form. Three are shown and five exist, and the last two sit behind Compare, so anyone who does not open it thinks they have seen everything.Round 2 · observed seam69% in the form and 75% in the chatbot, after looking at three cards on average. Whatever else people looked at, most took the one at the top.Round 2 · behavioural metricEach card has four things you can open to check it. Most people opened one or none before choosing.Round 2 · behavioural metricNot one person in 25 tapped to see how the match percentage was worked out. The number is believed on sight.Round 2 · behavioural metric

Research Impact

The study did not select a winning interface. It changed the product direction: guided, low-data personalisation with transparent, visibly complete choice.

72%
Took the top-ranked card

The top recommendation carried most of the decision weight, even after participants viewed a median of three cards.

92%
Made at most one check

Most people decided from information already visible on the card rather than opening supporting detail.

28%
Overrode the recommendation

A meaningful minority still used a personal priority, such as travel, credit building, or a flat-rate reward, to choose differently.

The recommendation became the decision

72% chose the top-ranked card and 92% made no more than one optional check. The first result and the information on its face carry most of the decision burden.

Conversation increased scrutiny rather than reducing it

The chatbot produced no blind selections and more validation activity than the form. The assumption that a warmer interaction would make people less careful was refuted.

Agency belonged in the recommendation logic, not extra controls

Form filters were used by 0 of 13 participants, while the most repeated request was to weight what mattered most. People wanted influence over the match, not more interface to operate.

Choice visibility was fixed, then revealed a smaller seam

Showing three cards solved the original single-card problem: every chatbot participant viewed at least two. But three of five available cards could still read as the complete set.

Prototypes

These are the rebuilt round-two experiences. Same cards, same figures, same match logic on both sides.

Form
Chatbot

How the Research Was Run

Round 1 · Identify what wasn't working

26 think-aloud sessions showed where people misunderstood money figures, missed available options, and hesitated to trust an unfamiliar financial experience.

Round 2 · Test the redesign against behaviour

25 sessions evaluated rebuilt, information-matched flows against pre-defined hypotheses, shifting the evidence from what participants said to what they actually did.

Trade-offs

Trust must be earned on the card face

People rarely opened supporting detail, so decision-critical information has to be visible before a tap. A deeper comparison tool cannot rescue a recommendation people have already accepted.

Guidance increased validation, but choice must look complete

The chatbot helped participants engage with the recommendation without producing blind choices. Its remaining risk was a shortlist that looked like the entire available set.

Visible choice did not translate into more agency

The form surfaced every card and enabled deeper comparison, but additional controls were ignored. Agency came from influencing the match logic, not navigating more UI.

Where I'd Take the Research Next

Measure decision quality, not just confidence

Use defined financial profiles with objectively better-fit cards, so confidence, behaviour, and recommendation accuracy can be evaluated together.

Isolate the interaction model

Hold information design constant from the outset, so the study can attribute differences to form versus conversation rather than content placement.

Observe more and prompt less

Use a small think-aloud sample to explain behaviour and a larger silent sample to measure it without changing it.

Scale before quoting rates

Increase and better match the sample before treating directional patterns in validation, time, or choice as conclusive rates.

The Next Experiment

That direction is one hybrid. It would make card count, rationale, and personal priority weighting visible before a person commits.

I would then test it silently at greater scale with known participant profiles. The question is no longer which interface feels better. It is whether the recommendation is accurate enough to deserve the trust people are already prepared to give it.

MATTHEW AHN © 2026