
Card Finder Research Study
Overview
A study evaluating form-based and conversational credit card shopping experiences, focused on understanding and evaluating trust, effort, and decision confidence.
Challenge
Two ways to discover a product: a form, or a conversation. The same six-step journey was built both ways under a fictional brand, driven by the same recommendation and personalisation engine, so the only real variable was how you were asked. Whether that changes perception, behaviour and interaction was the open question.
Approach
51 unmoderated think-aloud sessions across two rounds, with a senior and a staff researcher alongside me to frame the questions and shape the test plan. The first round was meant to be evaluative and turned out generative: we were still working out which questions would surface how someone actually shops for a credit card. The feedback was positive, but what people said did not quite match what they did, and that gap is what set up the second round. I refined both flows and we shifted from asking to watching: what people selected, what held their attention, what they checked before committing. Trust stopped being something reported and became something we could count.
Context
Choosing the right credit card can be surprisingly difficult. With dozens of options, unfamiliar terminology, and benefits that are hard to compare, people are often left sorting through information without knowing which features actually matter for their needs. Because it is a decision made so infrequently, most people never build the knowledge or confidence to navigate it easily.
So the same six-step journey was built twice under a fictional brand, Maison, chosen so no real issuer's reputation could tilt the answer.
How might we design a credit card shopping experience that helps customers find the card best suited for them?
How is trust and understanding built in the context of a credit card shopping experience?
What decision-making patterns and mental models are at play?
Methodology
Two rounds, 51 unmoderated think-aloud sessions on smartphones, deliberately built to answer different questions. The first round went looking for problems. The second measured whether the fixes worked, and it counted rather than asked.
Round one, generative
26 sessions, 14 on the form and 12 on the conversation. Recordings were transcribed locally, then every meaningful utterance was coded against a fixed taxonomy: an observation code, the stage it happened in, a controlled theme, a severity where it was a problem, a timestamp, and an id. The rule underneath it is that no claim floats free of its source, so any finding on this page can be walked back to the second of session it came from.
Round two, evaluative
25 sessions on rebuilt v2 designs, 13 on the form and 12 on the conversation. Eight hypotheses and the decision rules that would settle them were written down and frozen before anyone was recruited, so the analysis could not quietly move its own goalposts once the data arrived.
Prototypes
Both round-two builds, live and side by side at the size they were tested on. Same cards, same figures, same match logic. Work through either one and the comparison the study ran is the one in front of you.
Trade-offs
Neither paradigm won. Each one bought something real and charged for it somewhere else, and the useful output of the study is the price list rather than a verdict.
What conversation buys, and what it costs
The conversational frame is the more generous one and the more exposed one. Its warmth is a genuine differentiator, but it asks for something before it has proved anything, and it promises personalised math that people then want to see. A flow that offers more is held to more.
What the form buys, and what it costs
The form is the steadier of the two. It shows everything at once, never asks who you are, and its side-by-side comparison was the most-used feature in the study. What it gives up is attention: a number skimmed in a list of cards is a number read wrong, and the one it lost was the one that mattered most.
The two trade-offs underneath both
Scrutiny against friction. A verification affordance only earns its place if people open it. The score explanation was tapped by nobody in either arm across 25 sessions, while the fine print behind an expander was opened by 15% and 33%.
Decision-making
The mental model underneath every session was the same in both flows. People do not want to become experts in credit cards. They want a trustworthy expert to hand the problem to, and they keep a light veto for when the answer does not fit.
That is textbook bounded rationality: full comparison is expensive, so people satisfice and delegate to a signal they trust. What the behavioural round added was the ability to see it happening rather than infer it.
Across both arms, people committed having opened no more than one of the four optional checks on the card. Barely-checked picking is the norm, not the exception.
69% in the form and 75% in the chat chose the rank-one card, despite viewing a median of three. Viewing more did not move most people off the first.
Not one participant in either arm tapped to see how the match was calculated. The percentage is trusted as a number, never interrogated as a claim.
A real minority rejected the top pick on their own logic: a flat 2% beating a category 5%, building credit, travel. Delegation, with a veto held back.
Quick is not the same as careless. Almost nobody chose truly blind: the light checkers did look, they just read what was already open on the card rather than tapping to expand anything. But the people who opened nothing decided in a mean 67 seconds against 147 for those who opened at least one check, which is worth remembering the next time a fast, tap-free run gets read as a happy user.
The questions, the searching beat, and the resulting match made the pick feel built for them. This was the most-cited reason for trusting it, about 8 of 24 wrap-ups.
The match percentage reads as evidence that something reasoned. Nobody checked the reasoning, but its visible presence is what made leaning on it feel safe.
No forms, no credit check to browse, almost nothing personal asked. Restraint removed the threat, and it was the second-cited trust driver, about 5 of 24.
Insights & Feedback
Round one found the problems. Round two tested eight predictions against what people actually did, and the two it got wrong were more useful than the five it got right.
5 of 14 read the large estimated-rewards number as a cost. "I thought that was the annual fee." APR sat behind an expander and surfaced only after a card was chosen, too late for the people comparing.
It led with one strongest match. 8 of 12 did not realise more cards existed until they found "show another option", and the total count was never stated.
Asked for a name before showing anything, with no brand and no regulatory marks, several questioned whether it was a real financial institution. One typed fake information rather than trust it.
The match percentage paired with a plain-language reason for the shortlist. Alongside it, no credit check to browse was the strongest trust driver in the whole study.
The pre-registered prediction was that warmth would tax vigilance. The opposite happened: more checks opened in the chat, a median of 1 against the form's 0, and zero blind picks against the form's 8%. It survives removing the non-native speakers.
Its filter tabs went completely unused, 0 of 13, and cards viewed was identical in both arms. The assumption that a form affords more agency than a conversation did not hold.
Searching all 25 round-two transcripts for FDIC, accredited, legitimate, scam, fraud and never-heard-of returns nothing. Not one participant questioned who Maison was. Several things changed at once, so read the direction, not the mechanism.
A short searching beat before the results, checking your answers against rewards, fees and how you spend, drew unprompted delight. A labour-illusion moment that cost little and read as competence.
Everyone now viewed two or more cards, so the old illusion is behaviourally resolved. In its place, three shown out of five available reads as all of them unless you open Compare.
Let people weight what matters rather than tick everything equally. "I want low rates but rewards are more important to me." Equal-weight multi-select is the shared ceiling on how precise the match can be.
Of the eight pre-registered hypotheses, five held, two were refuted and one was confounded by its own instrument. Both refutations cut against the same intuition, that a friendly conversation dulls scrutiny and a form hands back control, and neither survived contact with the recordings. That is the finding that actually changes a roadmap decision, and it only exists because the prediction was written down before the data could soften it.
What We'd Do Better
A study is only worth as much as the limits you are willing to state about it. These are the ones that would change how the next round is run, listed in the order they would change it.
Round one's arms differed in interaction and in information design at the same time: APR on the card face in one and behind an expander in the other, one card surfaced against all of them. So "conversation fixes money legibility" could just as easily have been a layout decision any form could adopt. Round two enforced information parity, identical cards and figures and prominence, so only the frame varied.
Round one asked how sure people felt and never whether they chose well. No participant was given a profile with a correct answer to reach, so decision quality was unmeasurable. Confidence without accuracy was the study's structural blind spot, and it is the one thing a satisfied cohort can hide.
Asking people to say what felt clear, confusing, trustworthy, unnecessary or missing primes them to hunt for exactly those. Some desires were elicited by the question rather than felt in the moment. Round two moved every probe of that kind after the task and coded it as elicited.
Narrating makes people rationalise and over-explain, which can make a flow look better understood than it is, and it distorts every timing. Round two split into a silent instrumented cell for the numbers and a small think-aloud cell for the mechanism.
The chat arm drew four non-native English speakers the form arm did not, which inflated every time metric: a median 5:43 for that sub-group against 3:53 for the native speakers. Native-only, the two arms sit near parity, so no raw time comparison can be quoted. The verification findings survive the correction; the timings do not.
Enough to find usability problems and read a 20-point gap or a distribution shape. Not enough to quote a rate. Moving these numbers with confidence would take roughly 80 per arm, which is why the recommendation was a larger silent round rather than a few more sessions.
An intro that promised a passcode the build never asked for, an exit loop that cost one participant minutes, a foreign transaction fee shown as 3% on one screen and 2% on the next. These dinged genuine confidence ratings and would mostly vanish in a real build, so they had to be separated from design findings rather than counted alongside them.
Paid panellists lean toward telling the team what it wants to hear, so general praise was discounted and convergent, specific frictions were weighted instead. Round one also captured every SUS answer and never computed the score, which left the polarised-against-consistent claim resting on reading rather than on a number.
The recommendation that came out of it was not a winner. It was one build carrying the conversation's tone and low data ask alongside the form's honest surfacing of every option, tested quietly at 30 to 40 per arm so the rates could finally be quoted rather than gestured at. The work was never really about choosing an interface. It was about making the recommendation honest enough to deserve the trust people were already handing it.