I followed Airbnb's search-to-booking journey through moderated usability sessions, a PURE evaluation and an AI evaluation. The most useful result was not a longer issue list. It was learning which findings held up when checked against another kind of evidence.
- Key insight
- Visible information can still be difficult to interpret or discover.
- Design direction
- Clarify price units, preserve filter state, and explain commitment at the decision point.
- Outcome / status
- Three evaluation perspectives, checked assumptions, and concrete recommendations. No redesign outcome was measured.
What I contributed
- Planned and conducted the usability study and subsequent evaluations
- Compared behavioral observations, PURE judgments and AI findings
- Checked findings against the interface and developed design recommendations
- Analyzed and wrote the research reports
Methods & making
Individual study; participants and UX-trained evaluators contributed evidence. Individual contributions are described separately from collective work.
The question worth asking
Booking can be straightforward until someone needs to interpret a price, change their dates or understand what cancellation would mean. I examined where these decisions became harder than they needed to be.
One journey. Three decisions.

Five participants worked through search and filtering, listing comparison, and booking initiation. No task required an actual payment.
The sequence connected early budget choices with later interpretation of price and commitment.
My work: I planned and ran the study, analyzed the findings and wrote the reports.
Formative research across laptop and native mobile environments. The later PURE and AI work revisited desktop steps; this was not a controlled comparison.
Human. Evaluator. AI.
Each method made a different part of the experience visible. Select a perspective, then compare what it could—and could not—establish.
Human
Changing travel dates could reset filters, sending people back through work they had already done.
Behavior across a sequence: hesitation, repetition and uncertainty during real interaction.
5 participants · Mobile and laptop use · Small formative sample; device differences affect interpretation.
36 steps. Four ways to read the friction.
Start with the AI-assisted step inspection, trace the recurring themes, then compare the priorities from the human evaluators. Each original artifact preserves its method, ratings and annotations.
36 steps

I reused the evaluator task screens for a structured AI pass: 14 search and filtering steps, 11 listing-review steps, and 11 booking steps. This made it possible to trace an interpretation back to a specific part of the journey.
AI-assisted PURE-style evaluation: 14 search steps, 11 listing-review steps and 11 booking steps.Use the step map to choose where to inspect more closely, then check its claims against the interface and the other evidence.
This was an inspection of supplied screens, not autonomous interaction or a controlled comparison of methods. The earlier human sessions included mobile and laptop use.
My work: evaluation, cross-method synthesis and research reporting.
Follow the pattern back to the screen
These are the original interface captures with the research annotations intact. Each one connects a theme to a visible element; the interpretation still needs the human and evaluator context.
Price · reveal, then reconstruct

The total is visible, but its breakdown requires another interaction. The human study also recorded a nightly-versus-trip-total misunderstanding.
Original listing and price-breakdown capture with the report annotations preserved. The AI interpretation was checked against the human and evaluator findings.Make the unit explicit and make the breakdown easy to find.
RM4 p.9, Fig.5; human evidence in RM2 pp.14–16.
The red annotations belong to the original evaluation artifacts. They illustrate the study findings and hypotheses; they do not represent tested redesign outcomes.
A changed date should not erase the work.
Observed behavior exposed a sequence-dependent problem. A single screen did not tell the whole story.
Find the filter

The filter entry point appeared after search in the studied interface.
Source evidence illustration: the filter entry point appears after search. RM2, p.12, Figure 1.Make the next available action easier to discover.
My work: research, evaluation and synthesis. Original Airbnb interface; report annotations are proposals where labeled.
Where an action needed explanation
The AI pass also raised three interaction questions worth checking in a follow-up study. They remain hypotheses rather than newly observed participant failures.
Choosing dates

The calendar relies on an inferred check-in, then check-out sequence.
AI inspection hypothesis: selecting check-in and check-out relies on an inferred sequence. This is not a new user-test result.Test whether travellers can identify the active selection state without extra explanation.
AI hypothesis: RM4 p.19, Fig.13.
These suggested checks were not conducted as a post-redesign validation study.
Where the evidence met—and diverged
Compare the issue across methods. Agreement and disagreement are both useful.
Understanding the total cost
Human: Nightly and total-stay prices were confused during use. Evaluator: Multiple price elements increased interpretation effort. AI: Flagged price interpretation and breakdown visibility.
Agreement across different forms of evidence
Agreement helped. So did disagreement.

Price interpretation, cancellation clarity and comparison recur across the reports. Filter persistence is strongest in observed use, with different coverage in the later inspections.
Use differences to improve the next test: include changed dates and revised filters, rather than only a forward booking path.
My work: cross-method synthesis and critical comparison.
The matrix is qualitative synthesis. Sequential studies, reused materials and unequal device conditions do not isolate a method effect.
The feature already existed.
An AI suggestion meets the interface
“Add a numeric price input.”
The AI evaluation suggested a feature was missing. A plausible recommendation still needed checking.

From “missing” to easier to notice.
These are original source crops. The proposed cue is not a shipped or validated change.
Existing numeric fields

The original capture already contains minimum and maximum numeric inputs. The reports also describe typing as available.
Original interface evidence, cropped from RM2, p.15, Figure 4. The numeric fields were already present.Reject the absence claim and reframe the problem as discoverability.
My work: research, evaluation and synthesis. Original Airbnb interface; report annotations are proposals where labeled.
Make each recommendation testable.
Explain the total

The report makes a price-breakdown entry point more visible.
Original annotated proposal: make the price breakdown easier to discover. RM4, p.34, Figure 23. No follow-up validation is reported.Ask people to explain the amount before proceeding.
My work: research, evaluation and synthesis. Original Airbnb interface; report annotations are proposals where labeled.
What came out of it
A set of research-grounded recommendations, with the AI assumption corrected and the limits of each method made explicit.
What this work does—and doesn’t—show
This was not commissioned by Airbnb and was not a shipped redesign. Recommendations were not validated in a follow-up redesign test. The methods were not a controlled comparison: human sessions included mobile and laptop use, while later evaluations focused on desktop. The sample was formative; it does not establish population-level or commercial outcomes.
Read the original research report
For the full analysis, the 51-page academic report includes the step ratings, original annotated interfaces, cross-method synthesis and proposed recommendations.
Original 51-page research report. Public copy: eight private appendix links and their link labels removed; document metadata cleared. Original layout, screenshots, ratings, findings and proposed recommendations retained. The report’s original AI misinterpretation is retained alongside its later correction; recommendations remain unvalidated.
What I’d do next
I would test the proposed cues with the same decision journey, including changing dates after setting filters. I would check whether people can explain the total cost and cancellation deadline, find precise price entry, and retain their comparison context. Those would be new validation results—not outcomes of this study.
