Building a website or application without understanding how users naturally categorize information is like organizing a library by your own personal filing system and expecting everyone to find what they need. I have watched product teams spend weeks debating menu labels, arguing about category placement, and guessing at navigation hierarchies. They launch. Users cannot find basic features. The problem is rarely effort. It is almost always evidence, or rather a lack of it. IA decisions made in conference rooms without user data are just opinions dressed up as strategy. Card sorting and tree testing exist to replace those opinions with real behavioral data. They are the closest thing we have to a user's mental map, made visible. Here is a practical guide to using both methods, interpreting the results honestly, and building products people can actually navigate.
Why Information Architecture Fails Without Research
The most expensive mistake in IA is assuming that the way your team thinks about content is the way your users think about it. Product teams live inside their products. They know every feature, every edge case, every piece of internal terminology. That familiarity creates a blind spot. What seems obvious to a designer who has spent months inside a product is often invisible to a user encountering it for the first time.
📑 Table of Contents
- Why Information Architecture Fails Without Research
- What Is Card Sorting?
- Open vs. Closed Card Sorting: When to Use Each
- Designing a Card Sort That Produces Reliable Data
- How Many Participants Do You Actually Need?
- Analyzing Card Sort Results: From Raw Data to Actionable IA
- Five Common Card Sorting Mistakes and How to Avoid Them
- What Is Tree Testing?
- Designing a Tree Test That Measures Findability
- Card Sorting vs. Tree Testing: How They Work Together
- Analyzing Tree Test Results: Success Rates, Time, and Path Data
- Integrating IA Research Into Your Design Process
- The Limitations You Should Not Ignore
- References
Take an e-commerce site selling outdoor gear. To the product team, it makes perfect sense to organize inventory by technical specifications: waterproof ratings, denier counts, insulation types. But the average shopper wants to find gear for "camping with kids" or "backpacking in the rain." The mental model of a technical product manager and the mental model of a weekend hiker are structurally different. Neither is wrong. But only one of them determines whether the user finds what they need and completes a purchase.
I have seen teams spend six-figure budgets on visual design, copywriting, and development only to discover in post-launch analytics that users bounce because they cannot find the "Contact Support" page hiding under "Company > About Us > Locations." That is not a visual design failure. It is an IA failure. And it is entirely preventable with forty minutes of moderated card sorting and a tree test that takes participants under fifteen minutes to complete.
What Is Card Sorting?
Card sorting is a research method where participants organize individual topics, content items, or features into groups that make sense to them. It is one of the most direct ways to understand how users mentally categorize information. Instead of asking users what they think of your proposed navigation, you give them raw content items and ask them to create their own structure from scratch.
The method started with physical index cards. Researchers would write each content item on a card, spread them across a table, and ask participants to sort them into piles. Digital tools like OptimalSort, UserZoom, and Maze have made remote card sorting faster and more scalable, but the principle stays the same: let users reveal their mental models by observing how they naturally group related items.
When I run a card sort, I am not testing whether users agree with my proposed navigation. I am learning how their brains already organize the information. That distinction matters because effective navigation does not teach users a new system. It aligns with the system they already carry in their heads. The closer your IA matches their existing mental model, the less cognitive effort they need to find what they want.

Open vs. Closed Card Sorting: When to Use Each
Card sorting comes in two variants: open and closed. Choosing the right one depends on where you are in your design process and what question you need to answer.
Open card sorting
In an open card sort, participants get a set of cards and group them into categories they create and name themselves. No predefined labels. This works best during early IA design, when you are trying to understand how users naturally chunk information and what labels they would use for those chunks. Open card sorting reveals the categories that exist in users' minds, which is often different from what your internal team assumes.
A university redesigning its website might give participants fifty cards representing campus services: financial aid, housing applications, course registration, career counseling, library access. An open card sort might reveal that participants consistently group "scholarship applications," "grant information," and "financial aid forms" under a label like "Paying for School." The university's internal team had them scattered across "Student Services," "Administration," and "Academics." That insight is actionable immediately.
Closed card sorting
In a closed card sort, participants get both the cards and a fixed set of category labels. Their job is to assign each card to one of the provided categories. This is used later in the design process, after you have a proposed IA structure and need to validate whether users can correctly map content items to your categories.
Closed card sorts are essentially compatibility tests between your proposed IA and user expectations. If a significant percentage of participants place a card in a different category than you intended, that is a signal that your category label or placement needs rethinking. The data is easier to analyze statistically because you are measuring agreement with a predefined structure, but it limits your ability to discover unexpected mental models.
My advice: start with an open card sort to understand mental models. Use those findings to build your initial IA. Validate that IA with a closed card sort or tree test before committing. Skipping the open sort and jumping straight to validation means you might validate a structure that never aligned with user thinking in the first place.
Designing a Card Sort That Produces Reliable Data
The quality of your card sort results depends almost entirely on how carefully you design the study. A poorly designed card sort can send your IA in the wrong direction. Here is how to get reliable results.
Selecting your cards
Every card should represent a single, discrete content item or feature. Do not combine multiple concepts on one card. Participants need to sort each item independently. Aim for 30 to 50 cards. Fewer than 30 may not capture enough complexity to reveal meaningful patterns. More than 50 risks participant fatigue, which leads to careless sorting and unreliable data.
Each card label should be short and clear. Avoid internal jargon, branded terminology users would not know, and overly long descriptions. If a user needs to read a paragraph to understand what a card represents, your labels are too complex. Test your labels with a pilot session before launching the full study.
Avoiding identical words across cards
One of the most damaging mistakes in card sorting is using identical or very similar words across multiple cards. If you have cards labeled "Toyota Camry," "Toyota Corolla," and "Toyota RAV4," participants will group them mechanically by the word "Toyota" rather than thinking about deeper relationships like vehicle size or purpose. The groupings reflect label similarity, not conceptual relationships. That defeats the purpose.
When items share terminology, vary the phrasing to emphasize distinguishing attributes. Instead of three cards all starting with "Toyota," use "Camry (Midsize Sedan)," "Corolla (Compact Sedan)," and "RAV4 (Compact SUV)" so participants consider the actual differences.
Moderated vs. unmoderated sessions
Moderated card sorting, where a researcher observes and asks participants to think aloud, provides richer qualitative data. You learn not just where participants place cards, but why. You hear their reasoning, their uncertainty, the questions they ask themselves while sorting. That context is invaluable for understanding the cognitive process behind their groupings.
Unmoderated remote card sorting is faster, cheaper, and easier to scale. You can collect data from fifty participants in days rather than weeks. But you lose qualitative insights. You see the resulting groupings, not the reasoning behind them. For most projects, a hybrid approach works: run moderated sessions with five to eight participants to understand the why, then validate those findings with an unmoderated study of thirty to fifty participants for statistical confidence.
How Many Participants Do You Actually Need?
The sample size question comes up in every IA project. The answer depends on whether you are running a qualitative or quantitative card sort. For qualitative open card sorts where you are exploring mental models and gathering insights, NN Group research shows that fifteen participants is typically sufficient to identify the major patterns. Beyond fifteen, you see diminishing returns. The same patterns repeat. Additional participants provide confirmation rather than new discoveries.
For quantitative closed card sorts where you need statistical confidence to justify a major navigation restructuring to stakeholders, you need more participants. A minimum of thirty is standard. Fifty or more provides stronger reliability. The number to watch is not total participants but the agreement rate for each card. If seventy percent of participants place a card in the same category, you can be confident that placement aligns with the dominant mental model. If agreement falls below fifty percent, the category label or card placement needs reevaluation.
One practical consideration: the participants in your card sort must represent your actual user base, not a convenience sample of coworkers or friends. Internal team members have been exposed to your product's terminology and mental model. Their sorting patterns will not reflect what real users would do. Recruit from your actual audience, even if it takes more time. The data quality difference is dramatic.
Analyzing Card Sort Results: From Raw Data to Actionable IA
Analysis is where card sorting either delivers value or becomes an expensive exercise in confirmation bias. The data can be analyzed at multiple levels, each revealing different insights about user mental models.
Similarity matrices
A similarity matrix shows how often two cards were placed in the same group across all participants. If eighty percent of participants placed "Password Reset" and "Account Locked" in the same category, that is a strong signal that users see them as closely related. Similarity matrices help you identify the strongest conceptual relationships in your content, which should inform your top-level category groupings.
Dendrograms and cluster analysis
Cluster analysis builds a hierarchical tree showing how cards group together at different levels of similarity. The dendrogram visualization makes it easy to see where natural breakpoints occur in users' categorization. If the data shows a clear cluster of ten cards that most participants grouped together, those cards likely belong in the same top-level category.
I find cluster analysis especially useful for deciding how many top-level categories your navigation should have. If participants naturally create five clusters from your fifty cards, that suggests a five-category top-level navigation. If they create twelve, your content may require a more granular structure or deeper subcategories.
Category label analysis
In open card sorts, participants name their own categories. Analyzing those labels reveals the vocabulary users naturally reach for when describing groups of content. If most participants label a group of cards "Account Settings" rather than "Profile Management" or "User Preferences," that is strong evidence for using "Account Settings" as your navigation label. The language users choose is almost always better than the language your team invents.
Do not just look at the most common label. Look at the variety. If participants propose five different names for the same group of cards, that indicates your content grouping may be unclear, or that the items span multiple conceptual categories that need separation.
Five Common Card Sorting Mistakes and How to Avoid Them
I have made most of these mistakes myself. Here are the ones that cause the most damage, along with specific strategies to avoid them.
Mistake 1: Too many or too few cards
I mentioned the 30-to-50 card sweet spot already. The mistake I see most often is running a sort with only 15 cards because the team was in a hurry. With too few cards, participants create only a handful of categories, and you miss the nuanced relationships that make card sorting valuable. With too many cards, participants fatigue halfway through and start sorting randomly. If you need to cover a large content inventory, run multiple focused card sorts instead of trying to fit everything into one.
Mistake 2: Leading the participant
In moderated sessions, it is tempting to offer hints when a participant seems uncertain. Do not. The whole point is to see how they organize the information without your influence. If you help them, you are seeing your own mental model, not theirs. Sit silently. Let them struggle. That struggle is data.
Mistake 3: Ignoring the unsortable pile
Almost every card sort produces an "I do not know where this goes" pile. Some researchers treat this as noise and exclude those cards from analysis. That is a mistake. Items that participants cannot categorize are critically important. They indicate content that does not fit users' existing mental models. Those items may need restructuring, renaming, or different positioning. They might also indicate genuinely new concepts that require user education rather than intuitive placement. Whatever the case, do not discard them. Investigate why they were unsortable.
Mistake 4: Reading too much into small patterns
If three out of thirty participants grouped a particular set of cards together, that is not a pattern. It is noise. Focus on groupings that appear in at least fifty percent of participants. The strongest IA decisions come from the most consistent patterns. Individual variation is normal and interesting, but it should not drive structural decisions unless it reveals a genuine segmentation in your user base. For example, if new users consistently organize differently than power users, that is worth noting.
Mistake 5: Treating card sort results as final architecture
Card sorting tells you how users naturally categorize your content. It does not tell you whether the resulting navigation will actually work for real-world findability tasks. Users may create beautiful, logical categories in a card sort and then fail to find anything in a tree test because the category labels they chose create ambiguous information scent. Card sorting is a generative method. It produces hypotheses about navigation structure. Those hypotheses must be tested with tree testing before you finalize anything.
What Is Tree Testing?
Tree testing is a task-based research method that evaluates how easily users can find specific items within a proposed navigation hierarchy. Unlike card sorting, which asks users to create structure, tree testing asks users to navigate an existing structure and measures how successful they are at finding what they need.
The method presents participants with a text-based hierarchical menu: the "tree." No visual design, no color, no images, no layout. Just category and subcategory labels arranged in expandable accordion menus. Participants get tasks like "You want to reset your password. Where would you click?" and their clicks are tracked as they navigate through the tree to find the answer.
Because tree testing removes visual design entirely, it isolates the performance of your IA from the influence of good graphic design. You cannot hide behind beautiful buttons or clever iconography. If your IA is fundamentally broken, tree testing will reveal it without mercy.

Designing a Tree Test That Measures Findability
A tree test requires two things: the tree itself and the tasks participants will perform. Both must be carefully designed to produce reliable data.
Building the tree
Your tree should represent the full navigation hierarchy you intend to test, including all levels of subcategories. Do not limit it to top-level navigation only. Users need to see the full depth to make meaningful choices. If a top-level category has five subcategories and those have sub-subcategories, include them all.
Format your tree in a spreadsheet with each category on its own row. The top cell represents the homepage. Each subsequent column represents a deeper level in the hierarchy. Most tree testing tools accept this format directly and generate the clickable accordion menu automatically.
The number of items at each level matters. If a top-level category contains fifteen subcategories, users will struggle to scan through them regardless of how intuitive the labels are. Research consistently shows that humans can effectively scan around five to nine items before cognitive load increases dramatically. If any level of your tree exceeds nine items, consider reorganizing or introducing subcategories.
Writing effective tasks
Each task should ask participants to find something specific within the tree. The wording must avoid giving away the answer. If you ask "Where would you find the password reset option?" the word "password" might lead participants directly to a "Password & Security" category they would not have considered otherwise. Instead, describe the scenario: "You forgot your login credentials and need to regain access to your account. Where would you look?"
I recommend writing three types of tasks for a comprehensive tree test:
- Core tasks: items representing the most common user goals for your product. These should have clear, unambiguous correct answers within your tree.
- Sleight-of-hand tasks: items that are less common but test the information scent of specific category labels. For example, asking participants to find "information about charitable donation matching" on an intranet tests whether the label "HR," "Benefits," or "Culture" creates better scent for that content.
- New or controversial items: items placed based on stakeholder preference rather than user data. These are the ones most likely to fail, and tree testing provides objective evidence to challenge subjective decisions.
Include a warm-up task at the beginning to help participants understand the interface. A simple task like "Find the homepage" ensures participants are oriented before the actual test begins. If a participant fails the warm-up, you can exclude their data as likely careless responding.
Selecting the right metrics
The primary metrics in tree testing are success rate, directness, and time on task. Success rate measures whether the participant found the correct location. Directness measures whether they took a straight path or wandered through irrelevant categories first. Time on task measures how long they spent, which correlates with confidence and cognitive effort.
Directness is often more revealing than success rate. A participant who finds the correct item after clicking through four wrong categories has technically succeeded, but their wandering indicates that the information scent at each decision point was misleading. If multiple participants take the same wrong path before finding the correct answer, that wrong path reveals a structural problem in your labeling or hierarchy.
Card Sorting vs. Tree Testing: How They Work Together
I often hear designers ask which is better, card sorting or tree testing. The question misunderstands the relationship. They are complementary, not competing. Card sorting generates hypotheses about how to structure your IA. Tree testing validates whether that structure actually works.
Think of it this way: card sorting asks "How would you organize this content?" Tree testing asks "Can you find this content in the organization we created?" Both questions are essential. The first ensures your navigation aligns with user mental models. The second ensures that alignment translates into real findability.
The ideal sequence: open card sort to understand mental models and generate category ideas. Draft an IA based on those findings. Validate that draft with a closed card sort or tree test. Iterate based on the results. Tree test again to confirm improvements. Then design the visual interface around the validated structure. This prevents the most common IA mistake: designing a beautiful navigation that nobody can use.
In practice, running a single round of card sorting followed by two rounds of tree testing produces a dramatically better IA than either method alone. The card sort gives you the raw material. The first tree test shows where your initial structure breaks down. The second confirms that your fixes resolved those issues. This pattern rarely requires more than a week of research time and can save months of post-launch navigation fixes.
Analyzing Tree Test Results: Success Rates, Time, and Path Data
Tree testing tools provide a wealth of data. Knowing which numbers to focus on is the difference between actionable insights and analysis paralysis.
Success rate benchmarks
NN Group research suggests that a success rate above 80 percent for core tasks is reasonable for well-designed navigation. Below 70 percent indicates significant problems. If your most critical tasks fall below 70 percent, do not proceed with development until the IA is fixed.
These benchmarks depend on context. A complex enterprise product with thousands of features will have lower expected success rates than a simple e-commerce site. Compare against your own baseline and against similar categories within the same study, not against generic benchmarks.
First-click data
The path a participant takes through the tree is often more informative than their ultimate success or failure. Pay attention to first-click data: the first category they expand after reading the task. If the majority of participants expand the wrong top-level category on their first click, your top-level labels are creating the wrong information scent, even if participants eventually correct their path and find the answer.
I have seen tree test results where success rates looked acceptable at 85 percent, but first-click data showed that seventy percent of participants initially clicked the wrong top-level category. Those participants succeeded only because they had enough patience to keep exploring. Real users on a live site would not be so forgiving. They would bounce. First-click data often reveals problems that success rate alone hides.
Time on task as a proxy for cognitive load
Participants who take a long time to find an item are experiencing cognitive friction, even if they ultimately succeed. Long task times suggest that users are uncertain at each decision point, which indicates that your category labels are ambiguous or that your hierarchy does not match their mental model. If average task times for core tasks exceed thirty seconds, the IA needs refinement.
Be careful about comparing time across different task types. Finding "Contact Us" should be nearly instantaneous. Finding "Advanced Security Certificate Settings" will naturally take longer because it is more specialized. Compare time within similar task categories, not across them.
The PIR (Piercing Index Ratio)
Some tree testing tools provide a metric called the Piercing Index, or PIR, which measures how quickly and directly participants pierced through the tree to find the correct answer. A high PIR indicates strong information scent: users moved decisively toward the right location. A low PIR indicates that users had to explore broadly before narrowing in, suggesting weak or misleading labels at intermediate levels.
Integrating IA Research Into Your Design Process
The biggest barrier to better IA is not a lack of available research methods. It is the habit of treating IA as a design exercise rather than a research-driven discipline. Here is how to integrate card sorting and tree testing into a real product development workflow without adding weeks to your timeline.
Phase 1: Discovery (Days 1-3)
Run an open card sort with fifteen participants from your target audience. Use OptimalSort or a similar tool to keep it fast. Analyze the similarity matrix and cluster analysis to identify natural groupings. Draft a proposed IA with top-level categories based on the most consistent patterns. Identify three to five items that were frequently unsortable or placed inconsistently. These need special attention in the next phase.
Phase 2: Validation (Days 4-6)
Build a tree test in Treejack or UserZoom based on your proposed IA. Write eight to twelve tasks covering core user goals, edge-case items, and the problematic items from the card sort. Recruit thirty to fifty participants. Let the test run for 48 hours while you work on other aspects of the product.
Phase 3: Iteration (Days 7-8)
Analyze the tree test results. Identify every task with a success rate below 70 percent. Look for first-click patterns that reveal misleading labels. Revise the IA based on the data: rename categories, move items, restructure the hierarchy where necessary. Run a second tree test with the revised IA to confirm improvements.
Phase 4: Design (Days 9+)
Once the IA is validated through tree testing, move into visual design with confidence. The navigation structure has been tested and proven to support findability. Use the category labels from your research, not invented labels from a branding exercise. Design around the validated structure rather than forcing the structure to fit a visual concept.
This four-phase process takes approximately two weeks and requires minimal resources: a tree testing tool subscription and a participant recruitment budget. The return on that investment is measured in reduced support tickets, improved conversion rates, and users who can actually find what they need.

The Limitations You Should Not Ignore
Card sorting and tree testing are powerful methods, but they have limitations that every UX researcher should understand before relying on them too heavily.
Card sorting reveals how users categorize information in a controlled setting where they are actively thinking about organization. In the real world, users do not browse websites with the explicit goal of categorizing content. They search, they scan, they click around. The mental models revealed in card sorting are genuine, but they may not capture how users behave under time pressure, distraction, or emotional states like frustration.
Tree testing isolates navigation from visual design, which is useful for diagnosis but does not account for the many factors that influence real-world findability. Visual hierarchy, color coding, iconography, typographic contrast, and even search bar placement all affect how users navigate. A tree-validated IA can still fail in a real interface if the visual design is poor.
Neither method accounts for search behavior well. Users increasingly rely on site search rather than browsing navigation. A structure that scores poorly in tree testing might work fine if users find what they need through search. But relying on search as a crutch for bad navigation creates its own problems. Search results depend on content labeling and metadata quality, and users who cannot browse effectively will struggle to discover features they did not know existed.
Both methods also assume the content items included in the study are the right items. If your card sort or tree test is missing significant content categories, the resulting IA will be incomplete regardless of how well it tests. Content audits must precede IA research, and the content inventory must be comprehensive for the research to be meaningful.
Finally, neither method tells you about the emotional or experiential quality of navigation. A user might find what they need but feel frustrated, confused, or unwelcome during the process. Quantitative success rates do not capture that. Pair your IA research with usability testing that captures emotional responses, satisfaction ratings, and open-ended feedback.
References
- Nielsen Norman Group: Card Sorting: Uncover Users' Mental Models for Better Information Architecture
- Nielsen Norman Group: Tree Testing: Fast, Iterative Evaluation of Menu Labels and Categories
- Nielsen Norman Group: Card Sorting: How Many Users to Test
- Nielsen Norman Group: Information Architecture and Sitemaps
- Nielsen Norman Group: Mental Models
- Nielsen Norman Group: Information Scent
- Smashing Magazine: Card Sorting and Tree Testing: A Practical Guide
- UX Collective: Information Architecture Articles and Resources
- UX Booth: Complete Beginner's Guide to Information Architecture
- Laws of UX: Cognitive Psychology and UX Laws
This article was brought to you by Timothy Graf | GrafWeb via GrafWeb CUSO (grafwebcuso.com): design theory and practical UX research guidance for product designers, UX practitioners, and digital product teams.
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Card Sorting and Tree Testing for Information Architecture: A Practical UX Research Guide to Building Navigation Structures Users Can Actually Understand",
"description": "A practical guide to card sorting and tree testing, the two most important UX research methods for building information architecture that users can actually navigate.",
"author": {
"@type": "Person",
"name": "Timothy Graf"
},
"publisher": {
"@type": "Organization",
"name": "Timothy Graf | GrafWeb",
"url": "https://timgraf.com"
},
"datePublished": "2026-07-08",
"dateModified": "2026-07-08",
"wordCount": 4688,
"inLanguage": "en-US",
"isAccessibleForFree": true
}
