Taxonomy & Site Architecture
A clear taxonomy helps users, Google and AI search engines understand your site. How to build the hierarchy, the categories and the layers that carry your keywords.
Taxonomy is the practice of classification and inheritance: dividing things into groups, then dividing the groups. Biology is the familiar example, where organisms sit in kingdoms, then phyla, then classes, all the way down to one specific animal. A website works the same way, and how well it does so decides how much of your content ever gets found.
Architecture is the part of SEO that is cheapest to get right while the site is being built and most expensive to fix afterwards. It also does double duty: the same hierarchy that lets a visitor find a product in three clicks is what tells a search engine which of your pages is the authoritative one on a subject.
The four layers
Almost every site, from a small brochure site to a catalogue with 50,000 products, resolves into the same four layers.
The homepage. Reached on the bare domain, it collects what matters and sends people onwards. It carries the most authority on the site, which means an internal link from here passes more value than one from anywhere else. Spend that deliberately.
Pages and subpages. The core static content: what the company does, the services, contact. They sit directly off the root (/services) with subpages one level deeper (/services/keyword-analysis). Depth here should follow real hierarchy rather than habit.
Category and archive pages. These list everything belonging to a category, whether that is products or articles. On a clothing retailer, /mens-shoes lists every pair of men's shoes. This layer is where most commercial search volume actually lives, and it is the one most often treated as a piece of navigation instead of as a page in its own right.
Product and article pages. The bottom layer and the most specific: one product, one article. This is where a purchase happens or a question gets answered, which makes it the destination the other three layers exist to deliver people to. Sitting at the bottom of the hierarchy is not the same as sitting at the bottom of the priority list.
The layer decides the keyword. If you want to rank for "men's shoes", the category page is the page to work on. If you want to rank for one specific model, the product page is. Trying to win a category term on a product page, or the reverse, is among the most common causes of two of your own pages competing with each other.
Categories, subcategories and tags
Categories can nest, and subcategories can nest again. A product can also belong to more than one category, which retailers rely on constantly: the same jacket sits under a product type and under a brand.
Tags are the sideways axis. Where a category says what something is, a tag captures an attribute that cuts across categories: color, size, material, season. Keeping attributes in tags instead of forcing them into the category tree is what stops the hierarchy from becoming unnavigable. Most CMS platforms ship with both, and most sites use categories carefully and tags carelessly.
Two rules keep the tag layer from becoming a liability. A tag page needs enough content to be worth indexing on its own, and it needs a reason to exist beyond internal filtering. Apply the same test to the author and date archives many platforms generate automatically. A news site genuinely needs author archives. A webshop almost never does, and indexing them produces dozens of thin pages competing for attention with the ones that sell.
Depth, crawling and click distance
A crawler discovers pages by following links. The further a page sits from the homepage in clicks, the less often it gets crawled and the less authority reaches it. Three clicks to any important page is a reasonable target. A page that takes six is effectively hidden, even though it is technically live.
That makes flat and logical better than deep and precise on almost every site. A well-organized hierarchy with clear categories and internal links between related pages improves crawlability directly, which our guide to crawling and indexing covers in more depth.
Faceted navigation is where this most often goes wrong. Filters that generate a new URL for every combination of color, size and price can turn a catalogue of a few thousand products into hundreds of thousands of near-identical pages. Decide early which filter combinations deserve an indexable URL, because they map to real search demand, and block or canonicalize the rest.
Navigation, URLs and internal links
The navigation menu is the visible form of your taxonomy, and it should read as one. Descriptive labels, no more items than a person can scan, and the same grouping logic the URL structure uses. When the menu and the URL paths disagree about how the site is organized, both users and crawlers pay for it.
URLs should mirror the hierarchy: /category/product-name, or /category/subcategory/product-name where the extra level is real. The full set of rules for paths, separators and redirects sits in URL structure. Subdomains are the other decision worth taking consciously. Putting the shop on shop.example.com and the blog on blog.example.com splits authority across separate hosts, while example.com/shop keeps it in one place for no meaningful cost.
Then link across the structure, not only down it. Category pages linking to their strongest products, products linking to related products, articles linking to the commercial pages they support. That is the work described in internal linkbuilding, and it is what turns a hierarchy into something with topical depth instead of a filing cabinet.
Metadata at scale
Every layer needs a title and a meta description, including the category, tag and archive pages that get forgotten. On a site with a manageable number of pages, write them by hand. On a catalogue with thousands of products, templates with variables are the only realistic approach, and most CMS platforms support them: one pattern for product pages that pulls in the product name, another for categories that pulls in the category name, both overridable on any individual page that deserves the attention. What to actually put in them is covered in titles and meta descriptions.
Getting it right on an existing site
Architecture is rarely a greenfield problem. Usually the structure already exists, it grew organically, and it now has three overlapping category systems and a tag layer nobody has looked at since launch. Restructuring works, but it depends entirely on the redirect mapping being right, and that is the part that determines whether the rankings survive the move. A technical SEO audit establishes the current state first, and site migration is where we handle the larger changes.
Want to know how your current structure reads to a crawler? That is one of the first things we look at in a free SEO analysis.
