Dawn Anderson: SEO Through Information Retrieval Science

SEO Through Information Retrieval Science

Written for website owners who cannot get found. Not here: what to do before you have an offer. Traffic + Offer

Most SEO professionals learn the craft through pattern-matching. You observe what works, you repeat it, you report the results. That's pragmatic, and it produces traffic. It's also fundamentally limited, because it tells you what works without ever explaining why, it struggles the moment circumstances change, and it leaves you exposed to reading correlation as causation.

Dawn Anderson works from an entirely different foundation: Information Retrieval science, the academic discipline that studies how search engines store, index, classify and retrieve information. This isn't theory disconnected from practice. It's the science Google's engineers used to build the systems that rank pages. Understanding it means understanding how the system actually works, rather than how it looks from outside, or how people assume it works.

Who They Are

Anderson is the Managing Director of Bertey, a Manchester-based search marketing consultancy, and she has also lectured in digital marketing at Manchester Metropolitan University. That combination is unusual in SEO. Most senior practitioners don't carry academic credentials, and most academics don't work on client problems in real time. Her thinking gets validated against real systems and real constraints, rather than against a reading list.

The value of the IR foundation is that she translates abstract retrieval concepts into practical SEO implications. The translation works at more than one altitude, too. For a technical SEO, IR supplies the theory underneath the tactics. For a store owner or a B2B marketer, the practical version is simpler. A well-structured, technically sound site performs better because it aligns with how search engines are designed, not because it ticks boxes on an audit template.

What distinguishes the work is intellectual depth combined with practical application. She doesn't theorise at you. She explains how a specific IR concept shapes Google's behaviour, and then shows what that means for your site architecture, or your crawl budget, or your content structure. That builds credibility, because when she makes a claim about how Google processes language or allocates resources, she can point at the research that supports it. It isn't an assertion of opinion. It's established computer science, translated into strategy.

She's also known for challenging shallow SEO orthodoxy. When the industry settles on a certain truth about how SEO works, she has often pointed out that the academic literature suggests a different explanation, and those positions are worth taking seriously. The credibility there comes from rigour, not from being contrarian for its own sake.

What They Teach

Anderson's core teaching is how search engines actually work from a computational point of view, and crawl budget optimisation sits at the centre of it. Google doesn't crawl all of your pages with equal frequency. The crawler allocates resources based on a calculation of crawl efficiency, and that calculation rests on two factors. Host Load is the crawl rate your server can sustain before Google politely backs off. Crawl Demand is the perceived importance and freshness of your pages. Anderson teaches how the calculation works, how it shapes which pages get discovered, how often they get re-crawled, and what happens when your site structure makes crawling inefficient. Crawl budget isn't technical minutiae. It's foundational, because if Google doesn't know your pages exist, or doesn't re-crawl them when they change, performance collapses regardless of how well-optimised the content is.

The practical implications are significant. A site with poor crawl efficiency might hold 50,000 indexable pages with only 10,000 of them indexed, which means 40,000 pages are simply never discovered. You can create perfect content for those pages and it will have zero visibility, because the problem isn't the content. It's the crawl efficiency. The stakes climb with scale and churn. An ecommerce store with frequently changing inventory is the obvious case, and so is a B2B site with a sprawling resource section. Either can burn its budget on faceted navigation filters spawning near-infinite URL variations, while the pages that earn money wait to be re-crawled. Anderson teaches you to diagnose that by understanding how crawlers allocate resources, so you can identify which pages get crawled, which get ignored, and why. The diagnosis reveals opportunities that simply aren't visible without the framework.

Site architecture and indexability follow logically from crawl budget. How you structure a site determines how efficiently crawlers can navigate it, and a flat structure with everything equally accessible produces one crawl pattern while a deeply nested hierarchy produces another. Excessive internal linking, or a confusing taxonomy, produces inefficiency. Anderson teaches how to design site structure from a crawler's perspective, which isn't about what looks good to a human but about building an architecture a machine can navigate and understand efficiently. She covers pagination, faceted navigation, cross-linking strategies, and what each one does to your crawl patterns.

Her work on Natural Language Processing and semantic understanding bridges IR science and modern AI. She explains how BERT and the language models that followed it actually process text. Not as vague "understanding meaning" but as specific computational mechanisms: tokenisation, embedding, attention layers, semantic similarity. And how entities and the relationships between them get recognised and resolved. She connects that to SEO by showing how those mechanisms determine whether Google recognises your content as relevant to a query. They also decide whether it matches your page to the intent behind the query, rather than only to its words. She works the same seam as Bill Slawski, who read the patents, and Koray Tugberk Gubur, who built a methodology on top of it, while Anderson reads the academic literature underneath both. If you understand how semantic matching works, you understand what it means to optimise for relevance without keyword stuffing or content manipulation. Her teaching demystifies how Google can work out that "car", "automobile" and "motor vehicle" are semantically related without ever being told. That should reshape the writing, because the aim isn't an exact keyword match. It's writing in a way that lets a semantic model classify the content correctly.

Technical debt management is another pillar. Most sites accumulate technical problems over time: broken canonicals, duplicate content, poor mobile implementation, slow page load, crawl traps, orphaned pages. Those aren't individual small problems, they're symptoms of underlying architectural issues. Anderson teaches how to identify the root causes, how to prioritise the fixes when you can't do everything at once, and how to stop the accumulation starting again. It requires seeing your site as a system rather than as a collection of individual pages.

Her framework is systems thinking applied to SEO. Every element of a site, meaning architecture, content structure, technical implementation and crawl patterns, affects every other element. Change one thing and you have changed crawl efficiency, indexation, semantic relevance, user experience and ranking signals all at once. The job of strategy is optimising across those dimensions simultaneously, which requires understanding how they interact. The mindset she asks for is closer to information architecture than to marketing. You're designing how a body of information gets stored, connected and retrieved, and the retrieval system happens to be Google.

How It Maps to Opportunity and Authority

Anderson's work sits very high on both axes, perhaps 60 per cent Opportunity and 55 per cent Authority, because in her framework these aren't opposing forces. On the Opportunity side, crawl budget optimisation directly determines which content gets discovered at all. Site architecture determines which content combinations are eligible for search, and semantic understanding is what makes relevance matching accurate. A site with poor crawl efficiency will miss opportunities no matter how much content you produce.

The Authority component is equally strong, because technical soundness is a form of credibility. A site with proper canonicalisation, efficient crawling, semantic clarity and architectural logic signals reliability. Part of that signal is algorithmic: crawlers can trust that the site is well-maintained, that navigation is predictable, that content relationships are clear. Part of it goes to your readers, because a well-architected site is easier to navigate, loads faster, and doesn't break or redirect unexpectedly. That's Authority in the full sense, algorithmic signals and actual user experience together.

The framework connection is that Opportunity and Authority don't compete in Anderson's work. Technical excellence creates both. You don't sacrifice one to capture the other, you build a system that's technically sound, semantically clear and architecturally logical, and that system serves both goals. Which is why the work is valuable: it shows that good SEO is good systems thinking.

The Strategy Breakdown

Anderson's core strategies each pull on the two levers differently. Here's the split, so you know what you're buying when you adopt one.

Technical SEO fundamentals

Logical architecture, efficient crawlability, clean code, mobile-first implementation: engineering discipline applied to websites, with "avoiding technical debt" as the operating principle. Opportunity impact: very high. A page crawlers can't reach, whether that's through broken links, crawl traps or orphaning, has zero Opportunity no matter how good it is. Clean structure makes every relevant page discoverable, which for a large store means every product variation you actually want to sell. Diagnose with crawl logs and a crawler like Screaming Frog or Lumar, then work through technical SEO systematically. Authority impact: high. A well-architected, error-free site signals reliability to algorithms and users alike, and a site plagued by technical faults erodes trust faster than good content can rebuild it. For B2B software especially, technical stability reads as a proxy for competence, because a broken site selling reliability is a contradiction buyers notice.

Crawl budget optimisation

Manage the resources search engines allocate to crawling your site. Two things govern it: Host Load, meaning what your server can sustain, and Crawl Demand, meaning how important and fresh Google considers your pages. Opportunity impact: high, rising with site size. Efficient budget use means new and important content gets discovered and indexed promptly, so you can capture timely Opportunities. Budget wasted on unimportant or broken URLs delays the indexing of pages that matter, and unhandled faceted navigation is the classic offender. Small sites can mostly ignore this; sites with tens of thousands of pages cannot. Authority impact: moderate. Efficient crawlability contributes to site health and speed, which are positive signals. A site that manages its crawl surface demonstrates the kind of care that indirectly builds algorithmic trust. But this is a discovery lever first.

NLP and semantic search application

Structure and write content so that language models, BERT and its successors, accurately interpret meaning, entities and intent. Opportunity impact: high. Clear headings, semantic HTML, and precise, unambiguous language help search engines match your page to the right queries. This is the mechanism that makes keyword research pay off, because the topics you targeted only convert into visibility if the machine classifies your page the way you intended. A product description written in vague marketing language gets matched vaguely, and one written concretely gets matched to buyers. Authority impact: moderate to high. Semantically rich, well-structured content that handles entities and their relationships properly signals depth and expertise, which is topical authority in algorithmic form.

Information Retrieval principles

Apply the academic foundations, meaning indexing processes, relevance scoring, document analysis methods like TF-IDF and topic modelling, and distributional patterns like Zipf's Law, to practical strategy. Opportunity impact: moderate. IR knowledge is predictive rather than directly executable. It lets you reason about how a change to content or structure will affect retrieval and relevance scoring before you make it, rather than afterwards. Authority impact: high. A site built to align with how retrieval systems fundamentally work is logically structured and algorithmically legible, which is foundational Authority rather than borrowed signals. How directly something like Zipf's Law applies to SEO is debated, and the distinction is worth keeping, but the underlying principle of matching natural language patterns and logical information structure holds. I think that's why her recommendations tend to survive the algorithm updates that flatten tactic-driven sites.

When to Learn From Them

Learn from Anderson if your site has complex architecture. Complex doesn't necessarily mean broken, but complexity creates risk. Hundreds or thousands of pages, multiple category taxonomies, version variants like mobile against desktop, language variants or dated content, an international structure with regional variations. All of that is real complexity, and her framework helps you manage it without letting it quietly undermine SEO performance.

Go to her when your technical SEO has stalled. You've fixed the obvious problems, mobile, page speed, basic canonicalisation, and the rankings still aren't improving. The problem is probably deeper: crawl efficiency, semantic clarity, structural efficiency. Her diagnostic approach is built for finding those root causes.

She's the right read if you want to understand how Google actually processes and evaluates pages, rather than the surface-level advice. A lot of SEO writing is tactical, do this and avoid that. Her work explains the science underneath the tactics, which is valuable because it lets you apply the principle in a context nobody has written a checklist for. When Google makes an algorithmic change, you can reason through what changed and why, instead of simply reacting.

The alignment matters here too. If you believe understanding the science behind search leads to better outcomes than pattern-matching alone, and you're willing to put in the intellectual effort, her teaching pays dividends. If you prefer tactical checklist-driven SEO, it will feel overly theoretical, and that's a genuine mismatch rather than a failing on either side.

And she's where to start when technical debt has accumulated and you don't know where to begin. Her prioritisation framework and her systems thinking help you see the relationships between different technical problems, so you can address root causes rather than symptoms.

Where to Start

Her conference presentations are excellent entry points. Anderson is a skilled speaker, and the talks make IR concepts accessible while keeping the depth intact. Search specifically for the ones on crawl budget, site architecture and semantic understanding.

Her written work appears across the industry publications, so search for her byline. The posts tend to be longer and more technical than typical industry writing, which means they repay careful reading. Take notes. Come back to them.

Her university teaching materials, where they're available, provide structured education in IR science applied to SEO. That's rare. Most SEO education is either purely tactical or purely theoretical, and hers sits in the practical-theoretical middle ground.

Start by auditing your own site from a crawl efficiency perspective. How many pages does Google crawl each day, how frequently do the important ones get re-crawled, and is the budget going to pages that matter or to pages that don't? That assessment often reveals opportunities nobody knew existed, and once the current crawl patterns are clear, her framework helps you improve them systematically. Google Search Console shows crawl statistics. Screaming Frog shows your site's link structure. Server log analysis, or a platform like Lumar, shows what crawlers actually did, which is frequently different from what you assumed they did. Check the plumbing while you're in there: robots.txt rules, XML sitemaps that list what should be indexed, canonical tags that agree with each other. Combining those data points with her framework is how you find where crawl efficiency breaks down.

Then work backwards from your indexation goals. Count the pages you need indexed against the pages currently indexed, and the gap tells you something about crawl efficiency. If 80 per cent of your pages are indexed, that's fine; if only 40 per cent are indexed and all of them should be discoverable, crawl efficiency is your problem. Her framework helps you work out what improvement would close the gap.

If you have a specific technical SEO problem that feels intractable, search for whether she's written or spoken about it. The odds are good that she has, and the explanation will usually clarify the underlying issue rather than just the symptom.


Part of the Expert Series. Back to the framework or the diagnostic. Part of the Marketing Universe. Explore Traffic Plus Offer : The Trust Algorithm : 4-Quadrant AI. Read the book: Marketing Curious: Working the Noise.