What a Link Gap Analysis Can and Can't Tell You
A link gap or link intersect report performs one operation: it takes the referring domains of several sites, and returns the ones that link to some and not to others. It is a set difference over a vendor’s index. Everything people conclude from it — that those domains are reachable, relevant, or the reason a competitor outranks you — is inference the report did not make.
Read as a set operation with known input problems, it’s a good tool. Read as a list of missing links, it produces confidently wrong strategy.
What the report is computing
Three inputs, one operation.
The competitor set you chose. This is a judgement, and it determines everything downstream. Pick three sites that all sit in the same media ecosystem and the gap list will be that ecosystem. Pick sites that outrank you for different reasons and the intersection collapses to noise.
Each site’s referring domains, as one vendor’s crawl sees them. Every coverage and staleness problem in the underlying index propagates into the gap list. Domains missing from the crawl can’t appear in the gap; dead links still in the index will.
A threshold on “how many competitors.” “Links to at least 2 of 4” versus “links to all 4” produce very different lists, and the choice is usually a default nobody examined.
The output is a list of domains. Not opportunities, not endorsements, not causes.
The four inferences the report cannot support
That the link is obtainable. Many gap-list domains link to a competitor because of a merger, an investment, a personal relationship, a decade-old resource page, a sponsorship, or an arrangement you can’t or shouldn’t replicate. The report has no field for “why this link exists,” and that field is the only one that determines reachability.
That the domain is relevant to you. Set membership isn’t relevance. A competitor’s gap list will include their local chamber of commerce, their former employees’ blogs, and the conference they sponsored in 2019.
That the links explain the ranking difference. This is the big one. A gap list is cross-sectional: it describes two link profiles at one moment. Ranking differences have many causes, and links are one input with an unknown weight. See how to tell whether a link did anything for why isolating that contribution is not available to anyone outside the search engine.
That the count is the gap. “They have 340 domains we don’t” mixes reachable citations with structural links and index artefacts. The number is a set cardinality, and set cardinality is not a workload estimate.
What it does support, well
Discovering publications and communities you didn’t know about. This is the genuine value. If four competitors are all cited by the same trade publication and you’ve never heard of it, you’ve learned something about your market that no keyword tool would have told you.
Finding structural absences. If every competitor is listed in three industry directories and a comparison site, and you’re in none, that’s a specific, checkable, fixable observation. It’s rarely glamorous and it’s usually right.
Characterising how competitors get cited. Read the gap list qualitatively — what kinds of sites are these? Trade press? Academic? Local? User forums? That tells you what sort of thing earns citation in your space, which is more durable information than any individual row.
Prioritising within a set you already have. Given a list you were going to work through anyway, “cited by three of four competitors” is a reasonable ordering heuristic.
How to clean the list before believing it
Most gap lists are 70% unusable, and the unusable part is identifiable in a few passes.
- Drop what a vendor metric flags as near-zero authority, then spot-check twenty of the dropped rows to make sure you didn’t cut something real. Vendor scores are relative and imperfect — see Domain Rating versus Domain Authority.
- Drop scrapers and syndicators. If the linking page is a copy of a competitor’s content, the link is a byproduct of copying, not a citation.
- Drop link-farm shapes. Domains that link to everyone in the sector, with hundreds of outbound commercial links per page, are not opportunities and should not be treated as ones.
- Separate structural from editorial. Directories, review platforms, and profile pages are a different kind of work from a trade publication citing a study.
- Verify a sample by hand. Open twenty. Note why each link exists. That sample tells you the composition of the whole list, and composition is the finding.
Hypothetically: a 400-row gap list reduces to 60 after cleaning, of which the hand sample suggests roughly a third are structural listings, a third are trade or community sites where citation is plausible, and a third are relationship-dependent. That paragraph is a better deliverable than the 400-row export. (Illustrative figures.)
Reporting it honestly
Call it a set difference, not a gap. “Domains linking to at least two of these three competitors and not to us, per [vendor], as of 15 July: 412 before cleaning, 60 after.”
State the competitor set and why you chose it. The list is a function of that choice, and someone will eventually ask.
Never present the count as a target. “We need 412 links” is a sentence with no analytical content.
Don’t claim causation. “These domains cite competitors and not us” is the finding. “This is why they outrank us” is not, and the difference in credibility is large.
Add the composition sample. It’s the only part of the analysis that explains anything.
What nobody knows
Whether any of those domains would ever cite you, what fraction of the competitor’s links were earned versus arranged, and how much of the ranking difference is attributable to links at all. A gap list is a map of where citation currently sits. It is not a map of where it could sit, and it is definitely not a causal diagram.
Used as market reconnaissance it earns its cost. Used as a scoreboard it produces the most common failure in link reporting: a precise number standing in for an unanswered question.