Sufficient consensus
Passage retrieval scores sections of documents rather than whole documents, which lets long pages answer narrow questions. It is the basis of featured answers and of retrieval for generative systems.
Google ranks content pieces and excerpts.
What the result pages leave out
These belong to the subject. People ask about them. They are missing from the agreed coverage.
- Orientation inside the fragment. A reader who lands mid-page needs to know where they are, what they get and whether it is complete. Very few sections carry that.
- Headings as addresses. Each heading is a landing point for a different question. Most are written as decoration.
- Structure that survives extraction. A passage that depends on three paragraphs above it becomes wrong when it is lifted out.
- The long page that answers nothing precisely. Comprehensive pages often contain no section that is a complete answer to anything.
- What the machine quotes to a reader you never meet. Retrieved passages now feed answers that never link back. The passage is the whole relationship.
What people actually want to know
- Does each section of my page answer one question completely?
- If someone reads only this paragraph, are they misled?
- What is the address for each question I answer?
- Which of my sections would work as a spoken answer?
Every piece has to give the reader full orientation. Where am I here, what do I get here, is this true and complete.
More in information retrieval
Learning to rankThe system is trained on the behaviour it caused.Snippet generationIt is written by a machine, from your text, for a query you did not know about.Query expansionYou are answering a question that was quietly edited before it reached you.Relevance feedbackAbsence produces no signal.Query performance predictionA page of ten results looks the same whether the answer exists or not.Mean average precisionA score of 0.87 tells you how well the ranking matched a list somebody wrote down.Vector space modelIf the web never described your case, the vector for it is somebody else's.Inverted indexThe web is the part of knowledge that somebody bothered to publish and a crawler managed to reach.