Sufficient consensus
Query performance prediction estimates how well a system will do on a query before or after retrieval, using clarity scores, score distributions and similar signals. It is used to route queries and to trigger fallbacks.
The research field is well established.
What the result pages leave out
These belong to the subject. People ask about them. They are missing from the agreed coverage.
- The honest empty answer. A system that knows it is doing badly still returns ten results. Nothing in the interface says the corpus is thin.
- Thin topics as an opportunity. A query where every system performs badly is an open field for anyone willing to do first-hand work.
- The searcher's false confidence. A full result page reads as coverage. Most searchers take it as proof that they have seen what exists.
- Queries that are hard for a reason. Difficulty often marks a subject where practice and publication have diverged.
- Giving up too early. Most people stop searching at the point where the results start to look alike, which is long before the subject is exhausted.
What people actually want to know
- Is the answer I need actually written down anywhere?
- Why do these ten results all say the same thing?
- Which questions in my field have no good source at all?
- What would I have to do myself to answer this properly?
Most people give up the search too soon, because the results are aimed at the end of the search instead of the beginning of knowledge.
More in information retrieval
Learning to rankThe system is trained on the behaviour it caused.Snippet generationIt is written by a machine, from your text, for a query you did not know about.Query expansionYou are answering a question that was quietly edited before it reached you.Relevance feedbackAbsence produces no signal.Passage retrievalYour page is being read in fragments by a system that never sees the whole.Mean average precisionA score of 0.87 tells you how well the ranking matched a list somebody wrote down.Vector space modelIf the web never described your case, the vector for it is somebody else's.Inverted indexThe web is the part of knowledge that somebody bothered to publish and a crawler managed to reach.