Contents4
Sufficient consensus
Keyphrase extraction identifies the terms that represent a document, using statistical measures, graph methods or supervised models. It supports indexing, tagging and topic assignment.
C-value and similar measures favour repeated multi-word terms.
What the result pages leave out
These belong to the subject. People ask about them. They are missing from the agreed coverage.
- The subject you avoided naming. Writers often circle a term for style. The system reads the circling as absence.
- Extraction rewards repetition. A clear text that says the thing once can be classified as being about something else entirely.
- Terms you do not want. A page can be assigned to a neighbouring subject and compete in a market you never entered.
- The unnamed new thing. A concept without an established phrase cannot be extracted, so material about it is filed under the nearest familiar label.
- Alt text and captions. Terms that appear only inside images are invisible to extraction, which quietly removes them from the subject.
What people actually want to know
- What does a machine think my page is about?
- Is my central term actually written on the page, in those words?
- Which neighbouring subject am I being filed under?
- What do I call the thing I do, and what do my customers call it?
Name the subject plainly and early. Elegance that avoids the word costs the reading.
