Sufficient consensus
Topic modeling discovers latent themes across a document collection using methods such as LDA or embedding-based clustering. It is used for content audits, trend analysis and structuring large archives.
It works well and it scales.
What the result pages leave out
These belong to the subject. People ask about them. They are missing from the agreed coverage.
- Silence has no cluster. A theme that nobody wrote about produces no topic, and the report reads as complete coverage.
- Volume over importance. Frequently discussed themes dominate. A decisive subject mentioned by three sources gets absorbed into a neighbour.
- The number of topics is a choice. Set it low and distinctions vanish. Set it high and noise looks like insight. The setting is rarely reported.
- Labels are written by people. A cluster becomes a topic when someone names it, and the name carries their assumptions.
- Corpus selection decides the result. Analysing the first page of results produces the themes of the first page of results.
What people actually want to know
- What is nobody in this market writing about?
- Which subject appears in customer conversations and in no document?
- Did we choose this corpus, or did a ranking choose it for us?
- Who named these clusters?
The useful output of a content audit is the list of subjects that belong to the field and appear nowhere in it.
More in computational linguistics
Extractive summarizationThe summary is now the article for most of the audience.Keyphrase extractionIf your real subject is named once, it is not your subject as far as the system is concerned.Semantic searchThe engine got better at knowing what you meant.Named entity recognitionBeing an entity is a prerequisite for being understood.Question answeringA confident single answer is the format.Word sense disambiguationYour page competes in whichever meaning the system assigned to it.Sentiment analysisA dashboard of positive and negative tells you the weather, not the cause.