Contents67
Meaning is found again through words. A controlled vocabulary decides which ones.
Controlled vocabularies are the oldest working answer to synonyms, homographs and vague terms in search. This page states what is settled about them, what the usual treatment leaves out, and what follows for anybody who wants to be found and cited correctly.
Every heading is a question. The answer stands directly under it, in plain words.
What is subject indexing?
Subject indexing assigns terms from a controlled vocabulary to documents so that they can be retrieved by topic. It describes what a document is about, independent of the words the document happens to use. It is the step that makes a collection searchable by meaning.
What are subject headings?
Subject headings are standardised terms that name topic areas in an indexing scheme. They are often pre-coordinated, combining several concepts in one heading. They give every document on a subject the same entry point.
What is a thesaurus here?
A thesaurus in information science lists terms together with their equivalent, broader, narrower and related terms. It makes the relationships between concepts explicit. A search can follow those relationships instead of guessing at words.
What is a taxonomy?
A taxonomy is a knowledge organisation system that arranges terms into categories, usually as a hierarchy. It tells a system and a reader where a concept belongs. Most site navigation is a taxonomy, whether it was designed as one or not.
What is a preferred term?
A preferred term is the authorised label chosen for a concept in a controlled vocabulary. Each concept receives one such label, and each preferred term stands for one concept. That one-to-one relation removes the ambiguity of free language.
What is a natural language vocabulary?
A natural language vocabulary is the unrestricted language people actually use, with all its synonyms, homographs and words of several meanings. It is how searchers type and how authors write. A controlled vocabulary exists because of it.
What is a homograph?
A homograph is a word spelled the same as another word with a different meaning, such as bank. Controlled vocabularies separate homographs with qualifiers. Without that separation, a search for one meaning returns the other.
How are synonyms handled?
Synonyms are different words for the same concept. A controlled vocabulary maps all of them to one preferred term. A search for any of the words then finds every document on the concept.
What is a polyseme?
A polyseme is a word with several related meanings, such as head. Indexing with a controlled vocabulary assigns the intended meaning explicitly. Free text leaves the choice to whoever reads it, human or machine.
What is user warrant?
User warrant is the principle that terms are chosen according to the words searchers are most likely to use. It brings the vocabulary close to the search box. It is the part of vocabulary design closest to search intent.
What is literary warrant?
Literary warrant is the principle that terms are taken from the terminology actually used in the literature of a field. E. Wyndham Hulme formulated it in 1911. It anchors the vocabulary in studies and books, the lexical level on which a discipline records what it knows.
What is structural warrant?
Structural warrant is the principle that a term is admitted because the structure of the vocabulary needs it. It keeps hierarchies complete and consistent. A missing level makes neighbouring concepts impossible to place.
What is pre-coordination?
Pre-coordination combines several concepts into one heading when the index is made, such as a subject with a place and a period. The indexer decides the combination in advance. The searcher has to find the combination as written.
What is post-coordination?
Post-coordination keeps concepts as single terms and combines them at search time. The searcher builds the combination. Most web search works this way, with the combination left entirely to the person.
What is a broader term?
A broader term is the more general parent of a concept in a thesaurus hierarchy. It lets a search widen when the specific term finds too little. The relationship is marked BT in thesaurus notation.
What is a narrower term?
A narrower term is a more specific child concept in a thesaurus hierarchy. It lets a search become more exact. The relationship is marked NT.
What is a related term?
A related term is linked to a concept by association without being its parent or child. It points a searcher to neighbouring ideas. The relationship is marked RT.
What is precision in this context?
Precision is the share of retrieved documents that are relevant to the search. A controlled vocabulary raises it by removing wrong meanings of the same word. It measures what came back.
What is recall in this context?
Recall is the share of all relevant documents that a search actually retrieves. A controlled vocabulary raises it by gathering synonyms under one term. It measures what was missed.
What is faceted classification?
Faceted classification describes an item along several independent aspects, such as material, period and use. S. R. Ranganathan developed the approach in his Colon Classification. Each aspect can be searched and combined on its own.
What role does metadata play?
Metadata describes a resource, and controlled vocabularies supply the values for its subject fields. Consistent values make resources discoverable across a collection. Structured data on the web is a descendant of the same idea.
Who were search intermediaries?
Search intermediaries were specialist librarians who searched databases on behalf of the people who needed the information. They translated a request into the vocabulary of the system. Their interview before the search was a clarification of the user intent.
Who are indexers?
Indexers are trained professionals who analyse a document and assign controlled terms to it. Their judgement decides under which concepts a document can be found. Two indexers do not always decide alike.
What does the usual treatment of controlled vocabulary leave out?
The standards, the machines and the people who maintain vocabularies today. The usual treatment explains terms, relationships and warrants as they worked in the library. It leaves out the international standards for exchange, the mapping between vocabularies, the web formats in which vocabularies are published, the automatic methods that extract and link terms, and the professions that build and keep them.
What is ISO 25964?
ISO 25964 is the international standard for thesauri and their interoperability with other vocabularies, published in two parts in 2011 and 2013. It defines how thesauri are built and how they map to other schemes. A thesaurus built without it is hard to connect to anything else.
What is ANSI/NISO Z39.19?
ANSI/NISO Z39.19 is the United States standard for the construction of monolingual controlled vocabularies. It sets rules for term form, relationships and display. It is the practical rulebook for anyone building a vocabulary in one language.
What is a scope note?
A scope note states how a term is to be used and where its meaning ends. It separates a preferred term from its neighbours. Without it, two indexers apply the same term to different things.
What is a non-preferred term?
A non-preferred term is a word that is not used for indexing and points to the preferred term instead. It captures the words people actually type. It is how a vocabulary meets the searcher halfway.
If one preferred term were enough, the first row would be close to 100 %. For 50 common objects, even fifteen accepted words per object reached 60 to 80 % of the people asked.
| Words accepted per object | People whose own word is accepted |
|---|---|
| One word chosen by a designer | 12 % |
| The single most popular word | 26 to 28 % |
| Three words chosen by designers | 28 % |
| The three most popular words | 42 to 48 % |
| Fifteen optimally chosen words | 60 to 80 % |
Furnas, G. W., Landauer, T. K., Gomez, L. M. and Dumais, S. T., The vocabulary problem in human-system communication, Communications of the ACM 30(11), 1987. Common objects data: 50 objects, 337 college students. Ranges are the low and high estimates given in the paper.
What is the USE FOR relationship?
USE FOR is the thesaurus relationship that lists the non-preferred terms a preferred term replaces, the reverse of USE. The two together make equivalence visible from both sides. A vocabulary with only one direction loses searchers on the other.
What is a node label?
A node label is a heading in a thesaurus display that groups terms without being used for indexing itself. It explains why terms sit together. It organises the view without adding a concept.
What is a top term?
A top term is the broadest concept at the root of a hierarchy. It gives a vocabulary its entry points for browsing from the general to the specific. It is marked TT.
What is a polyhierarchy?
A polyhierarchy allows a concept to have more than one broader term. Many real concepts belong to several contexts at once. Forcing them under a single parent hides them from half their searchers.
What is a monohierarchy?
A monohierarchy allows each concept exactly one broader term, a strict tree. It is simple to display and maintain. It is also the reason many concepts end up in only one place.
What is semantic warrant?
Semantic warrant justifies a term by the meaning and conceptual structure of the domain itself. Clare Beghtol described the family of warrants in 1986. It asks whether a term is right for the concept, beyond whether people or texts use it.
What is cultural warrant?
Cultural warrant is the principle that a vocabulary should reflect the culture and values of the community it serves, a concept developed by Clare Beghtol. It exposes terms that are outdated or biased. Vocabularies inherit the prejudices of the catalogues they came from.
What is inter-indexer consistency?
Inter-indexer consistency measures how far different indexers assign the same terms to the same document. Studies since the 1960s have found it far from complete. Every indexed collection therefore carries the variation of the people who indexed it.
If trained indexers applied a controlled vocabulary identically, every row would read 100 %. On the same articles, two MEDLINE indexers agreed on fewer than half of all main headings.
| Terms compared | Consistency between two indexers |
|---|---|
| Checktags, such as human, male, female | 74.7 % |
| Main headings for the central concepts | 61.1 % |
| All main headings | 48.2 % |
| Main headings with their subheadings | 33.8 % |
Funk, M. E. and Reid, C. A., Indexing consistency in MEDLINE, Bulletin of the Medical Library Association 71(2), 1983. 760 articles from 42 journal issues (1974 to 1980) that the National Library of Medicine indexed twice, consistency by Hooper's measure.
What is vocabulary mapping?
Vocabulary mapping links the concepts of one controlled vocabulary to those of another. It lets a search in one system find material described in another. Without it, every vocabulary is an island.
What is vocabulary alignment?
Vocabulary alignment brings separate subject schemes into agreement at the level of their structure and meaning. It goes further than single mappings. Federated search across several collections depends on it.
What is term ambiguity?
Term ambiguity is the degree to which a term can be read in more than one way. It can be assessed, and it differs by domain and context. It is the measurable form of the problem a controlled vocabulary exists to solve.
What is the Simple Knowledge Organization System?
The Simple Knowledge Organization System, SKOS, is the W3C standard for publishing thesauri, taxonomies and subject headings on the web, a Recommendation since 2009. It turns a vocabulary into linked data. A vocabulary in SKOS can be used by any system that reads the web.
What is a lead-in term?
A lead-in term is an entry point that leads a searcher from the word they used to the preferred term. It works like a signpost in the index. On the web, it is the redirect from the query to the concept.
What is concept drift?
Concept drift is the change of a concept's meaning over time. A term that was precise twenty years ago can mean something else today. A vocabulary that is never revised indexes the past.
What is ontology matching?
Ontology matching finds correspondences between the concepts of different ontologies, usually with automatic methods. It connects knowledge models that were built separately. It is vocabulary mapping at machine scale.
What is entity linking?
Entity linking connects a mention in a text to the entry for that entity in a knowledge base. It turns a name into an identified thing. It is subject indexing performed by a machine on every sentence.
What is named entity recognition?
Named entity recognition finds and classifies names of people, organisations, places and other entities in text. It is the first step before linking them. It extracts from running text what an indexer would pick out by hand.
What is cross-language information retrieval?
Cross-language information retrieval finds documents in one language for a query in another. Multilingual vocabularies and mappings make it possible. A monolingual index hides everything written elsewhere.
What is automatic term extraction?
Automatic term extraction identifies candidate terms from a body of text by statistical and linguistic methods. It supports the building and updating of vocabularies. It finds the words, and a person still decides which ones become terms.
What is knowledge graph embedding?
Knowledge graph embedding represents the entities and relations of a graph as vectors that machine learning models can use. It carries the structure of a vocabulary into numerical form. The relations become computable.
What is faceted search?
Faceted search offers the facets of a classification as filters in a search interface. The searcher narrows results step by step along independent aspects. It is faceted classification made visible to the user.
What is semantic interoperability?
Semantic interoperability is the ability of systems to exchange data with its meaning intact. Shared vocabularies and mappings make it possible. Two systems can exchange records and still misunderstand each other without it.
What is a vocabulary server?
A vocabulary server publishes a controlled vocabulary through a web interface so that applications can query it. Terms and updates are fetched when needed. The vocabulary becomes a service instead of a file.
What is PoolParty Semantic Suite?
PoolParty Semantic Suite is commercial software for building and managing taxonomies, thesauri and knowledge graphs based on SKOS. Organisations use it to maintain vocabularies at enterprise scale. It is one of the established tools for the work.
What is VocBench?
VocBench is an open-source web platform for editing SKOS vocabularies and ontologies collaboratively. It was developed with the Food and Agriculture Organization of the United Nations. It is the open counterpart to commercial taxonomy software.
What does a taxonomist do?
A taxonomist designs and maintains taxonomies and controlled vocabularies for an organisation. The role joins information science with content and product work. Many organisations need the work without having named the role.
What does an ontologist do?
An ontologist builds formal models of a domain, with concepts, relations and rules that machines can reason over. The role sits between knowledge organisation and software engineering. It is the modern successor of the classification specialist.
What role does a subject matter expert play?
A subject matter expert checks whether terms and relationships are right for the field. Warrant from users and texts needs that check. A vocabulary validated only by its builders carries their blind spots.
What is a bounded context?
A bounded context is a concept from domain-driven design, described by Eric Evans in 2003, for the boundary within which a model and its terms have one meaning. The same word can mean different things in two contexts. It is the software engineer's version of a scope note.
What is ubiquitous language?
Ubiquitous language is the shared vocabulary that a software team and the domain experts use consistently, another concept from domain-driven design. It keeps code and conversation aligned. Teams that build one are building a controlled vocabulary under another name.
What is CDISC Controlled Terminology?
CDISC Controlled Terminology is the set of standard terms used in clinical research data submitted to regulators. It makes trial data comparable across studies. It is a controlled vocabulary with regulatory force.
What is SNOMED CT?
SNOMED CT is a comprehensive clinical terminology used in electronic health records worldwide. It codes diagnoses, findings and procedures as concepts with defined relationships. It is one of the largest controlled vocabularies in daily use.
What are Medical Subject Headings?
Medical Subject Headings, MeSH, is the controlled vocabulary of the United States National Library of Medicine, used to index the biomedical literature in MEDLINE. It is the best-known example of a subject heading system at large scale. Every search in PubMed runs partly on it.
What is AGROVOC?
AGROVOC is the multilingual thesaurus of the Food and Agriculture Organization of the United Nations, covering agriculture, food and related fields. It is published as linked data in many languages. It shows how one vocabulary can serve a whole domain across languages.
How does algorithmic bias enter a vocabulary?
Algorithmic bias here is the systematic distortion that terms and hierarchies carry into automated retrieval. Old vocabularies can encode outdated or prejudiced views, and automatic indexing reproduces them at scale. Precision and recall do not measure it.
What is zero-shot classification?
Zero-shot classification assigns categories to text that a model was never explicitly trained on, using the meaning of the category labels. It lets language models index against a vocabulary directly. The quality of the labels decides the quality of the result.
Why does a controlled vocabulary matter for language models?
Language models answer more reliably when they are given exact and correct source text, and most of all for rare knowledge. In one study, retrieval of source passages cut hallucinated dialogue responses from 68.2 % to 7.9 % (Kurt Shuster and colleagues, Findings of EMNLP 2021), and keyword search found evidence for rare names more often than a meaning-based retriever (Christopher Sciavolino and colleagues, EMNLP 2021). A controlled vocabulary keeps that exact text findable, although no vocabulary guarantees a correct answer.
What does this mean for SEO?
Semantic SEO works on meaning, and meaning is found again through words. A page that uses the preferred terms of its field, names their synonyms and states their scope gives search engines and language models an unambiguous hold. The vocabulary of the studies and books of a field is the lexical level on which that hold is built.
