Contents86
A screen shows ten results. A speaker reads one. There is no second place in a spoken answer.
Voice Search Optimization is the work of being the sentence the assistant says. This page states what is settled about it, what the settled advice leaves out, and the questions that arrive once customers start asking out loud.
Every heading here is a question. The answer stands directly under it, in plain words.
What is voice search optimization?
Voice search optimization is the practice of structuring content so that voice assistants and answer engines select it and read it aloud. The unit of success is a single spoken response rather than a placement in a list. Everything that is not selected is not heard, which makes the discipline unusually unforgiving.
What are conversational keywords?
Conversational keywords are the phrases people use when they speak rather than type. Spoken queries are longer, contain question words, and carry their intent openly: somebody types plumber Frankfurt and says who can fix a burst pipe near me tonight. Writing for the spoken form means writing the whole question the way a person would actually say it.
How do featured snippets relate to voice search?
Featured snippets are the extracted answers that assistants read out. The box on the screen and the sentence from the speaker are frequently the same text, selected by the same mechanism. Work done to win the box is largely the same work that wins the spoken answer.
Why does local SEO matter so much here?
A large share of spoken questions are about where, when and nearby: opening hours, directions, availability. Local SEO aligns your content with those geographic questions and keeps the underlying business data correct. A business whose hours are wrong in the places assistants read from will be described incorrectly out loud, confidently.
What role do FAQ pages play?
An FAQ page organizes concise answers to question-shaped queries in a form that is quick to extract. Its structure matches what the assistant needs: one question, one self-contained answer. That is why it remains the single most productive page type for spoken search.
What do digital assistants actually do?
A digital assistant parses a spoken question, decides what is being asked, retrieves candidate material and reads out the piece it selected. Three separate systems are involved before any of your words are spoken. A failure at any of them means silence about you.
What does schema markup contribute?
Schema markup supplies structured metadata that lets an engine categorize and extract page content without inferring it from prose. For spoken answers this matters more than on screen, because the assistant has no layout to fall back on. Facts that are labelled can be read out; facts embedded in a sentence may not be found at all.
What is different about voice search queries?
Spoken queries have natural language structure, specific question words, and more words than typed ones. They are also messier, with restarts and filler. The gap between how a query is spoken and how a page is written is where most spoken searches fail to find their answer.
What is the role of the smart speaker?
A smart speaker is the hardware that delivers the spoken answer, usually without any screen. There is no list to scan, no snippet to skim, and no way to compare sources. The device converts a ranking into a single verdict.
What is position zero in a spoken context?
Position zero is the featured answer block that voice optimization targets, because it is the block assistants read from. On screen it is prominent. Spoken, it is everything, since being second means not being said.
What does the common advice about voice search leave out?
The advice covers question keywords, FAQ pages, schema markup and local listings. All of it concerns the text on your page. It says nothing about how speech is captured, parsed, repaired, managed across turns or rendered back as audio, which is where the majority of spoken searches are actually decided.
What is slot filling, and why does it change how you write?
Slot filling is how a system turns a spoken request into structured parameters: a service, a location, a time, a quantity. Commercial advice treats a spoken query as a string of words to match. Written for slots, a page states each parameter explicitly and separately, so the system can fill them; a paragraph that bundles service, area and hours into one flowing sentence supplies none of them cleanly. This is the single largest practical difference between writing for a screen and writing for a voice system.
What is a task-oriented dialogue system, and why is one turn not the whole story?
A task-oriented dialogue system is built to complete something across several turns: book, order, reschedule, check. Marketing discussion reduces voice to a single informational question and an answer. Real spoken interactions ask a follow-up, correct a detail, and change their mind, and content that only answers the opening question drops out at the second turn. Writing for it means covering the sequence a request actually takes, including the corrections.
What is Speech Synthesis Markup Language, and why does it exist?
Speech Synthesis Markup Language is the standard for controlling how text is spoken: pauses, emphasis, how a name or a number is pronounced. Answer engine literature discusses Schema.org for visual engines and ignores the markup built for speech. The consequence is that product names, abbreviations and prices get mispronounced or run together, which damages a brand in a way no screen equivalent does.
What is acoustic model adaptation, and why does it affect you at all?
Acoustic model adaptation is the tuning of speech recognition to particular voices, accents and devices. Voice optimization resources analyze query text and treat recognition as solved. It is not: a brand name that is reliably misheard never reaches the retrieval stage, no matter how well the page is written. Knowing how your own name survives recognition is a basic check that almost nobody runs.
What is speech disfluency resolution?
Speech disfluency resolution is the repair of the restarts, fillers and self-corrections in real speech: um, I mean, no wait, the other one. Optimization material assumes spoken queries arrive in clean syntax. They do not, and the system's repair of them determines what query your page is actually matched against. Writing for the tidy version of a question is writing for a query nobody asks.
What is auditory display, and why is a page layout no help?
Auditory display is the study of presenting information through sound: ordering, chunking, what a listener can hold without a screen. Optimization articles assume a visual result layout. A list of seven items is readable and unlistenable, and a table read aloud is noise. Content intended to be heard has to be structured for a listener with no ability to scan back.
What is dynamic knowledge graph completion?
Dynamic knowledge graph completion is the inference of relationships that were never stated explicitly, by reasoning over the relations that were. Public sources treat entity relationships as fixed things written on pages. Answer engines infer, which means a business described consistently enough can be connected to questions it never answered directly, and one described loosely gets connected to nothing.
What is speaker verification, and why does it matter commercially?
Speaker verification is the recognition of who is speaking from the voice itself. Guidance treats every spoken query as anonymous and identical. Assistants increasingly know which household member is asking, which changes what gets recommended and what counts as a repeat customer. Personalization of spoken answers is already happening and is entirely absent from the advice.
What is multimodal dialogue management?
Multimodal dialogue management handles an interaction that moves between voice, screen and touch on the same device. Optimizers treat voice as a separate channel with its own page. Real devices answer aloud and display a card at the same time, and content that works in only one of the two modes is half-delivered.
What is a speech intelligibility metric?
A speech intelligibility metric measures how clearly synthesized audio is understood by a listener. Voice search guides never evaluate the audio at all. An answer that is technically selected and acoustically unclear fails at the last step, and nobody measures that step because nobody listens to the output.
What is ambient noise suppression, and why is it your problem?
Ambient noise suppression is the preprocessing that separates a spoken query from a kitchen, a car or a shop floor. Documentation implicitly assumes a silent room. Spoken queries about local services are asked in exactly the noisy places where recognition degrades, which is why short, distinctive, phonetically robust names and phrases carry further than clever ones.
What are voice user interface guidelines?
Voice user interface guidelines are the ergonomic standards for spoken interaction: how long an answer may be, how a system confirms, how it recovers from being misunderstood. Optimization content focuses on ranking tactics and ignores them. They matter because an answer that violates them gets abandoned mid-sentence by the listener, whatever its ranking.
Which pairs of these subjects are already treated together?
- Entity
- Voice Search Optimization
- Attribute
- established pairing
- Value
- conversational keywords
- Entity
- Voice Search Optimization
- Attribute
- established pairing
- Value
- featured snippets
- Entity
- Voice Search Optimization
- Attribute
- established pairing
- Value
- local SEO
- Entity
- Voice Search Optimization
- Attribute
- established pairing
- Value
- FAQ pages
- Entity
- featured snippets
- Attribute
- established pairing
- Value
- position zero
- Entity
- schema markup
- Attribute
- established pairing
- Value
- voice search queries
Six connections are established. Voice search optimization with conversational keywords, since spoken phrasing is the obvious starting point. Voice search optimization with featured snippets, because the box supplies the spoken answer. Voice search optimization with local SEO, since so many spoken questions are local. Voice search optimization with FAQ pages, the most productive format. Featured snippets with position zero, the answer and the place it occupies. And schema markup with voice search queries, the labels and the spoken request they serve. That is the settled ground. Everything below is a connection nobody has written.
What happens when conversational keywords are joined to slot filling?
- Entity
- conversational keywords
- Attribute
- missing pairing
- Value
- Slot Filling
- Approach
- Integration of Slot Filling Methods in Conversational Keyword Optimization Strategy
Conversational keywords describe how a question sounds. Slot filling describes what a system extracts from it. Keyword work produces long phrases that read naturally and still yield no parameters. Joining the two means writing so that each value a system needs to fill sits in its own clearly marked place: the service named plainly, the area named plainly, the hours stated as hours. A page can match the phrasing perfectly and still be unusable because nothing in it can be pulled into a slot.
What happens when position zero is joined to task-oriented dialogue systems?
- Entity
- position zero
- Attribute
- missing pairing
- Value
- Task-Oriented Dialogue System
- Approach
- Retrieval of Position Zero Results Within Task Oriented Dialogue Systems
Position zero is a single answer to a single question. A task-oriented dialogue system is working through a task across turns. The prize was defined for the first while users increasingly behave like the second. Connecting them asks what happens after your answer is read: does your content support the follow-up, the correction, the booking. Winning the first turn and having nothing for the second is a common and invisible failure.
What happens when FAQ pages are joined to Speech Synthesis Markup Language?
- Entity
- FAQ pages
- Attribute
- missing pairing
- Value
- Speech Synthesis Markup Language
- Approach
- Formatting FAQ Pages Content Using Speech Synthesis Markup Language Standards
FAQ pages are optimized as text. Speech synthesis markup controls how that text is spoken. The two live in separate worlds, so the most extracted content on most sites is never prepared for the medium it is extracted into. Formatting the answers for speech means writing numbers, units, names and abbreviations in the form they should be heard, and giving the assistant the pauses that make a two-clause answer comprehensible.
What happens when smart speakers are joined to ambient noise suppression?
- Entity
- smart speaker
- Attribute
- missing pairing
- Value
- Ambient Noise Suppression
- Approach
- Impact of Ambient Noise Suppression Efficiency on Smart Speaker Query Interpretation
Smart speakers are discussed as a distribution channel. Noise suppression is discussed in signal processing. Between them sits the fact that a query asked over a running tap is a different query by the time it is recognized. Bringing the two together changes naming decisions: what survives noisy capture is short, phonetically distinct and free of near-homophones, and this can be tested rather than guessed.
What happens when featured snippets are joined to dynamic knowledge graph completion?
- Entity
- featured snippets
- Attribute
- missing pairing
- Value
- Dynamic Knowledge Graph Completion
- Approach
- Connecting Featured Snippets Extraction to Dynamic Knowledge Graph Completion Models
Snippet extraction takes a passage from a page. Knowledge graph completion infers facts that no page states. As extraction gives way to inference, being the page that states something matters less than being an entity described consistently enough to be reasoned about. Writing that connection means describing your entities the same way everywhere, so the inferences drawn about you are the ones you would have written yourself.
What happens when voice search queries are joined to speech disfluency resolution?
- Entity
- voice search queries
- Attribute
- missing pairing
- Value
- Speech Disfluency Resolution
- Approach
- Processing Voice Search Queries Through Advanced Speech Disfluency Resolution Pipelines
Queries are studied as clean text. Disfluency resolution is how the messy original became that text. The repaired query is what your page is matched against, and repair discards material, sometimes the qualifier that mattered. Understanding this changes what you optimize for: the stable core of a spoken request rather than its full surface form, because the surface never arrives intact.
What happens when digital assistants are joined to multimodal dialogue management?
- Entity
- digital assistants
- Attribute
- missing pairing
- Value
- Multimodal Dialogue Management
- Approach
- Architectural Shift from Digital Assistants to Multimodal Dialogue Management
Assistants are treated as a voice-only channel. Multimodal management describes devices that speak and display at once and expect touch in reply. The shift from one to the other is already underway on phones and screened speakers. Content that answers aloud and has nothing to show, or shows a page that cannot be spoken, delivers half an answer on a device that offered both.
What happens when schema markup is joined to voice user interface guidelines?
- Entity
- schema markup
- Attribute
- missing pairing
- Value
- Voice User Interface Guidelines
- Approach
- Aligning Schema Markup Structured Data with Voice User Interface Guidelines
Markup describes what your content is. Voice interface guidelines describe what a spoken answer may be like: its length, its structure, how it confirms. Aligning them means marking up answers that are already the right shape to be spoken, rather than marking up paragraphs and hoping an assistant abridges them well. The alignment is straightforward and almost nobody does it, because the two standards are maintained by different communities.
What happens when acoustic model adaptation is joined to speech intelligibility?
Acoustic adaptation concerns how well your words are heard coming in. Intelligibility concerns how well the answer is understood going out. Both are audio quality questions bracketing the retrieval everybody concentrates on. Treating them together gives a complete picture of a spoken interaction, and it explains failures that look inexplicable when only the text is examined.
What happens when speaker verification is joined to local SEO?
- Entity
- Speaker Verification
- Attribute
- missing pairing
- Value
- local SEO
- Approach
- Personalizing Local SEO Content Delivery Via Speaker Verification Technologies
Local SEO targets a place. Speaker verification identifies a person. Together they allow spoken local answers to be personalized to a known individual with a history, which is a different competitive situation from anonymous local search. A returning customer recognized by voice can be routed back to you, and nothing in current local practice prepares for that.
What happens when featured snippets are joined to position zero as a retention problem?
- Entity
- featured snippets
- Attribute
- missing pairing
- Value
- position zero
- Approach
- Evaluating Featured Snippets Selection Criteria for Maintaining Position Zero Rankings
Winning the position is written about constantly. Keeping it is not. Selection criteria shift as competitors publish and as the engine re-evaluates, so a position won is a position under continuous review. The unwritten work is the maintenance schedule: which answers to re-check, how often, and what to change when one is lost.
What happens when auditory display is joined to schema markup?
- Entity
- Auditory Display
- Attribute
- missing pairing
- Value
- schema markup
- Approach
- Translating Schema Markup Attributes Into Structured Auditory Display Patterns
Markup attributes are structured for machines to read. Auditory display concerns how structure is conveyed in sound. Translating one into the other means deciding what a listener hears when a marked-up set of attributes is spoken: which fields, in what order, with what grouping. Nobody writes this, which is why marked-up content so often produces a spoken answer that lists everything and communicates nothing.
What makes companies look at this at all?
Nine situations bring it onto the table. A competitor becomes the spoken answer on smart speakers for local service questions. More than forty percent of the audience asks questions aloud on mobile rather than typing. Click-through falls because assistants read answers out without sending anyone onward. A marketing review finds that the site never triggers a featured snippet for spoken queries. An adviser recommends optimizing the FAQ for natural language. A new product launches and assistants give wrong details about it. Accessibility rules require the public portal to work with screen readers and voice interaction. Smart speaker adoption reaches the main demographic in the home region. And typed queries for the main services decline while spoken ones rise.
How quickly will store traffic drop if assistants answer without sending anyone to the site?
The decline follows the share of questions that can be fully answered aloud: opening hours, address and availability go first, while anything requiring comparison or judgment keeps its visits. Look at which of your entry queries are of the first kind, since that is the portion at risk. The remedy is to be the source of the spoken answer, because the alternative is a competitor being it.
How can my shop become the answer when neighbours ask their speaker for local services?
Get the business data exactly right and identical everywhere it appears, since assistants read from those records rather than from your design. Then write direct answers to the spoken local questions: what you do, where, when you are open, how quickly you can come. Short, complete, in the words a person would say.
Do we lack the writing staff to rewrite our questions into spoken phrasing?
The task is smaller than it sounds: take the questions you already answer and rewrite the headings as somebody would say them aloud. The answers underneath usually need only their first sentence tightened. It is an afternoon for the top ten questions, which carry most of the effect.
Does our technical team lack the training to set up data tags for spoken devices?
Basic markup for a business, its hours, location and services is a well-documented, small job that any web developer can do. The editorial part needs no developer at all. Do the editorial part now and the markup when somebody is available.
How much investment and staff time to update our business details so devices name us?
Correcting business records is mostly clerical: the listing, the hours, the categories, the service area, kept identical across the places that carry them. Budget hours rather than a project. It is consistently the cheapest intervention in this field and the one most often left undone.
How do we know whether summary boxes turn into customer calls if voice cannot be tracked?
Spoken interactions leave little in analytics, so the measurement has to be asked for rather than read: record how callers say they found you, with one extra question at the start of the call. Watch direct calls and direction requests, which move when spoken answers do. Attribution here is collected at the point of contact or not at all.
Why change our pages now when typed desktop search has always worked?
It still works, and its share is falling while spoken and assistant-mediated questions rise. The changes involved also improve the typed case, since a clear direct answer serves both. Nothing is given up by making the move early.
Am I wasting money on search ads as people switch to speaking?
Ads still reach the typed share, so the spend is not wasted, but it is covering a shrinking portion of the demand. The spoken portion cannot be bought the same way and is won by being the selected answer. Setting the two costs side by side usually reframes the budget conversation.
Does waiting make it harder to reclaim the top spot once rivals are the default?
Yes, because assistants reuse sources that have been used successfully, and listeners form habits about who answers. A default that has been established for months is harder to displace than an open question. The cost of entry rises quietly.
Have I gathered all the data I can about how people speak their queries?
The general patterns are known: longer, question-shaped, local, conversational. What remains unknown is how people speak about your particular services, and that comes from your own calls and enquiries rather than from further reading. Listen to a week of incoming calls and you will have better material than any keyword tool provides.
Do I already have enough answered questions on my pages to change for spoken search?
Most sites do. The content exists inside longer pages, unmarked, with the answer arriving in the third paragraph. Surfacing and re-heading it is the work, and it is far smaller than writing anything new.
Half my target group uses home speakers daily. Is my deadline already past?
The deadline passed when the behaviour became normal, and the position is still contestable because most competitors have done nothing either. Acting now is late and effective. Waiting makes it late and expensive.
If I do not decide now, am I handing clients to competitors?
A default forms regardless of whether anyone chooses it. When your material cannot be read aloud and a competitor's can, the assistant says their name and keeps saying it. Deciding nothing is how that becomes settled.
Will my options shrink if rivals lock in the preferred recommendations first?
The options remain and they get more expensive, because displacing an established source requires being clearly better rather than merely present. Early entry buys the cheap version of the same position. That is the whole of the timing argument.
Are my teams blocked on other projects until we settle how to structure spoken data?
Often yes, since the structure decision touches templates and templates touch everything downstream. Settle it in a one-page rule and the blockage lifts. Leaving it open is what quietly stalls the rest.
Should I wait for clearer data on which questions people ask aloud?
Clearer data arrives from publishing spoken-shaped answers and seeing which get used. Waiting produces more of the data you already have. The next useful fact is on the other side of a change.
Is it smarter to watch what my rival changes first?
Watching shows their result after they have secured it, which is the least useful moment to learn. Their material also sits on their domain with their history, so it transfers poorly. Testing your own is faster and more applicable.
Should I pause until the sales team agrees on the exact customer wording?
The sales team already has the wording in the calls they take every day, and a transcript of ten calls settles the question faster than a meeting. Agreement arrives quickly when there is evidence to agree about. Waiting for consensus without it is how this stalls.
Should we wait for our own assistant tool to finish testing before launching?
Your assistant tool and your published answers are separate things, and the second does not depend on the first. Publish the answers now, since they serve the platforms customers already use. The tool can arrive later without holding anything up.
Should we postpone rewriting our answer blocks until the redesign next month?
Text changes and design changes are independent, and the text carries the effect. Doing it now means the redesign inherits the improvement rather than delaying it. There is no saving in waiting.
Is now the best moment to make sure people hear our details when asking aloud?
It is a good moment for a specific reason: your business records are the input, correcting them is quick, and the effect appears as soon as the records propagate. Start there rather than with the website. Most spoken answers about a local business are read from those records.
Do we already have enough everyday phrasing written down to be found?
Check whether your headings sound like something a person would say out loud. Most sites are written in a register nobody speaks: solutions, offerings, portfolios. Rewriting headings into spoken questions is the fastest available improvement.
How does rewriting three questions today build traction for spoken queries?
Three well-chosen questions cover a surprising share of spoken demand, because spoken queries concentrate on a few practical concerns. They also produce a pattern the rest of the site can follow. Small and specific beats comprehensive and postponed.
Why is acting now cheaper than waiting until competitors dominate?
Now you are filling a space; later you are displacing an incumbent that the system has already used successfully many times. The work is the same and the resistance is not. That difference is the entire cost of delay.
Can we start by updating our customer question page without buying software?
Yes, and that is the right first step. Rewriting headings as spoken questions and putting the answer in the first sentence requires only your existing editor. No purchase is involved at any point in the editorial work.
How do standardized labels help our web team work with outside partners?
Markup forces an explicit, checkable agreement about what each page describes, which replaces discussion about wording with a shared specification. Agencies and internal teams can both verify it. Most teams find that clarity more valuable than the tags themselves.
Testing shows devices never read our site aloud. Is that proof enough to act?
It is direct evidence that the content is not usable in its current form, which is the clearest signal available. Look first at whether the business records are correct, then at whether any page answers a question in its opening sentence. One of those two is almost always the cause.
Almost half our audience asks aloud. Will staying idle hurt us?
Idleness while the channel shifts means the channel is handed to whoever did move, and assistants reinforce whoever they have already used. The loss accumulates instead of arriving at once. That is what makes it easy to underestimate.
We have help screens already. Can we reword them into questions and answers now?
Yes, and it needs no hire. Take each help screen, make its heading the question a person would ask aloud, and make the first sentence the complete answer. The existing material carries the rest.
Our new products are live but devices give wrong information. Should we fix that first?
Yes, before anything else. A wrong spoken answer is worse than no answer, because it is delivered confidently and the listener has no way to check it. Correct the source records and publish plain, correct statements of the facts being misreported.
Families here use home devices to find local businesses. Is this the moment?
It is, and the work is unusually concrete: correct records, right categories, right hours, plus short spoken answers to the practical questions. Local spoken search rewards accuracy over cleverness. That makes it winnable without a budget.
Is it legally impossible to format spoken responses if our platform blocks code edits?
A platform that prevents markup does not prevent the editorial work, which carries most of the effect. Where markup is genuinely required, the question becomes whether the platform still fits the business. Neither situation is a legal obstacle, and framing it as one delays the part you can do today.
Will our site fail with assistants if our platform cannot handle multi-step exchanges?
Your site does not conduct the conversation; the assistant does, using your content. What your platform needs to do is publish clear, self-contained answers that the assistant can use at each turn. Multi-turn capability on your own site is a separate product decision.
Could being chosen as the single answer cost us the visitors we have?
For questions with a short, complete answer, yes, and that trade is largely unavoidable, since the alternative is a competitor being the answer. Where a question genuinely requires depth, the spoken answer tends to send the listener onward. Decide which of your questions are which and stop expecting visits from the first kind.
Can I only learn how users ask about new products by launching spoken support?
No. The questions people ask about a new product arrive first by phone, email and chat, and they are the same questions. Collect them there, publish the answers, and the spoken channel is served without launching anything.
Does the new accessibility requirement force an immediate upgrade?
Accessibility obligations have dates attached, and they apply to the published portal rather than to a future redesign. The overlap with spoken search is substantial: clear structure, real headings, answers that make sense read aloud. Doing the accessibility work correctly delivers most of the voice work as a side effect.
Should we delay until IT understands how audio works with screen readers?
Screen reader compatibility is mostly a matter of correct structure, which is editorial and structural rather than audio engineering. Fixing headings, labels and reading order can begin immediately. The specialist knowledge is needed for a narrower part than it first appears.
Would waiting for engines to finalize audio features save rework later?
The features keep changing and the requirement underneath has been constant: a short, correct, self-contained answer in plain language. Work at that level does not need redoing. Waiting produces the rework it was meant to avoid, by arriving late to each successive change.
Is it safer to hold back until we are sure voices will not mispronounce product names?
Silence about your product is not the safer option, since assistants will answer with whatever they find elsewhere. Test how your names are spoken, and where they fail, publish a pronunciation-friendly form alongside the official one. That is a small, concrete fix rather than a reason to wait.
What proof is there that updating short answers increases how often devices select us?
The general case is well established: extractable, self-contained answers are selected and buried ones are not. The specific proof for your business comes from changing a defined set of answers and checking the spoken results for those questions before and after. Without the before, the after proves nothing.
What user feedback can we only get by publishing spoken-friendly answers now?
Which of your answers get chosen, which get shortened, and where the assistant substitutes somebody else. None of that is observable while the pages remain unchanged. Publishing is the instrument.
How does spoken design show that we care about accessible communication?
Content that works spoken is content that works for people who cannot see the screen, and the overlap is nearly complete. Clear structure, real headings and answers that stand alone serve both. It is one piece of work with two justifications, which is unusually easy to defend internally.
How does fixing audio clarity today meet our accessibility commitment?
Audio clarity, correct structure and comprehensible answers are the substance of the commitment rather than a proxy for it. Fixing them is the fulfilment, not a step towards it. It also produces a record of what was changed and when, which is what an audit asks for.
What new ways to interact with shoppers open once assistants follow step-by-step requests?
Multi-step handling makes ordinary transactions possible by voice: checking availability, reserving, rescheduling, reordering. The content requirement is that each step be answerable on its own, in order. Businesses whose material supports the sequence become usable in a way their competitors are not.
A neighbour business already answers spoken questions. Does acting now stop the loss?
It stops the loss from widening, and recovering the questions they now own takes longer than holding the ones still open. Start with the questions where no clear answer exists yet. Contest the settled ones after.
We lose visitors daily to assistants. Won't delay cost more?
Yes, and the cost is cumulative rather than fixed, because each spoken answer given by somebody else makes the next one more likely. The arithmetic favours acting on a small set of questions immediately. Completeness can follow.
Typed lookups keep falling. Should we reshape our rules now?
A sustained decline in the channel that carries your customers is the clearest mandate available for changing how content is written. Set the rule once, in one page, and apply it going forward. Retrofitting the archive can proceed gradually behind that.
Accessibility rules are active and our site failed the screen reader check. Is that proof?
It is both proof and an obligation with a date. The fixes it requires are the same fixes that make content usable by assistants. Treat one programme of work as satisfying both, and do it in the order the obligation sets.