Semantic Search (Dashboard → Settings → Search, member feature, on by default for members) makes the AI chat assistant understand what a visitor means, not just the words they typed — and puts the same engine one click away in site search, as a Keywords / Meaning toggle on the search overlay and the dashboard's ⌘K palette.
A visitor asking "my lights keep flickering" has described a problem in their own words. An electrician's site says "electrical repair", "panel replacement", "wiring". The two share no word at all, so keyword search returns nothing and the visitor concludes the site can't help. The chat assistant finds the page anyway, because retrieval compares meaning rather than spelling. So does the visitor who flips site search to Meaning.
Visitor search: keyword by default, meaning on request
Site search — the /search page, the header search bar, the Cmd-K overlay — runs keyword matching (Standard LIKE, MySQL FULLTEXT, or Meilisearch, whichever the Search settings page has picked) on every request, because a visitor typing into a search box wants speed. The semantic layer is the slow engine: it has to embed the query through a provider before it can score anything.
So it is never forced on a visitor. It is offered. The search overlay (the header search icon and ⌘K) carries a Keywords / Meaning toggle in its footer; a visitor whose keyword results look wrong flips it and accepts a moment's wait. Two switches on Settings → Search → Semantic Search → Site search shape that (VisitorSearchMode):
| Switch | Setting | Default | What it does |
|---|---|---|---|
| Let visitors switch to meaning-based search | search.semantic_visitor |
on | Renders the toggle on the overlay and honours mode=semantic on /_search and /search. |
| Open the search overlay in meaning mode | search.semantic_default |
off | The overlay opens on Meaning instead of Keywords. Inert while the first switch is off — withdrawing the offer withdraws the default with it. |
Both are inert unless the layer is actually usable (feature on and a provider that can embed), and the resolution is server-side: VisitorSearchMode::resolve() collapses any request for meaning mode to keyword when it isn't offered, and PublicSearch::search() re-resolves it rather than trusting the caller — so no visitor path reaches the provider on an install whose owner withdrew it. The defaults are what every install should start with: offered, not default. Nothing to migrate; an install that has never visited the settings page offers it the moment it can embed.
What meaning mode returns. The keyword pass still runs. Meaning mode then leads with a "Matched by meaning" group — one row per page or record, ranked by its best passage, the passage as the snippet (SemanticSearch::retrieveDocuments()) — and drops those URLs from the keyword groups that follow, so nothing lists twice and switching never loses a result. Pinned results still lead everything: an editor's exact-term curation outranks an inference. On a shop, the meaning group also keeps its lead over the product groups the typeahead normally hoists, because the row cap would otherwise cut the one group the visitor switched for. When meaning finds nothing — the site doesn't answer this, or the provider is mid-outage — the keyword results are left exactly as they were and the overlay says so ("Nothing matched by meaning — showing keyword matches."), so the toggle never makes a search worse than the mode it left and never looks like it did nothing. The query is scored against the spelling keyword search actually ran: an auto-corrected typo is embedded corrected.
Where the mode travels. The overlay's choice sticks for the tab (sessionStorage), debounces longer in meaning mode (450 ms vs 200 ms — one embedding call per distinct query, cached a week), and Enter with nothing highlighted carries mode=semantic onto /search, so the results page answers in the engine the visitor chose. The /_search JSON echoes the mode it actually ran in, which is what the "nothing matched" notice reads — a downgraded request is not a miss.
The dashboard palette (GlobalSearch / CommandPalette) has the same toggle, gated on the layer being usable rather than on the visitor switch (that Setting is about the public site). Meaning mode there adds the group straight after the "Go to" shortcuts, each hit pointing at its dashboard editor with the matched passage as the subtitle — for the page an editor remembers by what it says rather than by its title. Keyword is the default and the choice sticks for the session.
Headers with a built-in search bar get the same toggle without a single header row changing. The 36 design-library headers with a bar render a plain <form action="/search"> with a q input; everything interactive under it is injected at runtime — public.js clones the panel from the git-tracked search-suggest template and binds search-suggest.js to the form's own input, which is how the typeahead and typo correction reached every materialised header. The toggle is the panel's last row, so it appears the moment a visitor has results to judge. The form still submits natively: in meaning mode the component adds a hidden mode input to the form and the parameter to the "See all results" link, so Enter lands on /search?mode=semantic. The choice is shared with the overlay through the same sessionStorage key. Shop headers that opt in with data-search-suggest-form post to the product grid, where meaning mode does not apply, so the toggle is withdrawn there whatever the owner offers.
The /search results page has its own bar, which is a /search form and therefore gets the injected panel and toggle too. Two rules make it feel native there: the opening mode is the URL's, never the remembered one (the server already rendered the page in that mode, and a toggle that disagreed with the results under it would lie), and flipping the toggle with a query on screen re-submits the form so the whole page re-renders in the other engine, not just the dropdown. The baked results row needs no change for this.
History. Until 2026-08 a meaning-based rerank finished every /search request unconditionally: keyword search decided what matched and the semantic layer reordered it. It was removed because its cost-benefit never closed — an embedding round-trip per distinct query, on the surface where the visitor wants speed, and it could never ADD a result, only reshuffle ones keyword search had already found. Meaning mode is the opposite shape on both counts: it runs only when asked for, and what it adds is the result keyword search missed. Don't reintroduce an unconditional rerank on a visitor-facing path without confronting that history.
Two properties of retrieval do the safety work the old rerank's rules used to do:
- Retrieval is additive and separate. The chat tools run keyword search untouched and append a clearly-labelled "matched by meaning" group beside it; meaning mode in site search leads with the same group and dedupes rather than reordering. A keyword hit is never reordered, demoted, or dropped by a similarity score.
- Retrieval applies a relevance floor (below), because its output feeds a model — and a "matched by meaning" group that listed the nearest unrelated page would teach visitors the toggle is noise.
What gets embedded
| Source | Notes |
|---|---|
| Pages | From page_search_entries, which is already extracted, already per-language, and already carries required_role. |
| Content items | Blog, services, events, locations, and every custom content type. Published and listed only. |
| Documentation articles | Default version only, matching public search. |
| Products | Published only. |
Everything else is deliberately left out. Those are short, structured records where an exact filter already beats a similarity score — "3-bed under $400k" is a query for the listings filter, not for an embedding — and retrieval's one consumer, the chat assistant, wants prose passages that answer a question, which those records don't hold.
Role gating is not optional
Page copy can sit behind auth or a role gate. page_search_entries records that as required_role, and it is copied onto every embedded passage and enforced on every read.
Without it the public chat bot could quote an admin-only page to an anonymous visitor. That is a disclosure, not a ranking bug, so the default role is 0 — a caller that forgets to pass one gets the public slice, never everything.
A page that later moves behind auth keeps its text, so its content hash is unchanged and nothing would re-embed it. The indexer catches that separately and updates the stored role in place, at no API cost.
How indexing works
Indexing runs on the lazy cron (semantic:embed, every 5 minutes), so new and changed content is picked up within a few minutes of a save with no queue worker and no cron entry required.
Each pass is bounded — 40 passages by default — because the lazy cron executes inside a visitor's page request. Enumeration is lazy and abandoned the moment the budget is full, so the cost the limit exists to bound is not paid up front. Whatever is left is picked up on the next pass. Same reasoning as the Blog Agent's one-task-per-tick rule.
Staleness is a content hash and nothing else. There is no observer and no invalidation event: the sources are already kept correct by their own, so following them means an out-of-band change — a raw SQL edit, a snapshot restore, a content import — self-heals on the next pass instead of leaving a vector that confidently describes deleted copy.
Passages are chunked at roughly 900 characters with overlap, and every chunk carries its document's title. A paragraph three screens down often names none of the subject nouns ("we replace the panel and test the circuit") and is unfindable on its own; the title is the cheapest possible context.
A document scores as its best passage, never its average. Averaging would penalise a long page that answers the question in one paragraph, which is exactly the case chunking exists to handle.
Where the vectors live
In the embeddings table, as BLOBs of packed little-endian float32.
WebProCMS runs on MySQL, which has no vector column type — and laravel/ai's vector columns, whereVectorSimilarTo() and orderByVectorDistance() are all pgvector-only. So both the storage format and the distance metric are ours, and comparison happens in PHP.
That is only viable because the vectors are deliberately narrow: 256 dimensions, truncated from the model's full-width output, about a kilobyte each. A whole site is a couple of thousand passages, which PHP scores in single-digit milliseconds. Vectors are unit-normalized on write, so every comparison at read time is a bare dot product.
Vectors from different models or different widths are not comparable to each other. Changing either means a rebuild, which is what the Rebuild button on the Search settings page is for.
Where the embeddings come from
Almost always the managed AI proxy, on the house key, with no provider key on the install at all. Embeddings are a separate managed upstream from text and images — Claude has no embeddings endpoint, so a fleet brokering text through Anthropic needs a different house account for vectors, exactly as it does for pictures.
An install running its own keys falls back to its OpenAI key, then its Google key. Anthropic and DeepSeek can't embed at all, so an install whose text provider is Claude and which holds no other key will show "no provider can generate embeddings yet" and stay in keyword order — working, just not smarter.
Query embeddings are cached for a week. Scoring a search means embedding the query, which is the one place in this system a visitor waits on a provider round-trip, and real search traffic repeats itself heavily.
Retrieval for AI surfaces
SiteKnowledge::retrieve($query, $k) returns the top-k passages — not documents — with their titles and URLs, language-scoped and role-gated. It is the content counterpart to SiteKnowledge::forPrompt(): the Brief tells an AI surface who it is speaking for, and this tells it what the site actually says.
SemanticSearch::retrieveDocuments($query, $k) is the search-surface twin: the same scan, collapsed to one row per document (ranked and snippeted by its best passage, with the candidate pool scanned deeper to make up for the collapse). A long page answers a question in several adjacent chunks, and a results list that showed it three times would read as broken.
SiteKnowledge::retrievedForPrompt($query) returns the same thing formatted for a system prompt, with source URLs so the model can cite them.
Retrieval applies a relevance floor. Cosine similarity has no zero point — the nearest passage on a site that says nothing about the question still scores well above zero. On a real install, passages that genuinely answered a question scored 0.44–0.61 while the best match for a question the site never addresses scored 0.33. Handing an AI three unrelated passages is worse than handing it none, because it will answer from them and sound certain. So retrieval drops anything under 0.35 (tunable via the semantic_search.min_relevance Setting, since the useful value depends on the model).
The chat bot
The visitor-facing half of this is written up in ai-chat.md under "Search by meaning, not just keywords"; what follows is the mechanism.
search_site_content runs the plain keyword federation, then appends a "matched by meaning" group from retrieval on every call — deduped by path against the keyword results, role-gated to the current visitor, and degrading to nothing (never an error) when no embedding provider is available.
Both search tools take two inputs, because one string can't serve both engines. query stays short and keyword-shaped — the LIKE/FULLTEXT index narrows on every added word. question carries the visitor's need as one full sentence in their own words, and it is what retrieval scores: meaning-matching is at its best on exactly the phrasing keyword search is worst at ("someone tricked my employee into wiring money" finds the fraud-prevention content; the keyword "wiring" finds audio-visual pages). Before the split, the model's dutifully-compressed keyword query was also what fed retrieval, so the semantic layer never saw the sentence-shaped input it exists for. With no question passed, retrieval falls back to scoring the keyword query.
search_knowledge_base goes semantic-first, and it is the single biggest beneficiary — it was a bare LIKE '%term%' that didn't even use the search engine setting. Support questions almost never repeat the wording of the article that answers them: "it keeps logging me out" against a page titled "Session timeout" shares no term at all. Keyword remains the fallback, because it is still better for an exact lookup like an error code or a setting name.
What it does not do
Worth stating plainly, because all of these get assumed:
- It does not run on visitor search unless asked. Site search is keyword by default; meaning mode is the overlay's toggle (and the dashboard palette's), never an unconditional pass — see "Visitor search: keyword by default, meaning on request" above.
- It does not fix typos. Embeddings tolerate mild misspellings only incidentally, and fail on short queries. Typos are handled by a different layer: the database engines' spelling correction (see search.md "Typo correction"), which auto-corrects dead-end queries and offers "did you mean" links — or Meilisearch's native per-keystroke tolerance on the Smart Search add-on.
- It does not help with structured queries. "3-bed under $400k", "open after 6pm", "in stock under $50" — the existing filters are better, and those sources are deliberately not embedded.
- It cannot answer what the site never says. If the content isn't there, retrieval correctly returns nothing rather than the nearest unrelated page.
The whole feature is gated, data layer included
AI Knowledge gates its editing UI but deliberately leaves its data layer ungated, because free installs have migrated chat-bot instructions in there that must keep working.
Semantic Search does the opposite: the gate covers everything. Retrieval returns nothing and indexing does not run when the feature is off. The reason the two differ is that there is no pre-existing data here to protect — a free install has never had an embedding, so an ungated data layer would only mean buying vectors nobody can use. This was a decision, not an oversight; the ai_assistant split is the exception, not the pattern.
Operations
| Action | How |
|---|---|
| Build the index now | Search settings → Semantic Search → Build index (unlimited; the admin who clicked it is the one waiting) |
| Rebuild from scratch | Rebuild — required after an embeddings model or width change |
| From the CLI | php artisan semantic:embed --limit=0 (--flush to rebuild, --force to run with the feature off) |
| See whether it has caught up | The same card: "Still building the index…" while a pass is still working through the backlog, the provider's error when one is the reason, or "Caught up as of …". Deliberately point-in-time — content added since that pass is un-embedded until the next tick |
| Check volume | The passage / record / language counts on the same card. A count is not coverage: an install 28% through a backfill still reads a healthy-looking "1,350 passages indexed", which is why the catch-up line above exists |
| Offer meaning mode to visitors / open the overlay on it | The two Site search switches on the same card. Saved on change; each save clears the response cache, because the overlay bakes both flags into every public page |
Turning the feature off leaves the stored vectors alone; the chat assistant falls back to keyword-only matching immediately, and the Keywords / Meaning toggle disappears from the overlay and the palette.