<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" version="2.0">
  <channel>
    <title>Blog</title>
    <link>https://www.stevesale.com/blog</link>
    <description />
    <language>en</language>
    <pubDate>Wed, 16 Sep 2026 07:55:28 GMT</pubDate>
    <dc:date>2026-09-16T07:55:28Z</dc:date>
    <dc:language>en</dc:language>
    <item>
      <title>From Keywords to Meaning: A Practical Tour of Enterprise Search Engines</title>
      <link>https://www.stevesale.com/blog/from-keywords-to-meaning-a-practical-tour-of-enterprise-search-engines</link>
      <description>&lt;div class="hs-featured-image-wrapper"&gt; 
 &lt;a href="https://www.stevesale.com/blog/from-keywords-to-meaning-a-practical-tour-of-enterprise-search-engines" title="" class="hs-featured-image-link"&gt; &lt;img src="https://www.stevesale.com/hubfs/2026-09-15%20-%20From%20Keywords%20to%20Meaning%20-%20A%20Practical%20Tour%20of%20Enterprise%20Search%20Engines.svg" alt="From Keywords to Meaning: Statistical Search, Vector Search, Hybrid Search with Semantic Reranking" class="hs-featured-image" style="width:auto !important; max-width:50%; float:left; margin:0 15px 15px 0;"&gt; &lt;/a&gt; 
&lt;/div&gt; 
&lt;p&gt;If you've worked on an enterprise search project in the last few years, you've probably noticed the vocabulary shifting under your feet. "Relevance tuning" used to mean adjusting field boosts and stemming rules. Now it means choosing embedding models, tuning ANN indexes, and deciding how to fuse rankings from two completely different retrieval paradigms.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;If you've worked on an enterprise search project in the last few years, you've probably noticed the vocabulary shifting under your feet. "Relevance tuning" used to mean adjusting field boosts and stemming rules. Now it means choosing embedding models, tuning ANN indexes, and deciding how to fuse rankings from two completely different retrieval paradigms.&lt;/p&gt; 
&lt;p&gt;This post is a tour of that landscape;&amp;nbsp;where enterprise search has been, where it's going, and why the most effective systems today rarely pick just one approach. We'll move through three stages:&lt;/p&gt; 
&lt;ol&gt; 
 &lt;li&gt;&lt;strong&gt;Statistical (lexical) search&lt;/strong&gt; The workhorse of search for many, many years&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Vector (semantic) search&lt;/strong&gt; Retrieval based on meaning, not matching words&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Hybrid search with semantic reranking&lt;/strong&gt; Combining both, the way most production systems actually work today&lt;/li&gt; 
&lt;/ol&gt;  
&lt;h2&gt;Part 1: Statistical Search:&amp;nbsp;Search Built on Word Statistics&lt;/h2&gt; 
&lt;p&gt;Statistical, or "lexical," search engines work by matching the literal terms in a query against the literal terms in documents, then ranking results using statistics about how those terms are distributed across the corpus. No understanding of meaning is involved; it's pattern matching, made smart through mathematics.&lt;/p&gt; 
&lt;h3&gt;1.1 The Inverted Index&lt;/h3&gt; 
&lt;p&gt;Almost every statistical search engine (Such as Elasticsearch, OpenSearch, Solr, classic Lucene) is built on an &lt;span style="font-weight: normal;"&gt;inverted index:&lt;/span&gt; a mapping from each term to the list of documents (and positions) in which it appears. This is what makes lexical search fast, instead of scanning every document for a query term, the engine does a direct lookup.&lt;/p&gt; 
&lt;h3&gt;1.2 TF-IDF&lt;/h3&gt; 
&lt;p&gt;&lt;strong&gt;Term Frequency–Inverse Document Frequency&lt;/strong&gt; is the foundational scoring idea. It weighs a term's importance by two competing signals:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Term Frequency (TF):&lt;/strong&gt;&amp;nbsp;how often the term appears in &lt;em&gt;this&lt;/em&gt; document (more mentions, more relevant)&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Inverse Document Frequency (IDF):&lt;/strong&gt;&amp;nbsp;how rare the term is &lt;em&gt;across the whole corpus&lt;/em&gt; (common words like "the" or "system" get discounted; rare, distinctive words get boosted)&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;A document scores highly when it contains query terms frequently, and those terms are otherwise rare in the collection. TF-IDF is simple, explainable, and still conceptually underlies most lexical scoring today.&lt;/p&gt; 
&lt;h3&gt;1.3 BM25 — The Modern Default&lt;/h3&gt; 
&lt;p&gt;&lt;strong&gt;BM25 (Best Matching 25)&lt;/strong&gt; is a refinement of TF-IDF and is the default scoring algorithm in Elasticsearch, OpenSearch, and Solr. It improves on raw TF-IDF in two important ways:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Term frequency saturation&lt;/strong&gt; — the 10th occurrence of a word matters much less than the 2nd. BM25 diminishes returns on repeated terms instead of scaling linearly.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Document length normalisation&lt;/strong&gt; — a long document naturally contains more term matches by chance. BM25 penalises this so a short, focused document isn't unfairly outranked by a long one that happens to mention the term more.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;BM25 is fast, requires no training, is highly explainable (you can show exactly why a document scored the way it did), and remains extremely competitive — especially for precise, keyword-heavy queries: product codes, legal citations, exact names, error messages.&lt;/p&gt; 
&lt;h3&gt;1.4 Other Statistical Techniques Worth Mentioning&lt;/h3&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Boolean retrieval: &lt;/strong&gt;The pre-statistical ancestor (AND/OR/NOT matching), still underlies filtering logic in modern engines&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Field boosting and query-time weighting:&lt;/strong&gt; Manually telling the engine that a title match matters more than a body match&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Synonym expansion, stemming, and lemmatisation: &lt;/strong&gt;Techniques to bridge small vocabulary gaps ("run" / "running" / "ran") without leaving the lexical paradigm&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Language models for Information Retrieval&amp;nbsp;(e.g. query likelihood models)&lt;/strong&gt;&amp;nbsp;A&amp;nbsp;more probabilistic cousin of TF-IDF/BM25, less common in enterprise tooling but conceptually important.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;h3&gt;1.5 Where Statistical Search Falls Short&lt;/h3&gt; 
&lt;p&gt;The core weakness of every technique above is that they only match &lt;em&gt;words&lt;/em&gt;, not &lt;em&gt;meaning&lt;/em&gt;. A query for "employee exit process" will not match a document titled "offboarding procedure" unless someone has manually added a synonym rule. This is the&lt;span style="font-weight: bold;"&gt; &lt;/span&gt;&lt;span style="font-weight: normal;"&gt;vocabulary mismatch problem, and it's the single biggest driver behind the move to vecto&lt;/span&gt;r search.&lt;/p&gt;  
&lt;h2&gt;Part 2: Vector Search — Search Built on Meaning&lt;/h2&gt; 
&lt;h3&gt;2.1 The Core Idea: Embeddings&lt;/h3&gt; 
&lt;p&gt;Vector search represents text (or images, audio, etc.) as &lt;strong&gt;embeddings;&lt;/strong&gt;&amp;nbsp;dense numerical vectors produced by a neural network, positioned in a high-dimensional space such that semantically similar content ends up close together, regardless of the specific words used.&lt;/p&gt; 
&lt;p&gt;"Offboarding procedure" and "employee exit process" will land near each other in embedding space, even though they share no words — because the model has learned that they &lt;em&gt;mean&lt;/em&gt; similar things.&lt;/p&gt; 
&lt;h3&gt;2.2 How It Works, End to End&lt;/h3&gt; 
&lt;ol&gt; 
 &lt;li&gt;&lt;strong&gt;Chunking:&lt;/strong&gt; The document is first split into a series of passages small enough to be individually retrieved and fed to the model, but large enough to retain context. Chunking strategy is one of the most underrated variables in vector search quality: too small and you lose context,&amp;nbsp;too large and you dilute relevance and blow through token budgets.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Embedding generation: &lt;/strong&gt;An embedding model converts each document chunk into a vector.&amp;nbsp; These&amp;nbsp;typically have anything from 300 – 3500&amp;nbsp;dimensions.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Indexing: &lt;/strong&gt;Vectors are stored in a vector index, usually using an &lt;strong&gt;Approximate Nearest Neighbour (ANN)&lt;/strong&gt; algorithm such as &lt;strong&gt;HNSW&lt;/strong&gt; (Hierarchical Navigable Small World graphs), which trades a small amount of accuracy for dramatic speed gains over exact nearest-neighbour search.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Query time: &lt;/strong&gt;The user's query is embedded using the same model, and the engine finds the &lt;em&gt;k&lt;/em&gt; nearest vectors by a distance metric,&amp;nbsp;usually &lt;strong&gt;cosine similarity&lt;/strong&gt; or &lt;strong&gt;dot product&lt;/strong&gt;.&lt;/li&gt; 
&lt;/ol&gt; 
&lt;h3&gt;2.3 Why It Matters for Enterprise Search&lt;/h3&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Handles paraphrase and vocabulary mismatch: &lt;/strong&gt;The exact problem lexical search struggles with.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Cross-lingual retrieval: &lt;/strong&gt;Some embedding models place equivalent content from different languages near each other.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Works well on conversational, natural-language queries: &lt;/strong&gt;Increasingly the norm as users bring chatbot-style habits to internal search.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;h3&gt;2.4 Where Vector Search Falls Short&lt;/h3&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Precision on exact terms:&lt;/strong&gt;&amp;nbsp;Vector search can be &lt;em&gt;worse&lt;/em&gt; than lexical search at matching exact product SKUs, error codes, or names, because embeddings smooth over specifics in favour of general meaning.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Explainability: &lt;/strong&gt;It's much harder to tell a stakeholder &lt;em&gt;why&lt;/em&gt; a document scored highly; there's no simple term-overlap story.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Domain drift: &lt;/strong&gt;General-purpose embedding models may not capture the nuance of a specific enterprise's jargon unless fine-tuned.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Cost and infrastructure: &lt;/strong&gt;Embedding generation and ANN indexes add real compute and storage overhead compared to an inverted index&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;This complementary failure pattern, lexical is precise but rigid, vector is flexible but fuzzy,&amp;nbsp;is exactly why the two are increasingly combined rather than treated as competitors.&lt;/p&gt;  
&lt;h2&gt;Part 3: Hybrid Search and Semantic Reranking&lt;/h2&gt; 
&lt;h3&gt;3.1 The Case for Hybrid&lt;/h3&gt; 
&lt;p&gt;Neither paradigm alone is sufficient for most enterprise use cases. A legal team searching for a specific clause number needs BM25's precision. The same team searching "what happens if a supplier breaches confidentiality" needs vector search's grasp of meaning. &lt;strong&gt;Hybrid search&lt;/strong&gt; runs both retrieval methods in parallel and merges the results.&lt;/p&gt; 
&lt;h3&gt;3.2 Fusion Techniques&lt;/h3&gt; 
&lt;p&gt;The central challenge in hybrid search is that BM25 scores and cosine-similarity scores live on completely different scales and aren't directly comparable. Two common approaches:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Reciprocal Rank Fusion (RRF): &lt;/strong&gt;Instead of combining raw scores, RRF combines &lt;em&gt;rankings&lt;/em&gt;. Each document gets a score based on its rank position in each result list (1/(k + rank)), and these are summed. This sidesteps the scale problem entirely and is simple, robust, and requires no tuning.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Weighted score combination: &lt;/strong&gt;Normalising both score sets (e.g. min-max scaling) and combining them with a tunable weight (e.g. &lt;code&gt;0.3 × BM25 + 0.7 × vector&lt;/code&gt;). More flexible but requires tuning per use case and corpus.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;h3&gt;3.3 Semantic Reranking: The Second Stage&lt;/h3&gt; 
&lt;p&gt;Hybrid retrieval typically returns a good &lt;em&gt;candidate set&lt;/em&gt; (say, the top 50–100 documents), but the ranking within that set can still be improved. This is where &lt;strong&gt;semantic rerankers&lt;/strong&gt; come in, usually &lt;strong&gt;cross-encoder models&lt;/strong&gt;.&lt;/p&gt; 
&lt;p&gt;The key architectural distinction:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Bi-encoders&lt;/strong&gt; (used in vector search) embed the query and document &lt;em&gt;separately&lt;/em&gt;, then compare vectors. Fast, but the model never actually looks at the query and document together.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Cross-encoders&lt;/strong&gt; (used in reranking) feed the query and document &lt;em&gt;into the model together&lt;/em&gt;, letting it directly attend to how they relate. This produces much more accurate relevance judgments&amp;nbsp;but is far too slow to run over an entire corpus, which is why it's reserved for reranking a small candidate set rather than initial retrieval.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;A typical modern pipeline looks like this:&lt;/p&gt; 
&lt;ol&gt; 
 &lt;li&gt;&lt;strong&gt;Retrieve:&lt;/strong&gt;&amp;nbsp;BM25 and vector search each return their top N candidates&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Fuse:&lt;/strong&gt;&amp;nbsp;RRF or weighted fusion merges the two lists into one candidate set&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Rerank:&lt;/strong&gt; A cross-encoder (e.g. Cohere Rerank, BGE-reranker, or a fine-tuned in-house model) rescoring that candidate set for final relevance&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Return:&lt;/strong&gt; The top-k reranked results are served to the user or passed to a downstream LLM (in a RAG setup)&lt;/li&gt; 
&lt;/ol&gt; 
&lt;h3&gt;3.4 Why This Architecture Wins in Practice&lt;/h3&gt; 
&lt;ul&gt; 
 &lt;li&gt;&lt;strong&gt;Recall from two different signal types: &lt;/strong&gt;You catch documents that either method alone would miss.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Precision from reranking:&lt;/strong&gt; The expensive, accurate model is only run on a small candidate set, keeping latency manageable.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Graceful degradation:&lt;/strong&gt;&amp;nbsp;If the vector index is stale or the embedding model is poorly suited to a niche domain, lexical search still provides a solid baseline.&lt;/li&gt; 
 &lt;li&gt;&lt;strong&gt;Tunability at each stage: &lt;/strong&gt;You can adjust fusion weights, reranker choice, and candidate set size independently as you learn more about your users' query patterns.&lt;/li&gt; 
&lt;/ul&gt;  
&lt;img src="https://track-eu1.hubspot.com/__ptq.gif?a=148732941&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fwww.stevesale.com%2Fblog%2Ffrom-keywords-to-meaning-a-practical-tour-of-enterprise-search-engines&amp;amp;bu=https%253A%252F%252Fwww.stevesale.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>Hybrid Search</category>
      <category>Enterprise Search</category>
      <category>Statistical Search</category>
      <category>Vector Search</category>
      <pubDate>Tue, 15 Sep 2026 15:25:51 GMT</pubDate>
      <author>consulting@stevesale.com (Steve Sale)</author>
      <guid>https://www.stevesale.com/blog/from-keywords-to-meaning-a-practical-tour-of-enterprise-search-engines</guid>
      <dc:date>2026-09-15T15:25:51Z</dc:date>
    </item>
  </channel>
</rss>
