Laya as a Local Reranker for Internal Search and RAG

Key takeaways

Similarity search finds passages that look related. It does not find passages that actually answer the question. Here is how to add a local Laya reranking stage, how to measure whether it helped, and where it fits in SEO work.

There is a specific failure mode in search and retrieval that everyone hits and few people diagnose correctly.

You search your own content library for "how to reduce churn in a subscription business." The top result is a page about subscription pricing. It mentions churn. It is topically adjacent. It is not what you asked for.

The similarity model did its job. It found text that is close in vector space. What it could not do is judge whether the passage actually answers the question, because that is a different kind of judgment than "these two things are about similar topics."

That gap is what reranking fills, and it is a good fit for a local decision model. This article covers how to build it, how to measure it, and where it belongs in SEO work.

What you will finish with

  • A two-stage retrieval pipeline: cheap similarity search, then Laya relevance judgment
  • A relevance question written as a score rather than a yes/no
  • A recall-at-k measurement before and after reranking
  • A clear picture of which SEO tasks benefit and which do not

Who this is for: anyone building internal search, a content retrieval tool, or a RAG system over their own content.

Prerequisites:

  • Laya installed and running locally
  • A corpus of passages with an existing similarity or embedding search
  • A set of twenty to fifty real queries with known good answers

Definition of done: you can show that reranking improved recall at your chosen k, or you can show that it did not and stop.

Why similarity is not relevance

It is worth being precise about the difference, because the whole technique depends on it.

Similarity asks: how close are these two pieces of text in meaning or vocabulary?

Relevance asks: does this passage answer this question?

A passage can be highly similar and completely irrelevant. A page about subscription pricing is similar to a query about reducing churn, because they share vocabulary and topic. It does not answer the question.

The reverse also happens. A passage can be lexically distant and highly relevant. If someone asks "why do customers leave" and your best page says "the top driver of cancellations is onboarding friction," a keyword model may miss it entirely while a human immediately sees it as the answer.

This is why two-stage retrieval exists. The first stage optimizes for recall: get a broad set of plausible candidates cheaply. The second stage optimizes for precision: judge which of those candidates actually answer the question. The first stage is a similarity problem. The second stage is a relevance problem, and relevance is a decision.

The two-stage pattern

Code
query
  -> similarity search (cheap, broad, high recall)
  -> top N candidates
  -> Laya relevance judgment (cheap, precise)
  -> reranked results

The design principle is the same one that runs through this whole series: put the cheap filter first, and the judgment second.

Stage one is your existing search. Embeddings, TF-IDF, BM25, whatever you already have. Retrieve more candidates than you need, because stage two will cut them down. Twenty to fifty is a reasonable range.

Stage two is Laya. For each candidate, ask how well it answers the query. Then sort by the answer.

The reason this is affordable is that stage one is free and stage two is local. If you were paying per call for stage two, reranking fifty candidates per query would get expensive fast. Locally, it is a rounding error.

Write the relevance question as a score

This is the design decision that determines whether reranking works.

Do not ask a yes/no question. "Does this passage answer the query?" collapses too much. Almost everything in a top-fifty candidate set is sort of related, and a binary judgment forces the model to draw a line that you cannot inspect.

Ask for a score instead.

Code
How well does this passage answer the query?

0 — not related to the query at all
1 — same topic, but does not answer the question
2 — partially answers the question
3 — answers the question, but not completely or not directly
4 — directly and completely answers the question

Four things to notice about this scale.

Level 1 is the important one. "Same topic, but does not answer the question" is exactly the failure mode similarity search produces. Giving it its own level means the model can separate topical adjacency from actual answers, which is the whole point.

The levels are ordered and described. Each level has a concrete meaning, not just a number. Vague scales produce vague scores.

Four levels, not ten. More granularity does not mean more precision. It means more noise, because the model cannot reliably distinguish level 6 from level 7.

You can threshold it. After reranking, you can drop everything below 2, or show everything but sort by score. The scale gives you that choice.

Diagram showing a query going through similarity search to produce fifty candidates, then through a Laya relevance score, then sorted into a reranked list.

Stage one optimizes recall. Stage two optimizes precision. They are different jobs.

Build it

Step 1: Capture your baseline

Before you change anything, measure what you have.

Take your twenty to fifty test queries. For each one, record which passages a human would call correct answers. Then run your existing similarity search and record where those correct answers appear in the ranking.

The metric you want is recall at k: of the passages a human would call correct, how many appear in the top k results?

Measure it at a few values of k. Recall at 5, at 10, and at 20. You will use these as your baseline.

Step 2: Retrieve more candidates than you need

Change your retrieval to return more results than you currently show. If you display ten, retrieve fifty.

This costs nothing, because stage one is cheap, and it gives stage two something to work with. Reranking cannot recover a relevant passage that stage one never returned.

Step 3: Score every candidate

For each query, send every candidate passage through the relevance question.

json
{
  "query": "how to reduce churn in a subscription business",
  "passage": "The top driver of cancellations is onboarding friction..."
}

Store the score for every candidate. Do not filter yet.

Step 4: Sort and cut

Sort candidates by score, descending. Then decide your cutoff.

Two options:

Threshold. Drop everything below 2. Simple, but you may drop relevant results on queries where everything scores low.

Top k. Keep the top ten regardless of score. Consistent result count, but you may surface weak results on queries with no good answer.

A practical hybrid: keep everything at 3 or above, then fill up to k with the highest-scoring 2s. That guarantees a result set without padding it with level 1 noise.

Step 5: Re-measure

Run the same recall-at-k measurement on the reranked results. Compare to your baseline.

This is the only number that tells you whether the work was worth doing.

Measure it properly

A reranker that "feels better" is not a result. Here is how to get a number.

Use real queries. Not queries you invented to test the system. Pull them from your actual search logs or your sales team's questions.

Label the correct answers by hand. For each query, list which passages are genuinely correct. This is the tedious part and it is not optional.

Measure recall at k, not average score. The average relevance score will almost always go up after reranking, because you are sorting by that score. That tells you nothing. Recall at k tells you whether the correct answers moved into view.

Look at the failures. Pull the queries where reranking made things worse. If a correct answer dropped out of the top k, understand why. Usually the passage answers the question in a way the score criteria did not anticipate.

A worked example, with invented numbers:

Metric

Baseline

After reranking

Recall at 5

0.52

0.71

Recall at 10

0.68

0.84

Recall at 20

0.81

0.89

The gains are largest at small k, which is exactly where you want them. Reranking is most valuable when you are showing few results.

If recall at 20 barely changes, your reranker is reordering results without finding anything new. That is still useful for the top of the list, but it is a smaller win than the numbers above suggest.

Where this helps SEO work

Reranking is not only for customer-facing search. Four SEO applications.

Content audits

When you are deciding whether an existing page already covers a topic, you are doing retrieval. Reranking makes the "does a page already answer this" judgment more reliable, which directly improves keep-or-merge decisions.

The candidate-narrowing step in internal linking is retrieval. You find passages that might link to a page, then judge whether the link makes sense. A reranker improves the quality of the candidate set before the link judgment runs, which means fewer wasted decisions downstream.

Support content matching

If your support content and your marketing content live in different systems, reranking across both surfaces the page that actually answers the question rather than the page that shares vocabulary with it.

RAG systems over your own content

If you have built anything that answers questions from your own documentation, reranking is the single highest-leverage improvement you can make to retrieval quality. It is also the cheapest, because the model is local.

Panel showing recall at 5, 10, and 20 before and after reranking, with the largest gain at k equals 5.

Recall at k, before and after. Illustrative numbers only.

Where it does not help

When your corpus is small. If you have fifty pages total, a human can look at all of them. Reranking adds machinery for a problem you do not have.

When your stage-one recall is bad. Reranking cannot recover what retrieval never returned. If your correct answers are not in the top fifty, fix retrieval first.

When your queries are navigational. If someone is looking for a specific page by name, similarity search already handles it. Reranking adds latency for no gain.

When you have not labeled a test set. Without hand-labeled correct answers, you cannot tell whether reranking helped. Do not ship it on vibes.

Maintain it

Re-measure when your content changes. A reranker tuned on last quarter's corpus may not hold after a large content migration.

Re-measure when the model changes. A new checkpoint is a new model, and your score thresholds were set against the old one.

Grow your test set. Twenty queries is enough to start. Every time you find a failure in production, add it to the test set. After a few months you have a regression suite.

Watch the score distribution. If most candidates start coming back as 2, your criteria have drifted or your corpus has changed. Either is worth investigating.

What to do next

Take twenty real queries and hand-label the correct answers. Measure your current recall at 5, 10, and 20. Then add the Laya reranking stage and measure again.

If recall at 5 improves meaningfully, you have a reranker worth keeping. If it does not, check your stage-one recall first, because reranking cannot fix a retrieval problem.

And if the improvement is real, the same pattern applies to every other retrieval-shaped task in your SEO work: content audits, internal link candidates, and any system that answers questions from your own content.

Read the rest of the series

This article is part of a thirteen-part series on using Laya for SEO and GEO work.

Author: Julian Mercer, 14-Year Technical SEO Practitioner at Auspia. Julian writes about crawlability, site architecture, internal linking, and the technical foundations that make content readable to both search engines and AI systems.

Explore this topic

Keep following the same growth thread