Fine-Tuning Laya on Your Own SEO Labels

Key takeaways

Fine-tuning is the one capability that makes an open decision model genuinely different from a hosted one. Here is how to build a label set from your own review decisions, train on a laptop, and avoid the trap that inflates every published benchmark.

Everything in this series so far has used Laya as it ships. That is the right way to start, and for some tasks it is enough.

But there is a version of this workflow that a hosted model cannot offer you at all, and it is the reason to choose an open model in the first place: you can train Laya on your own decisions.

That changes the accuracy question completely. Instead of asking "how good is this model at my task," you ask "how good can I make it at my task." Those are very different questions, and the second one has a much better answer.

This article covers how to do it, how to tell whether it worked, and the specific trap that makes most published fine-tuning numbers look better than they are.

What you will finish with

  • A labeled dataset built from decisions you have already made
  • A fine-tuned checkpoint trained on your own hardware
  • An honest accuracy measurement, including a holdout set you did not train on
  • A decision about whether fine-tuning helped enough to keep

Who this is for: anyone who has run a Laya workflow, measured the accuracy, and found it not good enough to trust.

Prerequisites:

  • Laya running locally
  • A few hundred decisions you have already reviewed by hand
  • A machine that can run a training job, which in practice means a laptop is enough for a model this size
  • Patience for the measurement step, which is the part people skip

Definition of done: you have a fine-tuned checkpoint, an accuracy number measured on data it never saw, and a clear answer to whether it beat the base model.

Why fine-tuning is the whole point

Recall the accuracy picture from earlier in this series. Laya's base checkpoint scores 0.362 on the typed-decisions benchmark its own authors published, which they describe as close to random. The headline number people quote, 0.766, came from a checkpoint fine-tuned on that benchmark's training data.

That is not a scandal. It is the design. Laya is built to be a fine-tuning base, not a general-purpose classifier. The base model gives you a working architecture and a reasonable starting point. Your data gives you the accuracy.

For SEO work, this matters because your taxonomy is not anyone else's. Your intent labels, your definition of "thin," your rules for when two pages should merge: none of that is in any training set. A general model has to infer your rules from your prompt every time. A fine-tuned model has learned them.

There is also a second benefit that gets less attention. A fine-tuned model can be faster, because it needs less context to make the same decision. One developer reported fine-tuning Laya on an application-specific dataset and cutting inference latency sixfold while matching results, all on a laptop. When the model already knows your taxonomy, you do not have to explain it in every call.

Build the dataset from decisions you already made

The good news is that you have probably been generating training data for months without calling it that.

Every time you reviewed a Laya output and corrected it, you created a labeled example. Every time you hand-tagged a query, you created one. Every time you decided a page should be merged rather than updated, you created one.

Where the labels come from

  • Your review queue. The rows you sent to human review and then corrected are the highest-value examples, because they are exactly the cases the model got wrong.
  • Your historical work. Old audits, old intent tags, old keep-or-kill decisions. If you can reconstruct the input and the decision, you have a training example.
  • Deliberate labeling. Sit down and label a few hundred examples specifically for this purpose. Tedious, but it is the fastest way to get a balanced set.

How much data you need

There is no universal number, but the shape of the answer is consistent: a few hundred examples per label is a reasonable starting point for a narrow task, and more is better for wide taxonomies.

What matters more than raw count is balance. If 90% of your examples are one intent label, the model will learn to guess that label and your accuracy number will look fine while the model is useless. Count your examples per label before you train, and fix the imbalance first.

Format the data

Each example is a state and a target answer.

json
{
  "state": {
    "query": "hubspot vs salesforce for small teams"
  },
  "question": "Which family does this query belong to?",
  "target": "compare"
}

Keep the question text identical to what you use in production. If your training question and your inference question differ, you have trained a model for a task you are not running.

Expand synthetically, carefully

If you do not have enough examples, you can generate variations with a language model. Take a real labeled example, ask a model to produce paraphrases that keep the same label, and add them.

Two rules:

Never let the generator change the label. If you cannot verify that a paraphrase still belongs to the same class, do not include it.

Never generate your holdout set. Synthetic examples are variations of your real ones, so they are not independent. Keep your holdout set entirely hand-labeled and real.

Train it

The mechanics are simpler than most people expect, because the model is small.

The setup

  • Load the base checkpoint
  • Attach the decision head for your question type
  • Train on your labeled examples
  • Validate on a held-out slice

A laptop is enough. The model is in the hundreds of millions of parameters, not billions, and the training job is a classification task rather than a generative one.

The parameters that matter

Learning rate. Too high and the model forgets what it knew. Too low and it does not learn your task. Start conservative and adjust.

Epochs. More is not better. Watch your validation accuracy, not your training accuracy, and stop when validation stops improving.

Class balance. If you cannot fix the imbalance in the data, weight the loss so rare classes count more.

The holdout split. This is the one that decides whether your result is real. More on it below.

The trap: benchmark-split inflation

This is the most important section in the article, and it is the reason most published fine-tuning numbers should be read with suspicion.

When you fine-tune on a benchmark's training split and then report accuracy on that benchmark's test split, you are measuring something real but narrow: how well the model learned that specific distribution. You are not measuring whether it generalizes.

There is a specific warning sign to watch for. If your fine-tuned model scores higher than the teacher model that generated your labels, something is wrong. The teacher's self-consistency is a ceiling on what its labels can teach. Exceeding it usually means the model memorized patterns in your training data, including the noise, rather than learning the underlying task.

In published Laya results, the fine-tuned checkpoint exceeded the teacher model's self-consistency ceiling. That does not make the result fake. It does mean the number should not be read as "better than the teacher," because that is not what it demonstrates.

How to avoid the trap

Hold out real examples. Your test set must be examples the model never saw, drawn from the same distribution as your production data, and labeled by a human.

Split before you augment. If you generate synthetic examples, split first, then augment only the training portion. Otherwise your paraphrases leak into your test set and your accuracy number becomes meaningless.

Report the baseline. Always report the base checkpoint's accuracy on the same holdout set. A fine-tuned number without a baseline tells you nothing about whether fine-tuning helped.

Test on a second distribution if you can. If your model works on last month's queries but not this month's, you have overfit to a moment in time.

Measure honestly

Here is the measurement procedure that produces a number you can actually use.

  1. Split your labeled data. Reserve 20% as a holdout. Do this before any augmentation.
  2. Measure the base model on the holdout. This is your baseline.
  3. Fine-tune on the training portion only.
  4. Measure the fine-tuned model on the same holdout.
  5. Compare. If the fine-tuned model is not clearly better, fine-tuning did not help.
  6. Look at the per-label breakdown. An aggregate improvement can hide a regression on a specific label, which matters if that label is the one you care about.

A worked example, with invented numbers:

Label

Base accuracy

Fine-tuned accuracy

learn

0.71

0.88

compare

0.64

0.86

buy

0.58

0.81

navigate

0.92

0.94

problem

0.49

0.79

Overall

0.67

0.86

The overall number improved, but the interesting part is problem, which went from 0.49 to 0.79. That is the label the base model could not handle, and it is exactly the case fine-tuning is for.

Diagram contrasting the benchmark-split trap with the correct holdout measurement approach.

The benchmark-split trap, and the honest alternative.

What to do when fine-tuning does not help

Sometimes it does not. Here is how to diagnose it.

Check the data before the model. Nine times out of ten, a fine-tune that does not improve is a labeling problem. If two people labeled the same examples differently, the model cannot learn a consistent rule.

Check for label leakage. If your state contains a field that correlates with the label in a way that will not hold in production, the model will learn the shortcut. A common version: including the URL when the URL itself encodes the category.

Check the question text. If your training question and your inference question differ, you have trained for the wrong task.

Check the class balance. A model trained on imbalanced data will guess the majority class and look accurate on aggregate.

Accept that the task may be too hard. Some judgments are genuinely ambiguous, and no amount of training will make them consistent. If your own human labelers disagree with each other on 30% of cases, that is your accuracy ceiling, and it is not a model problem.

Panel showing per-label accuracy before and after fine-tuning, with the problem label showing the largest gain.

Per-label accuracy before and after. Illustrative numbers only.

Maintain the model

A fine-tuned checkpoint is a new artifact you now own.

Version it. Keep the base checkpoint, the training data, and the fine-tuned weights together. In six months you will need to know what produced what.

Re-measure after every retrain. A new checkpoint is a new model, and your confidence threshold from the previous version does not transfer.

Retrain when your taxonomy changes. If you add an intent label, the old model does not know it. Retrain rather than hoping the prompt will cover it.

Watch for drift. Your query mix changes over time. A model trained on last year's queries may quietly degrade on this year's. Re-measure periodically against fresh hand-labeled examples.

What to do next

Take the review queue from your existing workflow and turn it into a labeled dataset. You probably already have a few hundred examples without realizing it.

Then run the honest measurement: base model on the holdout, fine-tuned model on the same holdout, per-label breakdown. If the fine-tuned model is clearly better on the labels you care about, you have something worth keeping.

And if it is not, you have learned that your task is either ambiguous or already handled well enough by the base model. Both are useful answers, and both are better than shipping a model whose accuracy you never measured.

Read the rest of the series

This article is part of a thirteen-part series on using Laya for SEO and GEO work.

Author: Marcus Ellery, Growth Experimenter Behind 150+ SEO Tests at Auspia. Marcus writes about experiments, benchmarks, learning loops, and evidence-led growth content.

Explore this topic

Keep following the same growth thread