Name the type of failure first
“Let us fine-tune the model on our documents” sounds like a universal fix. But the documents represent two different problems. Price lists, return policies, and shipping schedules change; the structure of a support reply, an email tone, or field extraction from an inquiry should remain consistent. Mix up those needs and you may pay to train a model that confidently quotes yesterday's price. Our little robot intern has already collected a stack of printouts, but the manager should first decide where the current source of truth lives.
The first problem usually calls for access to a current source: retrieval-augmented generation (RAG), an API call to a system of record, or both. The second may need only a clear instruction and examples; fine-tuning is justified only after a persistent behavioral error has been measured. Microsoft explicitly separates the patterns: RAG for private or frequently changing data, fine-tuning for model behavior, style, and task performance. This is a starting point, not a universal quality guarantee for a particular implementation.
What happens when knowledge changes
With RAG, the document remains outside the model's weights. The system receives the latest version, splits it into passages, indexes it, retrieves passages for a query, and gives them to the model with the question. For a product price or stock balance, a direct ERP or catalog query is often better: a search index can also lag. Amazon Bedrock's documentation notes that a knowledge base must be synchronized after files are added, changed, or removed; its incremental sync processes the changed documents. The same operational principle applies to a local deployment, even if the components differ.
Fine-tuning changes model parameters or adds an adapter. It can reinforce a desired format, terminology, and way of handling a repeated task. But changing one line in the source price list does not update the trained weights. You need fresh data, a training run, evaluation, and another release. A study published at EMNLP 2024 compared knowledge-injection methods and found a retrieval advantage in its tested settings. That is not a claim that every RAG system beats every tuned model on every business task.
A practical decision tree for an SMB
- **You need fresh facts with provenance:** return policy, product range, current contract terms. Use RAG for documents and a system-of-record API for exact operational values. The answer should carry a document version, date, and, where possible, a passage reference.
- **You need consistent behavior:** predictable JSON for inbound requests, classification of contact reasons, or a short email in a prescribed tone. Start with an instruction, a schema, and a few checked examples. Measure failures before and after.
- **You need both:** keep facts in RAG or the system of record and use an adapter only for stable formatting or domain language. This hybrid is more expensive to operate, so prove that the simpler pattern fails your acceptance criteria first.
- **You need calculations or transactions:** the model should not “remember” the final balance, discount, or warehouse availability. Verified functions and databases provide those values; the model can explain them.
The word “fine-tuning” does not replace diagnosis. Sometimes the error is not in the model but in retrieval: old and new documents sit side by side, a table was parsed into fragments, access rights were not passed into a filter, or the relevant passage never reached the context. Before training, inspect the trace: which source was selected, what passage the model saw, and which part of the answer it supports.
Data and infrastructure
A minimal RAG workflow needs a document owner, versioning rules, ingestion, an index, access control, a response model, and an error log. The index is a search copy, not the master record: deletions and updates in the primary system must reach it within a defined delay. Personal or commercially sensitive material needs the same access policy at retrieval as at the source. A prompt saying “do not reveal other people's data” is not access control.
Fine-tuning needs a properly licensed base model, a curated set of good inputs and desired outputs, a separate validation set, compute, adapter versioning, and regression tests. Hugging Face explains that LoRA and other PEFT methods reduce trainable parameters and gradient/optimizer memory, but do not make all training costs vanish: base weights and activations remain. Expert time for labeling and investigating failures remains too.
Running locally does not remove those duties. It changes the data-transfer boundary but leaves CPU/GPU, storage, upgrades, monitoring, and access control in the budget. If a sales department changes commercial terms daily, a reliable path from “approved source → update → availability check” matters more than another training epoch.
Model calculation: price the change, not just the training run
Consider a hypothetical team receiving 2,000 inquiries a month with 40 document changes. This is **a model calculation, not a company's observed result**. Assume an hour of qualified staff time costs RUB 2,000. Exclude infrastructure and API costs for both approaches for now because they vary strongly by model, load, and deployment. For RAG, allow eight hours of initial setup, then 15 minutes to validate each change and six hours a month to sample answers. That means RUB 16,000 in setup labor and RUB 32,000 in monthly labor: 40 × 0.25 × 2,000 + 6 × 2,000.
For a “retrain at every update” approach, assume even a modest two hours to prepare, train, test, and release each version, plus the same six hours of monthly review. That is RUB 172,000 in monthly labor: 40 × 2 × 2,000 + 6 × 2,000, excluding compute and initial labeling. This does not mean RAG is always cheaper. If documents rarely change, retrieval is difficult, and a behavioral skill is used millions of times, the ratio can change. The calculation shows why update frequency belongs in the budget separately from per-answer cost.
For a real decision, calculate cost per **accepted** answer: data preparation, indexing or training, inference, human review, fixes, outages, and new versions. Do not compare only token price with a GPU rental. The same 2,000 inquiries can require very different review effort and carry very different error costs.
A low-cost pilot
Take 50–80 anonymized real inquiries. Split them into questions about current facts, format/tone tasks, and system-of-record requests. For each, record the expected source, tolerable staleness, critical error, and whether a human must approve it. Test a base model with a good instruction first; add RAG or API calls only where required. If remaining errors really concern behavior and persist on an independent holdout set, estimate a small adapter and its full lifecycle.
Decide on fine-tuning only after testing new cases that were not used in training. If a document edit repeatedly triggers dataset preparation and a model release, mutable knowledge probably lives in the wrong place. Let the robot find the latest approved page instead of memorizing every new price list.
