Where the cost begins

When a company builds a local knowledge base from contracts, manuals, price lists and service records, the first substantial expense is often not the language model. Documents must be collected, parsed and connected to a version, an owner and access rights. If a table becomes a line of text without column headings, or OCR drops a digit, search may return a convincing but wrong passage. Making that process faster is not necessarily saving money.

Docling is one open-source option for the preparation stage. According to its documentation, the converter selects a format-specific backend and pipeline, produces a unified Docling Document and lets the user export or chunk it. Supported inputs include PDF and Office files; some older binary formats require extra components. Docling is not a complete enterprise knowledge base. Indexing, permissions, refresh policies and answer verification must be designed separately.

The benefit of local processing is control over the data path, not miraculous recognition quality. Docling documents how to prefetch model artifacts into a local directory. Sending user data to remote services requires an explicit opt-in option. Yet “local” needs to be verified across the whole chain: the OCR engine, model downloads, logs, storage, search index and final LLM, not merely the parser process.

What to measure before buying hardware

Consider a model scenario, not results observed at any real company: 1,200 documents each month, averaging four pages. A staff member currently spends four minutes per document making it searchable: checking its name, type, date and important fields, then filing it in the right area. At a fully loaded labour cost of RUB 900 per hour, this takes 80 hours or RUB 72,000 a month. It excludes the initial scanning of paper and the work of answering users' questions.

Suppose local parsing leaves 15% of documents as exceptions, each taking five minutes to resolve, and another 10% are sampled for a two-minute check. Exceptions consume 15 hours and sampling four hours. The 19 hours of labour cost RUB 17,100. In this model, we assign a further RUB 18,000 per month to the amortisation of existing equipment, operation, storage and support. The hypothetical monthly total becomes RUB 35,100, a difference of RUB 36,900 compared with the manual scenario.

If integration costs RUB 180,000, dividing that one-off expense by the assumed monthly difference gives roughly 4.9 months. This is neither a promise of Docling payback nor a market quotation for implementation. Every rate, volume, exception share and cost here is an assumption used to show the calculation. Actual numbers may reverse the conclusion. If clean, structured files already arrive and manual preparation takes only one minute, there may be little to save. If the stream consists of poor scans and complicated tables, the review queue may exceed the original workload.

Above all, do not hide the cost of errors inside “average accuracy”. In the same model, 12 incorrectly extracted documents causing RUB 5,000 of direct loss each would add RUB 60,000 of expense and wipe out all of the calculated saving. Twelve incidents is another unmeasured stress-case assumption. It illustrates why critical fields must be evaluated separately from attractive Markdown output. A wrong part number, amount, validity date or contract clause has a different price from a missing line break.

Where parsing ends and RAG begins

A workable route starts when a file enters controlled storage and receives an identifier, a version, a source and a list of authorised staff. A parser then extracts text, structure and tables. A validation stage flags failed pages and critical fields. Only then do chunks enter the index. Retrieval must filter by the user's permissions and the current document version. The answer generator receives retrieved passages and source references, rather than an unfiltered archive.

Each stage owns different failures. Docling does not decide who may read a contract; storage and retrieval enforce that. It does not guarantee that an obsolete price list has been removed from the index; a refresh and deletion mechanism is needed. It cannot establish the legal authority of extracted text; disputed amounts and conditions require a person to inspect the original. Evaluate the economics of the full “document — validation — retrieval — answer” chain, not the price of an OCR call alone.

Updates have a cost as well. Define whether a newly uploaded file is a new version, how duplicates are handled, whether stale chunks are deleted and who is accountable for a wrong source. For an initial trial, use one document class and one process owner. Indexing an entire file server immediately tends to increase volume before anyone can establish measurable quality.

Limitations visible in tests

A small, open, independent DocCrush test compared four parsers across six PDFs. Docling completed all six and performed well on financial tables and a clean scan, but poorly on a complex image-heavy report. That is a useful warning, not a universal ranking: the sample is small, and parts of the structure scoring were AI-assisted and still await human review. Its results cannot be assumed to apply to Russian contracts, delivery notes or scans without testing local documents.

Check OCR languages, source scan quality, handwriting, rotated pages, merged cells and tables spanning pages separately. A fast parser may be sufficient for straightforward text PDFs, while a heavier pipeline is justified only where structure matters. Compare not just seconds per page but the cost of correcting output, the fraction of unreadable files and the time staff need to find an answer.

For an isolated environment, decide in advance where model weights live and how they are updated, which components can make outbound connections, how much memory and CPU time batch processing consumes, and what is written to logs. Docling's MIT licence permits commercial use of its code, but licences for individual models and dependencies require separate review. Leaving that until after hardware procurement makes any monthly cost estimate incomplete.

A pilot without a large budget

Collect 100–200 documents from one process, including ordinary files, scans, tables and deliberately awkward exceptions. Preserve the originals and create a ground truth for important fields and 20–30 genuine staff questions. Measure today's manual time per document. Then run the local pipeline, recording machine time, human review time, exception rate and critical errors separately. Answers should identify their document and page; an answer without a verifiable source should count as a failure.

Make the rollout decision at two gates. The first is quality: the process owner sets the acceptable rate of missing critical fields and incorrect answers in advance. The second is cost: measured staff-time savings must cover operation, integration and correction costs within an acceptable period. If quality fails, an attractive hours-saved calculation is irrelevant. If the economic gate fails, local RAG may still be a useful search experiment, but it should not be sold as a cost-saving project.

Photo: Skot / Wikimedia Commons, CC BY-SA 4.0. The square cover is a crop of the photograph and is shared under the same licence. It depicts book digitisation, not Docling running or a result from the hypothetical calculation above.