What actually happened
Nine small and medium-sized companies took part in a Canton of Zurich program in which they and the Swiss Data Science Center (SDSC) built an open RAG prototype. The practical scenario was finding and reconciling documents about the environmental properties of packaging: product declarations, supplier materials, company reports and regulatory texts. The project presents this through a PrimePack AG use case. This was collaborative prototyping, not a verified production rollout at all nine companies.
The problem is familiar to a Russian supplier of packaging, components or building materials. A customer asks whether a specific product attribute has been substantiated. The catalog contains a supplier claim, a declaration uses different wording, and the relevant document revision sits in an email inbox. A fast but unsubstantiated response can cost more than a slow one. RAG is useful here not as a legal expert, but as a way to find evidence and prepare a response for the accountable employee.
According to SDSC, the baseline pipeline covers document chunking, embeddings, storage, retrieval and answer generation. The team developed four tracks on top: quality evaluation, structured answers, hybrid retrieval and multi-step questions. The code is published under the MIT license. SDSC reports a working prototype and a shared architecture, but does not provide verified production accuracy, time savings or monetary returns. Those results must not be attributed to the participating companies.
Why attaching a link is not enough
A document link does not make a conclusion correct. A product card may contain a manufacturer's claim, a report may describe tests on a different batch, and a regulation may state a requirement without proving that this product meets it. In the demonstration scenario, the prototype is intended to distinguish verified information, claims, missing evidence and mixed or conflicting evidence. The repository shows the labels VERIFIED, CLAIMED, MISSING and MIXED in a structured JSON response.
For a business, this implies a separate verification chain:
- record the owner, date and revision of every document;
- associate each excerpt with the correct SKU, supplier and relevant batch;
- show the employee the exact passage, not merely a filename;
- flag an unsupported statement for review;
- preserve the human decision and the version of evidence used.
The final two items are our recommendations for a production workflow, not features claimed for every participant's prototype. A quality or procurement employee remains responsible for the decision and for signing off an external response. If a source is withdrawn or replaced, an earlier generated response must not be treated as current.
What a smaller company needs to replicate it
Start with one narrow question, such as: “What evidence supports the environmental attributes of this box range?” You need a governed set of documents, not thousands of files: customer contract requirements, product records, declarations, supplier letters and internal review policies. Assign each document a source, version, effective date, evidence type and access rights. Without that metadata, vector search will find similar words but not necessarily current evidence.
Next come PDF and spreadsheet ingestion, an update log, lexical and semantic search, an answer generator and a review screen. SDSC compares vector search, keyword search and a hybrid approach using Reciprocal Rank Fusion. This is not a reason to choose the most complex option by default: the pilot should compare which documents and passages are actually retrieved for typical queries, including part numbers, abbreviations and supplier-specific wording.
At first, ERP or CRM integration can be read-only: retrieve the SKU and product status, but do not let the model alter master data or send a customer response without approval. A decision record should hold the product identifier, question, document-version references, reviewer and timestamp. Only after quality is established should the team consider automatically creating a draft in the operational workflow.
A local model is not the same as a fully local system
The SDSC repository demonstrates generation through Ollama. However, the README describes OpenAI embeddings for the demo text-index build and also presents local-component alternatives. Switching the generator to Ollama alone therefore does not establish that documents and queries stay off external services. This matters for businesses handling confidential contracts and specifications.
Before deployment, inspect the complete chain: parsing, embedding, query rewriting, synthetic-test generation, LLM-as-judge evaluation, telemetry and backups. Record each component's endpoint and permitted data class. If the requirement is a closed environment, replace external calls with local alternatives and inspect outbound network connections in a test setup. This is a technical check, not a claim of compliance with any particular regulatory regime.
Measure value without inventing outcomes
The prototype includes evaluation for faithfulness, answer relevance, context precision and context recall, plus a synthetic question set. These metrics help compare configurations, but synthetic questions cannot replace genuinely difficult customer requests. The independent NIST Generative AI Profile likewise emphasizes testing, evaluation and validation appropriate to the use case.
For a pilot, 30–50 anonymized queries from actual correspondence are enough to start, including cases where the right answer is “insufficient evidence.” An expert identifies the required documents and acceptable conclusion in advance. Then compare two workflows: time spent without the assistant versus with it, the share of answers needing correction, links to the wrong document revision and unsafe claims stopped by a reviewer. Include document preparation and index maintenance in the cost.
Here is an explicitly illustrative calculation. Assume 100 such requests per month, 12 minutes of manual searching and checking per request, and seven minutes to review a RAG draft. The potential gross saving is 500 minutes, or eight hours and 20 minutes per month, before ingestion, corrections and ongoing support. This is neither a result of the Swiss project nor a forecast for a particular business. At low request volume a dedicated server may not pay for itself; when an incorrect claim carries a high cost, evidence quality matters more than a few minutes saved.
Next step
Choose one product group and one type of customer question. Assemble 30–50 documents with owners and revisions; mark sources that are forbidden or obsolete. For two weeks, compare ordinary search with RAG in draft mode, while an employee alone sends external answers. If the system cannot point to exact evidence or acknowledge that evidence is missing, it is too early to expand it across the catalog.
The cardboard photograph is illustrative and does not depict participants in the Swiss project. Photo: Henry Söderlund; CC BY 2.0; square crop applied.
