The task: answer from documents without mixing up customers
DocsBot helps companies build chat assistants on their own support materials: documentation, knowledge bases and internal pages. For this kind of product, the question is not merely whether search finds the right paragraph. The service must also ensure that one customer's assistant never receives another customer's paragraph. For a small developer, that boundary matters more than an attractive model demo.
In Weaviate's published case study, DocsBot is described as a product built by solo founder Aaron Edwards. After a sudden influx of users, the initial approach to many separate indexes no longer met its scalability and operating-cost needs. A one-person business needed relevant document retrieval, customer-data isolation and manageable operations. This is an infrastructure vendor's account of its own customer, not an independent business audit.
What was actually built
According to the case study, DocsBot ingests documents, splits them into chunks and generates embeddings itself. Vectors and metadata are stored in Weaviate Cloud. When a question arrives, the service uses semantic or hybrid search, passes the retrieved context to a language model and generates an answer. Multi-tenancy lets it serve many isolated datasets in one cluster.
Weaviate reports more than 50,000 tenants in one cluster and more than 6.1 million customer questions answered in a year. These are vendor-reported case-study figures: the counting method, answer quality, agent time saved and return on investment are not disclosed. They must not become a promise to another company. Agent actions such as routing requests or triggering workflows are described as the product's next direction, not a verified outcome of this deployment.
The useful detail is not the tenant count by itself. DocsBot chose a managed cloud database because its founder needed to spend time improving the product rather than continuously operating dedicated search infrastructure. A Russian business with secure-environment requirements may make the opposite choice, but then the cost of administration, redundancy and monitoring stays with the business.
Why one shared index is not access control
RAG is often drawn as a short chain: document, retrieval, model, answer. A multi-customer service needs another mandatory step before retrieval: the server must determine which customer and user the request represents. It cannot simply trust a parameter that a browser or integration can change. OWASP explicitly recommends binding tenant context to verified identity and separately checking authorization.
Even if a vector database can separate tenants, the boundary must hold throughout the pipeline:
- ingestion assigns a verified owner and permissions to each document;
- processing and re-indexing do not mix different customers' queues;
- retrieval is restricted to the authorized tenant before passages reach the model;
- caches, files, logs and backups preserve the same boundary;
- customer offboarding removes or isolates data according to the agreed retention period.
Otherwise, the robot may be remarkably diligent: it brings exactly the right folder, only from the neighboring cabinet. That works as an illustration, but an access-control test must return a strict zero.
What transfers to an on-premises setup
A Russian integrator does not need to copy DocsBot's cloud configuration or scale a pilot to 50,000 tenants. The transferable part is the order of decisions. First list the owners of the knowledge: branches, dealers, separate legal entities or SaaS customers. Then decide which boundary they require: a separate database, a collection, a tenant within a collection, or stronger physical separation. The choice depends on access requirements, data volume, load and the acceptable blast radius of an error.
Two or three trial tenants with de-identified but clearly different documents are enough for a pilot. User A asks about a fact that exists only in tenant B. The system must refuse or say that the information is unavailable; B's semantically similar passage must not reach either the model prompt or the answer citation. Repeat the test after re-indexing, a role change, session expiry, backup restoration and cache use. Measure not only answer accuracy but also the proportion of retrieved documents owned by the wrong tenant; the target there is zero.
The pilot also needs normal operating metrics: 95th-percentile latency, the share of answers with verifiable sources, refusals caused by insufficient context, cost per question, index size per customer and time to delete data. A local model may be appropriate when data cannot leave the organization or the query flow is predictable enough, but local deployment alone will not repair a retrieval authorization bug.
Economics: cost the entire question path
The case study does not disclose DocsBot's budget, so claiming that this approach is cheaper would be invented. Compare at least two scenarios for your own company on the same question set: a managed service and an on-premises stack. Include document preparation, embeddings, vector storage, model compute, backups, monitoring, updates and specialist labor. Assess the impact of a tenant-isolation failure separately; it should not be hidden in an average token price.
There is no universal break-even point. With few questions and policies that allow an external provider, owned infrastructure may sit idle. With high volume and data that must remain inside, a local option may be more sensible, but load and maintenance cost must be measured. Hold availability and retrieval-quality requirements constant across the comparison; otherwise a price table is comparing different products.
A next step without a long project
Choose one workflow, such as answers from instructions for two dealers, and map who uploads documents, who asks questions, who may see each folder and when it must be deleted. On a de-identified sample, test ten ordinary questions and ten attempts to reach another tenant's information. If access isolation holds and answers cite authorized sources, expand the pilot. If not, postpone model selection and fix the access boundary. That turns the DocsBot lesson into a testable decision rather than a race to reproduce someone else's scale figures.
