Data first, then a camera with an “intelligent” answer

A small production line wants to find defects in photos of parts. Choosing a model seems like the first step. In practice, the team must first agree on what counts as a defect, whether it can be seen in the image, and who resolves an ambiguous case. Otherwise, even a careful detector will reproduce disagreements in the labels: one shift marks a scratch; another considers it acceptable.

CVAT is a tool for annotating images and video. Its Community edition can be self-hosted. Teams can create tasks and label objects; automatic pre-annotation is possible if an appropriate model is connected to the instance. CVAT documentation distinguishes the editions: Nuclio-based functions are available to Community, while third-party and native functions have different edition availability. CVAT is not a ready-made factory inspection system; it is a workspace for preparing and reviewing data.

What follows is an editorial pilot design for a small manufacturer, not a story about a particular company's realized savings. The cover shows an official CVAT automatic-annotation settings screenshot with label mapping, a threshold, and a region of interest. The screenshot comes from project documentation in the repository under the MIT License. The original is placed intact on a square canvas over a blurred extension of itself.

Define the task before adding a model

Choose one operation: detecting a missing hole, a large scratch, or a shifted label, for example. “Find any defect” is too broad for a first dataset. Each class needs a positive example, a counterexample, and a rule for borderline cases. Annotation instructions should state whether to mark the entire object or only the defect, the minimum size, how to handle reflections, and when photo quality is insufficient for a decision.

The photos should reflect real conditions: different batches, shifts, backgrounds, lighting, angles, and cameras. Several nearly identical frames of one part are not a substitute for diversity. Do not collect only obvious defects; include difficult normal parts that resemble faulty ones. Otherwise, the demo may work while operators on the line receive a stream of false alerts.

Define the decision process: a model proposes a label, an employee confirms or rejects it. At first, the model should not automatically reject a product. Where an error is costly or dangerous, a person remains the final decision-maker until safety and economic effects have been evaluated separately.

What to do in CVAT

Create a project with a short class list and tasks for the photos. Label some images manually as reference examples and review disputed cases with a process engineer. The editor supports rectangles, polygons, and other object types, but choose the simplest representation sufficient for the decision. If the requirement is only “defect present or absent,” pixel-perfect outlines may add cost without value.

CVAT documentation describes a Review mode for manually finding and correcting annotation problems. An important distinction: the CVAT page labels automated quality assessment via Ground Truth and Honeypots as a CVAT Online and Enterprise feature. Do not promise it to the owner of a local Community installation. In Community, a team can arrange manual double-checking of a selected sample and record disagreements separately.

If a suitable detector is available, pre-annotation can reduce routine work. CVAT's official instructions require that the model actually be connected to the instance; users then map its labels to project labels, choose a threshold, and optionally specify a region of interest. For Community, deploying Nuclio functions is a separate technical task, not an “enable AI” toggle. Even after automatic annotation, a person must check missed defects and unnecessary boxes.

The cover screenshot shows `person` and `car` labels as examples of the interface, not a proposed inspection model. Your application needs suitable labels and a separate review of the model weights' usage rights. A permissive software license does not automatically grant commercial rights to every detector connected to the tool.

Split data so the test is honest

Randomly splitting individual files is risky: adjacent frames of one part can end up in both training and test sets. The result then measures recognition of the same scene, not performance on a new batch. Split by batch, time, or physical object so images of one part do not cross sets. Keep another set of images captured after a lighting or camera change.

Track at least two types of error. A miss lets a defective part through; a false alarm sends a normal part for extra inspection. Their costs can differ by an order of magnitude, so one overall accuracy percentage is inadequate. A small business needs a simple table: total parts, actual defects, model misses, unnecessary stops, and controller minutes spent reviewing.

NIST's AI risk measurement guidance emphasizes data quality and representativeness and evaluating performance in the conditions of intended use. In a pilot, this means testing not only attractive photos from the annotation set but also new batches, weaker lighting, and a dirty camera. The metric must belong to the actual workflow, not to a generic benchmark.

Infrastructure and data

For a first cycle, a camera with fixed geometry, stable lighting, an image store, and an annotation owner are enough. CVAT Community can be hosted in the company's own environment when product photos must not go to an external service. Administration, backups, updates, and access restrictions for original images are still required. The photographs may reveal product design, serial numbers, or a workplace, so agree on storage and transfer rules before uploading them.

CVAT can export data in several formats, but test the chosen format on a small task before annotating at scale: masks, boxes, and attributes are transferred differently. Keep a task backup and original photos separately from the training archive. Changing a defect category halfway through the project calls for a versioned label scheme and a review of previously annotated frames.

Running a detector locally on the production line is a later stage, separate from using CVAT. GPU, latency, and controller-integration requirements depend on the chosen model and conveyor speed. Do not buy a GPU based on annotation project size; first test the actual stream, including image transfer and operator action time.

Economics and a two-week pilot

An illustrative calculation, not a CVAT result: suppose a line photographs 2,000 parts per shift and 1% are defective—20 faulty parts. If the detector misses 10% of defects, two defects remain unnoticed. If it falsely flags 5% of normal parts, about 99 parts need manual reinspection each shift. At 30 seconds per check, that is almost 50 minutes of extra work. The economics depend on the cost of a miss, the cost of a false alert, and the actual defect rate; every number here is an assumption for testing the formula.

For two weeks, choose one part and one visible defect. In week one, collect photos from several batches, agree on annotation rules, and manually review disputed examples. In week two, test a simple baseline or an already available model on a held-out batch without changing the inspector's decision. Record misses, false alarms, review time, and causes of errors. If the data cannot reliably distinguish the defect, improve lighting and camera angle before making the model more complex.

A manager needs more than an “AI accuracy” percentage: the decision is whether to continue annotation, improve capture, or stop trying to automate this particular defect. CVAT helps make the data manageable; people remain responsible for product quality and investment decisions.

Image credit: CVAT.ai Corporation, official CVAT repository, MIT License.