- Artificial intelligence
- Metadata
AI keywording: what it takes over and where you have to look
Automatic image recognition opens up legacy archives in hours instead of months. It also invents things that are not in the picture. A sober assessment.

The value of automatic keywording comes down to one number. An archive of 80,000 images that would have taken two people about nine months by hand is opened up by machine in one night. That is not an efficiency gain, that is the difference between "will get done" and "will never get done".
At the same time the temptation to treat the result as finished is strong. It is not finished. It is a very good first draft.
AI image keywording: what works reliably
Objects and scenes. Vehicle, factory hall, child, conference table, coastal landscape. On clearly recognisable subjects the hit rate in our evaluations stays consistently above 90 percent.
Text in the image. Signs, labels, shirt numbers, packaging print. This is often the single most valuable hit, because that text appears in no other field — a trade fair booth becomes findable through the booth name nobody would have captured.
People and faces. Technically reliable, but the most sensitive part of keywording. We recommend enabling face recognition separately and documenting it separately, not as a side effect of regular image keywording — whoever uses it should decide to do so deliberately, not have it run along as a default setting.
Limits of automatic keywording: where it systematically goes wrong
Context that is not in the picture. A model sees a group of people at a table. Whether that is a client meeting, a works council or an award ceremony is not decided by the image but by knowing about it. Keywords like these get guessed, and they sound convincing while doing it.
For exactly this case VAULO has a dedicated context field: it lets you tell the AI up front in what setting an image was taken, so it can take that information into account during keywording instead of guessing.
Anything containing a judgement. "Modern", "premium", "trustworthy", "sustainable". These come back surprisingly often and are worthless for search, because nobody searches for them and because they distinguish nothing.
Domain vocabulary from your own industry. A model recognises a metal part. Whether it is a pinion, a sprocket or a feather key it does not know, and your own staff search with exactly those words.
Negations. No model reliably tags what is *not* visible. That sounds academic until somebody needs "product shot without people".
Alt text and accessibility: what automatic image description means
Automatically generated image descriptions are a good starting point for alt text, but not a replacement. Alt text does not describe what is in the image, it describes what the image stands for in this particular place — the same photo needs different text in a press area than in a product catalogue.
The workable route: the machine description as a pre-fill, the editorial adjustment as a mandatory step before publication. That is considerably less work than writing from scratch and considerably more reliable than accepting it unchecked.
In VAULO this can be mapped directly as an approval step: an asset whose alt text is machine-generated only stays in an "unreviewed" state and is released for publication only once someone has confirmed or adjusted the text.
Conclusion: the honest summary
Automatic keywording does the part of the work that would otherwise be left undone, and it does it well. It does not replace the person who decides which terms mean anything inside your organisation. Run it as a suggestion engine and you gain months. Run it as an authority and you build an archive full of plausible wrong answers.
FAQ
Can AI keywording completely replace manual keywording?
No. It reliably handles what is visible in the image — objects, scenes, text in the image. Context, judgements and your industry's domain vocabulary it cannot deliver reliably, because the knowledge for them is not in the picture. It replaces the groundwork, not the decision about which terms mean anything inside your organisation.
How does the AI get information that is not visible in the image, such as the occasion of a photo?
Through a context field, as offered by VAULO for example: it lets you store background information about the image that the AI can take into account during keywording instead of guessing it.
Should I let face recognition run automatically alongside?
We would not recommend it. Face recognition should be enabled deliberately and separately from regular image keywording, with its own documentation of who uses it and why — not as an automatic side effect.
Does an automatically generated image description replace alt text?
No. A machine description is a good pre-fill, but not finished alt text, because alt text describes what the image stands for in its particular place, not just what is visible in it. The editorial adjustment before publication therefore remains a separate, mandatory step.


