Skip to content
VAULO
Back to the blog
  • Artificial intelligence
  • Metadata

AI keywording: what it takes over and where you have to look

Automatic image recognition opens up legacy archives in hours instead of months. It also invents things that are not in the picture. A sober assessment.

Marc ConzelmannManaging Director4 min read
A workshop photograph on a light table, marked with colored dots and a magnifying glass

The value of automatic keywording comes down to one number. An archive of 80,000 images that would have taken two people about nine months by hand is opened up by machine in one night. That is not an efficiency gain, that is the difference between "will get done" and "will never get done".

At the same time the temptation to treat the result as finished is strong. It is not finished. It is a very good first draft.

AI image keywording: what works reliably

Objects and scenes. Vehicle, factory hall, child, conference table, coastal landscape. On clearly recognisable subjects the hit rate in our evaluations stays consistently above 90 percent.

Text in the image. Signs, labels, shirt numbers, packaging print. This is often the single most valuable hit, because that text appears in no other field — a trade fair booth becomes findable through the booth name nobody would have captured.

People and faces. Technically reliable, but the most sensitive part of keywording. We recommend enabling face recognition separately and documenting it separately, not as a side effect of regular image keywording — whoever uses it should decide to do so deliberately, not have it run along as a default setting.

Limits of automatic keywording: where it systematically goes wrong

Context that is not in the picture. A model sees a group of people at a table. Whether that is a client meeting, a works council or an award ceremony is not decided by the image but by knowing about it. Keywords like these get guessed, and they sound convincing while doing it.

For exactly this case VAULO has a dedicated context field: it lets you tell the AI up front in what setting an image was taken, so it can take that information into account during keywording instead of guessing.

Anything containing a judgement. "Modern", "premium", "trustworthy", "sustainable". These come back surprisingly often and are worthless for search, because nobody searches for them and because they distinguish nothing.

Domain vocabulary from your own industry. A model recognises a metal part. Whether it is a pinion, a sprocket or a feather key it does not know, and your own staff search with exactly those words.

Negations. No model reliably tags what is *not* visible. That sounds academic until somebody needs "product shot without people".

Alt text and accessibility: what automatic image description means

Automatically generated image descriptions are a good starting point for alt text, but not a replacement. Alt text does not describe what is in the image, it describes what the image stands for in this particular place — the same photo needs different text in a press area than in a product catalogue.

The workable route: the machine description as a pre-fill, the editorial adjustment as a mandatory step before publication. That is considerably less work than writing from scratch and considerably more reliable than accepting it unchecked.

In VAULO this can be mapped directly as an approval step: an asset whose alt text is machine-generated only stays in an "unreviewed" state and is released for publication only once someone has confirmed or adjusted the text.

Conclusion: the honest summary

Automatic keywording does the part of the work that would otherwise be left undone, and it does it well. It does not replace the person who decides which terms mean anything inside your organisation. Run it as a suggestion engine and you gain months. Run it as an authority and you build an archive full of plausible wrong answers.

FAQ

Can AI keywording completely replace manual keywording?

No. It reliably handles what is visible in the image — objects, scenes, text in the image. Context, judgements and your industry's domain vocabulary it cannot deliver reliably, because the knowledge for them is not in the picture. It replaces the groundwork, not the decision about which terms mean anything inside your organisation.

How does the AI get information that is not visible in the image, such as the occasion of a photo?

Through a context field, as offered by VAULO for example: it lets you store background information about the image that the AI can take into account during keywording instead of guessing it.

Should I let face recognition run automatically alongside?

We would not recommend it. Face recognition should be enabled deliberately and separately from regular image keywording, with its own documentation of who uses it and why — not as an automatic side effect.

Does an automatically generated image description replace alt text?

No. A machine description is a good pre-fill, but not finished alt text, because alt text describes what the image stands for in its particular place, not just what is visible in it. The editorial adjustment before publication therefore remains a separate, mandatory step.

Blog

Keep reading

Two trays on a light table — eight neat index cards on the left, an overflowing stack on the right
  • Metadata
  • DAM basics

Metadata that pays off: eight fields instead of eighty

The most common way a DAM rollout fails is a metadata model that is too good. How to find the fields that actually get maintained — and leave out the rest.

Marc ConzelmannManaging Director4 min read
A telephoto camera and sports prints with marker flags on a light table
  • Events
  • Artificial intelligence
  • Workflow

Pictures in seconds, not minutes: DAM at the sports event

From the camera chip straight into the system, automatically keyworded, instantly findable: how LiveUpload, VisionAI and VisionFace turn event media work from a bottleneck into a strength.

Marc ConzelmannManaging Director2 min read

Try it yourself

See how AI keywording with VisionAI performs on your own photographs.