Skip to content
VAULO
Back to the blog
  • Metadata
  • DAM basics

Metadata that pays off: eight fields instead of eighty

The most common way a DAM rollout fails is a metadata model that is too good. How to find the fields that actually get maintained — and leave out the rest.

Marc ConzelmannManaging Director4 min read
Two trays on a light table — eight neat index cards on the left, an overflowing stack on the right

The first workshops reliably produce a metadata model everybody is proud of. It has fields for campaign, product line, region, photographer, agency, usage type, colour world, subject category, approval status, contact person, cost centre and occasion. It is complete, it is well thought through, and six months later four of them are filled.

That is not a matter of discipline. It is that a field only gets maintained when the person filling it in gets something out of it. Everything else is work for somebody else, and in day-to-day work that loses every prioritisation.

The three origins

Before you talk about fields, sort them by who supplies them. Almost everything follows from that.

Automatic, from the file. Capture date, camera model, focal length, resolution, colour space, GPS position, embedded IPTC data from the image editor. These fields cost nobody any time and are still the basis of most good searches. Take everything the file offers.

Automatic, from the context. Who uploaded it, into which area, under which process, with which approval. That also happens without effort, provided the upload route is defined properly.

By hand. Everything somebody has to type. This is the scarce resource, and that is why a hard ceiling applies here.

The ceiling

Our rule of thumb: at most eight mandatory fields filled by hand, and for every single one somebody has to be able to answer which search fails without it.

If the answer is "finance will need that at some point", it is not a mandatory field. It is an optional field, and optional fields stay empty — which is fine, as long as nobody later builds a report on them.

A model that has held up across several projects:

FieldOriginWhy
TitleBy handWhat appears in result lists. Without it nobody reads the hits.
DescriptionHand or AIFull text search works here, not on the filename.
KeywordsAI, curatedThe actual entry point into the archive.
Rights statusHand, picklistDecides usability. Never a free text field.
Expiry dateHand, only where neededWithout this field nobody can warn before expiry.
CreatorHand or IPTCLegally required on publication.
AreaContextDrives visibility and permissions.
Capture dateAutomaticThe most frequently used sort order there is.

Anything beyond that starts as an optional field and becomes mandatory once it has proven itself. Not the other way round.

Picklists instead of free text

The second biggest lever after the field count. A free text field for rights status produces "free", "released", "approval granted", "ok", "OK per Miller" and "see contract" within a year. After that no reporting is possible, and cleaning up is manual work.

Where a finite set of answers exists, it belongs in a picklist. That applies to rights status, usage type, approval level and area. It explicitly does not apply to keywords — those should be allowed to grow, but they need a person who consolidates them regularly.

After two years we had 11,000 distinct keywords, 3,400 of them used exactly once. Consolidating took four days and nearly doubled the hit rate.

Archive lead, public sector client

What AI can take over

Image content, objects, scenes, text in the image, colour world, and to some degree mood. For a new archive that saves most of the capture work, and it makes legacy archives searchable at all.

What AI cannot take over is everything that sits outside the image: which campaign it belongs to, who paid for it, how long it may be used. That is exactly why the list above stays as short as it is — and why it is worth knowing the limits of automatic keywording before you design a model around it.

The six month test

Take a sample of a hundred assets from live operation and count the fill rate per field. Anything below 70 percent is either not a mandatory field or not a sensible field. Delete it or make it optional, rather than writing reminder emails.

A model with eight maintained fields beats one with thirty half-empty ones on every measure that counts: hit rate, trust in the search, and the willingness to upload anything at all.

Blog

Keep reading

AI-generated still life on a light table: on the left a loose pile of photo prints and notes, on the right the same prints neatly sorted into an archive grid with tabs and colored markers
  • DAM basics

What is a DAM? Digital asset management explained simply

A DAM system is the central software that companies use to organize, find, share and securely manage images, videos, logos and documents — instead of losing them in folders, emails and clouds.

Marc ConzelmannManaging Director11 min read
A workshop photograph on a light table, marked with colored dots and a magnifying glass
  • Artificial intelligence
  • Metadata

AI keywording: what it takes over and where you have to look

Automatic image recognition opens up legacy archives in hours instead of months. It also invents things that are not in the picture. A sober assessment.

Marc ConzelmannManaging Director4 min read
Four folders holding the same architectural prints, stacked messily on a light table
  • DAM basics
  • Workflow

A shared cloud folder is not a media archive

SharePoint, Dropbox and the network drive solve storing files. They do not solve finding them again. How to recognise the point where a folder stops carrying the weight.

Sebastian KüstersManaging Director4 min read

See VAULO against an archive of your own

Thirty days, your own files, no credit card. After that you will know whether the search holds up to what this article promises.