The Fields This Archive Cannot Count, and What That Costs
2,592 technical facts wear 1,598 different labels and 1,754 ticket prices wear 1,355. A fact written in prose is not data, and here is what that costs.
Records in this piece: Ewer signed by Yunus ibn Yusuf al-Mawsili, Bowl, Potpourri Vase (Vase potpourri a vaisseau), The Archimedes Palimpsest, The Ardabil Carpet

On this page 10 sections
ContentsClose
In short
- The 293 published records carry 2,592 rows of technical evidence under 1,598 different labels. The commonest label, Dimensions, is used 168 times; most labels are used once, so almost none of that evidence can be counted or compared across objects.
- The split is often a spelling. Inscription appears 41 times and Inscriptions 21, Material 31 and Materials 20. Nothing in the data says those are the same field, so any query for one silently loses the other.
- The venues corpus is worse. Its 353 museums carry 1,754 admission rows under 1,355 distinct tier labels, and the ordinary adult ticket is written at least as Adult, Adults, General admission, General, Adult (18+) and Full price.
- Founded is not a date. Of 353 museums, 224 give a bare year and 129 give a passage of prose holding several different events, so a median founding year computed from this field is an artefact of whichever year happens to appear first in each sentence.
- Two records carry an attribution status of documented, which is not one of the attribution values at all. It is borrowed from the provenance vocabulary, where documented, disputed and unknown are the three grades.
A fact written in prose is not data
This archive was built on the rule that every claim carries a source, and it keeps that rule well. What it keeps badly is the separate promise implied by a structured record, which is that facts of the same kind sit in the same place under the same name. A reader does not notice the difference. A question does, the moment anyone asks how many, how often or compared with what. A record that reads well and counts badly is a common outcome, and it is invisible from the reading side, which is why it survives for years in archives that are otherwise careful.
The house has known this pattern for a while and named it: facts live in prose, not in fields. What follows is the measurement. It is not flattering and it is the sort of thing an archive should publish about itself, because the alternative is quietly serving statistics that its own data cannot support.
2,592 technical facts under 1,598 labels
Every published record carries a technical section, a list of label and value pairs holding what is physically known about the object: support, medium, weight, inscription, condition. Across 293 records there are 2,592 such rows and they are written under 1,598 distinct labels. That ratio is the whole problem in one line, because it means the typical label is used once and never again.
| Label | Rows | Near-duplicate also in use |
|---|---|---|
| Dimensions | 168 | Dimensions and weight, 16 |
| Support | 66 | Support and medium, 8 |
| Medium | 45 | Material and technique, 10 |
| Inscription | 41 | Inscriptions, 21 |
| Condition | 41 | none in the top labels |
| Material | 31 | Materials, 20 |
| Signature | 17 | Signature and date, 7 |
Ask this archive how many objects carry an inscription and the honest answer is that it cannot tell you, because Inscription and Inscriptions are two fields as far as the data is concerned and neither one is the whole answer. The same is true of material, of signature and of dimensions. The information is all there and a person reading a page sees it perfectly well. A count does not.
The adult ticket has at least six names
The venues corpus holds 353 museums and 1,754 admission rows, and those rows carry 1,355 distinct tier labels. The ordinary full-price adult ticket, the single most useful number a visitor could want, is written as Adult 87 times, General admission 25 times, General 10 times, Adult (18+) 6 times and Full price 6 times, with Adults and Foreign adult alongside. A strict match on the obvious spellings identifies 127 adult prices out of 1,754 rows.
That is not a rounding problem. It means a question as basic as what it costs to get into these museums can be answered for about seven per cent of the recorded prices without a human reading the rest. The prices themselves are correct, sourced and dated. It is the label above them that stops them being usable, and the label was free text because nobody decided it should not be.

Founded in is not a date, it is several events
Of the 353 museums here, 224 give a founding date as a bare year and 129 give a passage of prose. The prose is better history and worse data. Museum Boijmans Van Beuningen records a foundation on 3 July 1849 built on the bequest of a collector who died in 1847, and a renaming in 1958 after a second collector added his own holdings. The Getty entry holds 1954, when J. Paul Getty opened his ranch house, 1974 for the Villa and 1997 for the Center.
Each of those sentences is accurate and each contains three or four candidate answers to the question of when the museum was founded. Take the first four-digit year from every entry, as any naive count must, and the corpus reports a median founding year of 1926. That number is not wrong so much as meaningless: for a third of the museums it is whichever event the writer happened to mention first.

One institution, written two ways, in the same column
The artists corpus records where each maker's work can be seen: 879 holdings rows across 106 artists, naming 528 distinct institutions. Two of those 528 are the Metropolitan Museum of Art, which appears 45 times under that name and 11 more as The Metropolitan Museum of Art. Two more are the National Gallery, 13 times without the article and 8 with it.
A ranking of the institutions that hold the most work by the makers in this archive therefore understates the Met by about a fifth and puts the National Gallery in the wrong place entirely. The fix is an hour of normalisation and a rule about the definite article. The lesson is that nobody would have found it by reading the pages, because on a page the two spellings look like the same museum, which is exactly what they are.
A value borrowed from the wrong vocabulary
Every record grades its attribution, and the grades are meant to be a small closed set. Across 293 records the values run: accepted 176, disputed 52, unknown 46, workshop 8, attributed 8, documented 2 and rejected 1. Six of those are attribution language. Documented is not.
Documented belongs to the provenance vocabulary, where every link is graded documented, disputed or unknown. It reached the attribution field on the Ardabil carpet and the Archimedes palimpsest, and in both cases the writer meant something true and specific: the object carries a signature and a date that can be read. The right way to say that was accepted, with the inscription in the technical section doing the work. Two overlapping vocabularies with one shared word between them is all it takes.

The count nobody could run
The test that settled this was an ordinary question: how many objects in this archive carry an inscription. It should take one query. It takes a person, because Inscription and Inscriptions are separate labels, because seven more records file the same evidence under Signature and Signature and date, and because a further group put the reading of the inscription inside a Dimensions row where the object's measurements and its text are described in one paragraph.
The same question asked of the provenance data answers instantly, and the reason is instructive. Provenance links carry a certainty graded documented, disputed or unknown, from a list of exactly three values decided once and applied everywhere. Across 293 records that produces 2,057 graded links, 1,567 of them documented, 310 unknown and 180 disputed, and those totals can be trusted because there was never a fourth spelling available to anyone. The difference between the two halves of this archive is not effort or care. It is whether somebody wrote the permitted values down before the first row was entered.
What the object does better than the database
The objects in this archive are frequently better structured than the records describing them. A Mosul metalworker signs in the same place with the same formula, maker then role then date. The Sevres manufactory marks the underside with the year and the decorator. Abu Zayd signs and dates a bowl in 1187 and the Met's own record of it, eight centuries later, has no field in which to say who owned it.
None of those makers had a data model. They had a convention, applied every time, in a place everybody knew to look, and that turns out to be the entire trick. A field is only a field if the same thing always goes in it under the same name. Where that holds, a thirteenth-century inscription is queryable. Where it does not, a record written last month is not.
What to fix, and what to leave alone
The repair is not to force everything into a form. Technical evidence about art is genuinely various and a schema that admits only twelve labels would lose the interesting rows, which are the ones that exist once because that object is unusual. The fix is a small controlled vocabulary for the labels that recur, a rule that new labels are checked against it before they are written, and a free-text label kept for the genuine one-offs, marked as such so a count can exclude it honestly.
Admission tiers should be a closed list, because a ticket price is not various. Founded should be a year with the prose beside it rather than inside it. Institution names should be normalised once against an authority. And any statistic this archive publishes off a free-text field should say which field it came from, so a reader can discount it. Publishing the median founding year without that warning would have been the real error, not the field.
Why this archive is free to read
We do this for everyone who loves art. The people who have studied it for years, and the people who just stopped to look. Iconotheca exists so that what we learn belongs to all of them.
Nobody pays us for this. No ads, no sponsors, no commercial interest. These works belong to all of us, and we would rather share them with as many people as we can than keep them to ourselves.
If it gave you something today, tell us to keep going. Follow us, leave a like, or write a positive comment. We read every one, and they are what keeps us going.
Questions
Because sourcing and structure are different disciplines. A sourced claim proves that a statement is true; a structured field proves that two statements are about the same kind of thing. This archive enforces the first rigorously and the second hardly at all, so it can defend any single sentence and cannot reliably total a column.
A fixed list of permitted values for a field, checked when the value is written rather than when it is read. Museums use them for object type, medium and role, most often against a published authority such as the Getty vocabularies, so that a search for one term returns every record that means it.
No. Every row measured here carries its source and its value is correct. What the labelling costs is comparison: the archive cannot say how many objects carry an inscription, or what the typical adult admission is, without a person reading the rows, because the same fact is filed under different names.
Because 129 of the 353 founding entries are prose holding several events, such as a collector's bequest, a building's completion and a reopening under a new name. A count that takes the first four-digit year in each entry is averaging whichever event each writer chose to mention first, not the foundings.
No. Technical evidence about art is genuinely various, and a label used once is often the most interesting row on the page. The workable rule is a closed list for the labels that recur plus an explicitly marked free-text label for the one-offs, so that counts can exclude them honestly instead of silently.
Sources
- 1The Metropolitan Museum of Art, Open Access Collection API, object 451753, accession 64.178.2, the Abu Zayd bowl. Queried directly on 24 September 2026. The response carries medium, date and dimensions as fields and has no objectHistory and no provenance key at all. collectionapi.metmuseum.org/public/...
- 2The Walters Art Museum, object record 54.456, Ewer, 644 AH/AD 1246-1247, acquired by Henry Walters 1917. Source for the signature formula quoted in the hero caption and for the museum's own transcription of the inscription. art.thewalters.org/object/...
- 3The Walters Art Museum, object record 48.559, Potpourri Vase (Vase potpourri a vaisseau), 1764, acquired by Henry Walters 1928. Read in the in-app browser on 24 September 2026. Source for the designer, the decorator and the museum's own marks section. art.thewalters.org/object/...
- 4The Archimedes Palimpsest Project, The History of the Archimedes Manuscript. Read 24 September 2026. Source for the scribe, the 1229 dating of the prayer book and the imaging programme behind the plate reproduced here. archimedespalimpsest.org/about/...
- 5The Metropolitan Museum of Art, press release, The Met Makes Its Images of Public-Domain Artworks Freely Available through New Open Access Policy, 7 February 2017. Source for the CC0 designation and for the structured collection data released alongside the images. metmuseum.org/press-releases/...
More in Research and Method
The partThe First Number in a Hoard Is a Counting Rule
The Staffordshire Hoard is 4,600 pieces, or just under 4,000, or around 700, or more than 600, depending on which official page you read.
The First Link Is in the Maker's Own Hand
A maker's bill, workbook or memoir is the only provenance document that fixes date, cost and first owner at once.
The date on the object dates one layer of it
Seven dated objects in this archive show that a date belongs to one layer of the thing: the parchment, the copying, the printing, the compilation.