Skip to content
GDELT Project
← All Posts

Gemini For Museums: Rethinking What Metadata Museums Should Collect In An AI Era

The museum and cultural heritage worlds have long explored the concept of how to represent highly complex items in simple concise descriptive language and metadata as item records in their catalogs. Unlike mass-produced books that have standardized citations, museum objects by their very nature are often one-of-a-kind or have unusual characteristics. Libraries and archives encounter the same issues with manuscript (handwritten) works given their singular nature. Over the past half-century, the rise of computerized search has led to a growing focus on how to make records and their associated metadata increasingly searchable and how best to represent an item in digital form. How is all of this changing in the AI era? In short, instead of the traditional focus on standardization, museums should focus more on describing what is unique about their specific copy.

  • Standardization Is Both Obsolete & Solved. Institutions, especially libraries and archives, have gone to great lengths to standardize their item records, which has unlocked incredible new possibilities for cross-institutional search. In classical relational database-style catalog search systems, precise standardized item records and descriptions were critically important to allow for search. If a book's precise publication date was unknown beyond somewhere in the early 1560's, if one library listed it as 1560, another as "circa 1560", another as "1560-1565" and yet another as 1562 due to the year they accessioned it, a traditional catalog search would be unable to link these four records. Modern LLMs are exceptionally good at looking across all of these differences: not only realizing these all likely refer to the same item, but capable of interactively iterating with the user to refine their query and precisely confirm which copies match their search. While standardization is always helpful, the relentless focus on precise formatting is far less useful. Moreover, for use cases where standardization is a requirement, modern LLMs are capable of flawlessly reformatting items into infinite formats and standards. However, in some areas of standardization, like antiquarian book sizes, standardization is actually becoming increasingly problematic as search engines need to distinguish the actual size of a "4to", "quarto" or "large folio".
  • Translation Can Be Automated. One of the greatest obstacles to cross-national research is that item records are typically in the native language of the institution's host country or the primary language of the item's creator and/or scholarship and often have been edited over the years to add in material in multiple languages that the curator did not have the expertise to authoritatively translate or wished to preserve as-is. This means that a Dutch engraving published by a French publisher held by a German museum that represented a religious topic (Latin) and whom a major study was written by a Spanish scholar might have an item record might contain text in all five languages, plus likely excerpts in several more. The Louvre, for example, notes that its item records are updated daily and thus it cannot provide English translations. Some museums have invested heavily in translating their records, either from other languages into their host country's language or to expand their accessibility to an increasingly multilingual visitorship (for example, many US museums have begun adding Spanish translations to their gallery placards). The latest translation and domain-tuned translation LLM models have largely eliminated meaningful hallucination and can yield high-quality translation at scale. In our own work across more than 400 languages and a wealth of topics, we have observed consistently strong translations when using appropriate models and parameters. The end result is that museums can mass-translate their holdings with spot checking and monitor their web traffic for the most-accessed items to continually refine high-traffic translations.
  • The Failure Of Copy-Paste Item Records. I cannot begin to describe the number of items and institutions we've come across that over the last few decades simply copy-pasted their item records from the holdings of other institutions without bothering to check whether their own copy contained all of the described inclusions. This is especially problematic when it comes to maps, engravings, title pages, foldouts, coloring, etc, where we've traveled to an institution or had to pay a copy fee for material, only to find the institution's copy does not match their description. Worse, not one of the tier one institutions we've notified afterwards has actually corrected their catalog. This has only accelerated in the LLM era, with museums, universities, auctions and dealers alike using LLMs to write their item descriptions and failing to verify the generated description against their actual item. Remarkably, we've seen a huge surge in tier one institutions using LLMs to generate their item descriptions: a previously unthinkable development.
  • The Need For Specificity. Museums used to invest heavily in breaking apart and codifying all of the various characteristics of a work. For books, this meant splitting all of its relevant fields to different columns in a database. LLMs encode much of this knowledge now and can easily ingest it from authoritative sources, meaning a user searching for "the first edition of Dore's Bible" can find a museum record for "Bible. French. Bourassé-Janvier. 1866. La Sainte Bible: traduction nouvelle selon la Vulgate, par MM. J.-J. Bourassé et P. Janvier; dessins de Gustave Doré; ornementation du texte par H. Giacomelli. Tours: Alfred Mame et Fils, 1866. 2 vols., large folio (43–44 cm), 228 engraved plates." Instead of focusing on standardized metadata, institutions must refocus their efforts on describing what about their specific edition is DIFFERENT. For example, a copy with unusually large margins, a proof copy, "before the nails" or "before text", "first state of engraving 4", etc. The inclusion of a specific map and its completeness, coloration and its vintage, etc. For example, a researcher attempting to locate an institutional copy of Peter Force's American Archives to explore how his William Stone Declaration of Independence engraving was tipped in needs to know if Volume 5 contains the foldout, but many institutions fail to note whether their copy has it or not. Institutions need to focus on describing their own copy, not generating a standardized generic catalog entry.
  • The Automation Of Specificity. One of the most powerful aspects of modern AI models is that visual reasoning has advanced to the point where models can generate extremely detailed working understandings of a given copy of a work based purely on its scanned imagery. Asked for copies of Peter Force's American Archives that contain the William Stone Declaration of Independence engraving, modern LLMs know to navigate to Volume 5, know which page range to check and can report back whether it is present and even estimate its completeness and state – and do so at scale across every digitized copy worldwide. Similarly, scanning for marginalia, contemporary coloring or gilding, specific bindings, printing errors, even collation can all now be fully automated. Institutions can now generate extremely detailed item-specific records that describe all of the unique characteristics of their specific copy. Most importantly, since the models encode all of the key information about each work, they know specifically which things to look for, such as specific foldouts, colored plates, states, etc, that a typical curator or accession specialist might be unfamiliar with. However, it is important to note that AI models can only observe what they can see: if only a few pages of a book have been scanned, if the binding wasn't scanned or if the scans don't include a ruler or color bar (to assess size and color temperature of the lighting), the models can't measure those things. This suggests institutions should double down on digitizing their items, while making sure to include rulers and other information for context. It is worth noting that institutions that prefer to maintain human electronic record creation can still leverage the domain knowledge of items to use the LLM to prompt the records specialist with the list of frequently-missing items for a given work for them to verify their copy.
  • Moving From Vague Terms To Precise Attributes. Archives tend to describe books in terms of classical terms like "4to" or "quarto" or "folio" or "imperial", etc. While giving important information about their construction and a vague idea of size, these terms are highly imprecise: a "folio" could be anything from 10 inches tall to 40 inches tall. Worse, these terms are typically used incredibly loosely. Even terms of the trade like "blind tooled" that have specific meaning are less helpful when searching for a specific blind tooled pattern to identify a specific binding style. Again, providing AI models with high-quality scanned imagery can help with all of this.
  • Institutions Should Focus On Capturing What Can't Be Seen. A recent research project saw us looking for engravings and manuscript documents within specific size ranges. It turns out that few institutions record the size of their oversized manuscripts in inches and non-art institutions often don't record the size of their engravings or record only the expected size, rather than the size of their specific copy, making it impossible to locate outliers. AI models can discern any required visible attribute from scanned imagery, but they can't discern what they can't see. Institutions should focus on capturing things that aren't visible, such as size, material composition (if it can't be reliably determined from the image) and capture paper watermarks, multispectral imagery, xrays, etc for works where those are relevant. Codifying precise size, material and inclusions (for printed works) are perhaps the most important. Ideally, institutions should just attach all known information they have about an item to an "appendix" section of their records for AI models to read through to refine searches.
  • Customized Descriptions. Some institutions have invested heavily in creating different versions of their item records for different user constituencies, from expert scholars to schoolchildren. AI can automate much of this at scale, allowing institutions to expand these specialized item records and focus more on what kinds of information to include rather than the rote rewriting of creating them.