
What if you could have a conversation with a medieval sculpture? Ask what kind of wood it is made from, what the object it holds in its hand means, or where its gilding has gone? At The Metropolitan Museum of Art, this is now a reality. Since July 6, thanks to the Town Square AI project, four statues in the medieval gallery have been answering visitors’ questions, both orally and in writing, in English, French, Spanish, and Chinese.
The Metropolitan Museum of Art has just transformed its medieval gallery. Visitors can now move freely around the sculptures, come within a few inches of them, and observe them from every angle, with no glass case between them and the works.
This degree of proximity is exceptional. A medieval sculpture is not only discovered from the front: it is read in space, through a gesture, a fold of fabric, a raised hand, or a trace of almost-erased polychromy.
But this new freedom also reveals another difficulty. One can look at a sculpture for a long time without understanding what it is telling us. Why are two figures each holding a book? Why this crown? Why has the gilding disappeared? The clues are visible, but the codes that once made them legible are no longer familiar to most visitors.
It was around this challenge that The Met and Ask Mona designed Town Square AI: how can visitors be helped to look more closely, without making their phone the center of the experience?
The first choice was not technological. It was museographic.
Visitors do not speak with Mary Magdalene, Saint Augustine, or Balthazar. They speak with the sculptures that represent them.
This distinction allows each work to tell two stories at once.
The first is the story of the object itself: its material, how it was made, the traces left by time, its restoration, and its provenance.
The second is the story of the figure represented: their history, the symbols that allow them to be identified, and what they embodied for medieval visitors.
By adopting the point of view of the sculpture rather than that of the character, a single conversation brings these two dimensions together without ever opposing them.
Each answer was designed to bring the visitor’s gaze back to the sculpture.
When a visitor asks why a statue is holding a book or what a crown represents, the agent does more than provide an explanation. It invites the visitor to observe a specific detail: an attribute, a posture, a trace of gilding, or the way a garment is carved.
The experience thus becomes a guide for looking.
This principle lies at the heart of the entire experience. An answer from the conversational agent never replaces observation; it makes visitors want to return to the work to check, compare, or discover a detail they may have missed just seconds before.
The conversation does not stop with the sculpture being questioned. When there is a connection with a neighboring work, the agent invites the visitor to continue their journey through the gallery. The four sculptures thus become gateways into the rest of the medieval gallery.
This design choice led to a second challenge: how do you write the voice of sculptures representing saints without turning the conversation into religious discourse?
The work was carried out with the curatorial teams at The Metropolitan Museum of Art. The content is based on the museum’s scholarly research, documentation work, and the knowledge produced by its curators.
The experience therefore makes it possible to address both the making of a sculpture and the history of the figure it represents, its evolution, or its journey through the collections.
Some works illustrate this richness particularly well. The sculpture long presented as an “African prince” is now identified as Balthazar, one of the Magi. A conversation can now explain this evolution in scholarship, just as it can tell the story of the work’s journey before it entered the museum.
All the answers are based on content prepared and validated by the teams at The Metropolitan Museum of Art. For each response, the AI agent draws on this knowledge base to identify the right information to share.
It is a closed system that prevents the AI agent from sharing any information beyond what has been validated. In this way, the experience gives access, in conversational form, to the scholarly knowledge already produced by the museum.
Beyond this essential dimension of information control, we worked in depth on the tone of the conversation. This is what gives the experience its editorial richness: it is not a dialogue with a database, but an embodied conversation. The AI speaks differently depending on the statue with which the visitor is interacting. Each statue can therefore offer a different perspective on this period.
The ambition of the Town Square AI project is not so much to add a technological layer to the visit as to create the conditions for a more attentive gaze at the works.
Throughout the experience, the conversation remains in the service of what is happening in the gallery: the physical encounter with the sculptures, their presence, their materiality, and the traces of their history. It does not seek to replace the visitor’s gaze, but to accompany it, sometimes to slow it down, and to provide points of entry into a work whose codes are no longer immediately legible.
This is also what gives the experience its singularity: it does not produce the same experience for everyone. Each visitor can begin with what they notice, what surprises them, or what they do not yet understand. The conversation then becomes a personal space of exploration, where one moves forward through observation and questions.
The Met’s teams describe the ambition of the project as making Town Square AI “a complementary interpretation tool, in the service of curiosity, close observation, and access to the museum’s knowledge.”
This sentence neatly sums up what the experience seeks to make possible: an experience in which AI does not take the place of the artwork, but helps make the encounter with it more attentive, more accessible, and more profound.