CLARIAH Summer School 2026
Four days at the CLARIAH summer school in Amsterdam: SPARQL, RDF triples, and turning a Chilean documentary archive into linked open data.
The 2026 edition of the CLARIAH summer school has just finished, and having now done it twice (I took part in 2025 as well), I can say it is a course that continues to evolve. It brings humanities scholars into the world of cultural data through well-chosen entry points, or strands, that feel increasingly necessary for the academic study of the arts and humanities. Rather than a task to hand over to technicians, producing and safeguarding the quality of cultural data is fast becoming a fundamental edge that enables these disciplines to contribute to broader society, and CLARIAH is one of the places actually doing that work.
Last year, I followed the audiovisual collections strand, learning to use the Media Suite, the tool that CLARIAH has built for exploring audiovisual collections held by Dutch institutions associated with this research initiative. This year I was in the linked open data strand, learning to query with SPARQL and to model data as RDF triples, the subject-predicate-object statements that linked open data is built from. Technical stuff, yes, but it clarified something that, however unglamorous, matters a great deal: how to produce good data and keep it usable far into the future, following the principles now on every arts and humanities scholar’s lips: that data should be findable, accessible, interoperable, and reusable, rather than merely stored.

The archive lifecycle
What stayed with me was how the school, through its three strands, tries to hold the whole lifecycle of a digital cultural archive in view at once.
It begins with capture, which one can’t separate from ethics: asking permission, securing long-term legal arrangements that keep material accessible, and thinking through whose voices are included and on what terms. Then comes archiving and structuring, where the linked open data work lives, turning annotations into RDF triples and building databases that can talk to other open datasets. Finally, reuse is where researchers navigate, annotate, and make things with the collections through a specialised graphical interface.
Each strand went deep into one of these stages, while the plenary sessions kept pulling the tracks back together, a reminder that these are not separate problems but a single continuous piece of scholarly work.
Learning in real time
Summer school was long days, and no matter how much coffee we could squeeze from the machines, it was sometimes hard to stay sharp. SPARQL queries and RDF structures were new to me, and they took some time to digest, but by the end, something had clicked. I now understand how these databases actually work, how unstructured material gets shaped into something queryable, which archives can be shared, and why that matters. There is plenty left to learn, but I am happy with what I managed to assimilate. I left with the particular clarity that only comes from wrestling and grappling with the unfamiliar.
Mapping MAFI
For the exercises, I worked with a sample of the MAFI collection, the Chilean documentary archive I helped start back in 2010 and whose artistic afterlife I’m now working on. That made the abstractions concrete: the triples I was learning to write described films I already knew. Over the days, I built a small set of data stories inside the Institute’s system, turning the sample into things you could read at a glance. Some were plain infographics drawn from the MAFI data, with the archive represented as cinemetrics, such as the distribution of shot scales across the shorts.

The one graph I kept returning to combined two datasets into a single map: the places where MAFI films were shot and every Chilean city with more than 100,000 inhabitants, the latter pulled from DBpedia. Click a point, and a card opens with the relevant details: a film’s title and director or a city’s population. Two archives that had never met joined in a single view by a few lines of SPARQL. It is a small thing, but it is the exercise the strand was aiming at: making archives move from being closed boxes to start talking to the rest of the open web.
Why CLARIAH matters
The thing I came away holding from this summer camp is that data stewardship for cultural archives is serious scholarly work, not speculation or infrastructure nobody asked for. While the humanities don’t chase breakthroughs in computational efficiency, they do care, deeply, about how technology shapes what gets remembered, who gets heard, and what futures we are quietly building into our tools. CLARIAH takes that seriously. It puts rigorous technical work, ethical reflection, and humanistic purpose in the same room and refuses to treat any of them as optional.
Data stewardship for cultural archives is serious scholarly work, not speculation or infrastructure nobody asked for.
If you’re curious about how cultural memory gets annotated, how archives can be reused for exhibitions or artistic work, or how humanistic scholarship can engage with data and infrastructure on its own terms, CLARIAH events, and future editions of this summer school are a great place to start.