This dataset provides a systematic evaluation of twelve multimodal large language models on the reading, analysis and interpretation of statistical maps. Each model answered nine questions about each of sixteen statistical maps of Europe in batch mode, with all nine questions submitted in a single prompt; four of them repeated the full set in iterative mode, one question at a time in a continuing conversation. Together this yields 2,304 records collected in March 2025.
Each record contains the question, the verbatim model response, fourteen sub-criterion scores grouped into four criteria (correctness and relevance; organization and conciseness; data support; response justification), a composite Response Quality Index on a 0-5 scale, and metadata describing the model (provider country, access tier), the map (cartographic methods, symbol scaling, number of methods combined, symbol structure, spatial level, source) and the question.
The design supports comparison across nine interaction primitives of map use: identify, locate and retrieve value (reading); compare, cluster and associate (analysis); interpret, cause/effect and predict (interpretation). All maps depict Europe using NUTS enumeration units and are in English or bilingual English-Polish.
The dataset can serve as a benchmark for evaluating later model generations against a fixed set of cartographic tasks, as material for work on automated scoring of open-ended responses, and as a reference point for human-model comparisons in map use research.
(2025-12)