THE ADDITION OF DATA — SOCIOPLASTICS 9589 — Anto Lloveras — LAPIEZA-LAB, Transdisciplinary Research Laboratory — Madrid, 2026



Contemporary institutions often answer uncertainty by adding data. More sensors, documents, images, citations, behavioural records and model parameters are expected to produce a more faithful account of reality, yet accumulation eventually changes the character of the problem. Beyond a certain threshold, information scarcity gives way to difficulty of orientation: the decisive questions become how records are classified, weighted, connected, updated and made available for comparison. Torrecilla and Romo argue that large quantities do not remove the need for statistical reasoning, while Klein and D’Ignazio show that data practices must examine power, context, labour, environmental impact and consent rather than treating scale as an automatic path to better knowledge. Birhane, Prabhu and Kahembwe demonstrate the consequences of internet-scale collection when malignant stereotypes, explicit content and weak curation become embedded in multimodal datasets. These works shift attention from the romance of abundance toward the architecture of selection. A city reconstructed through mobile-phone transactions will privilege movement that leaves a commercial trace; a cultural field mapped through citations will reproduce institutions already recognised by citation databases; a health system trained on documented treatment may overlook people whose exclusion prevented documentation in the first place. Data does not abolish selection. It relocates selection into collection protocols, schemas, interfaces, missing values and decisions over which phenomena count as comparable. RecurrenceMass becomes useful when repeated records acquire enough density to reveal patterns that no isolated item contains, but recurrence must not be mistaken for truth. What happens frequently may be an effect of the apparatus that observes it, the behaviour encouraged by an existing policy or the accumulated consequence of inequality. The relevant advance is therefore not simply a larger dataset but a better organised one: provenance attached to records, categories open to revision, scales kept distinct, anomalies retained rather than cleaned away, and routes of inference that can be reconstructed. At ten records, a collection may be read individually; at ten thousand, architecture becomes unavoidable. Coordinates, identifiers, relations and internal partitions determine whether density produces knowledge or opacity. The ideal corpus does not attempt to absorb the whole world. It describes its exclusions, preserves the difference between evidence and interpretation, and allows new questions to reorganise the material without destroying earlier arrangements. Data acquires intellectual force through structured difference, not through quantity alone.



Klein, L. and D’Ignazio, C. (2024). Data Feminism for AI.
Torrecilla, J. L. and Romo, J. (2018). Data Learning from Big Data.
Birhane, A., Prabhu, V. U. and Kahembwe, E. (2021). Multimodal Datasets.
Gitelman, L. (ed.) (2013). Raw Data Is an Oxymoron. MIT Press.