Developing and implementing a superconnector of producers in the printing industry to facilitate book historical research
Guido Thys, Antwerp Bibliophile Society, Antwerp, Belgium (guido@nextnext.nl)
Published in HT '23: 34th ACM Conference on Hypertext and Social Media · DOI: 10.1145/3603163.3609061 · License: © Copyright held by the owner/author(s).
Authors: Guido Thys
Session: Interactive Media: Art and Design
Abstract
The field of book history is on a head-on collision course with the requirements for the next phase in its digital development. This study first explores the notion of "digitalisation", the forthcoming phase, by drawing analogies from the forerunners in the Digital Revolution. The second section pinpoints the primary arenas where current practices could obstruct a seamless progression towards the ensuing phase. Subsequently, 'superconnectors' are introduced as potential mitigators of these barriers. A particular instance of such a superconnector is then delineated. This article lays the theoretical groundwork for TAPITTA, an exemplary superconnector that was demonstrated at the Hypertext 23 Conference in Rome, from Sep 4-8, 2023. Commencing Jan 1, 2024, it is accessible online at
hflps://tapifla.be."
Ccs Concepts
•Information systems~Data management systems~Information integration~Mediators and data integration
Guido Thys: Developing and implementing a superconnector of producers in the printing industry to facilitate book historical research. In 34th ACM Conference on Hypertext and Social Media (HT ’23), September 4–8, 2023, Rome, Italy. ACM, New York, NY, USA, 6 pages. hflps://doi.org/10.1145/3603163.3609061
1 Introduction
In the context of the Digital Revolution, the field of GLAM (Galleries, Libraries, Archives, and Museums) has not assumed a
leading position. This circumstance provides a singular opportunity to foresee the stages of its development with appreciable precision and to gain insights from the trials and resolutions of the innovators. A substantial parallel can be drawn between the Logistics sector's stocktaking practices on the one hand and data management within the Cultural Heritage domain on the other, with the initial concern in both spheres being the accessibility of data which were initially only available on paper (or microfilm) records. Within the Logistics sector, the progressive integration of automation, Internet technologies, and AI has facilitated a swift transformation from Inventory Taking to Enterprise Resource Management, culminating in Supply Chain Management. These transitions have been instrumental in advancing from mere data management to process efficiency, and subsequently, to an orientation centred on the user's perspective [1]. Similarly, the burgeoning interest in Linked Open Data (LOD) and related concepts indicates that the traditional practices of archiving and cataloguing heritage materials (digitisation) are slowly evolving to include the facilitation of research processes (digitalisation). Thus, Cultural Heritage Management finds itself on the brink of its Second Technological Revolution, with substantial potential to draw lessons from the developments that have transpired in the Logistics sector. This paper highlights the challenges as they are becoming apparent at present and their solutions, focussing on book history while a number of aspects apply to the GLAM area at large.
2 From digitisation to digitalisation
In general, digitisation may be characterized as "the material process of converting individual analogue streams of information into digital bits" [2]. In contrast, digitalisation emphasizes the utilization of this digitised information in various processes, with the objective of "increasing the operational efficiency, lowering the costs, enhancing the quality of the product, and, naturally, facilitating the simplification of people's lives" [3]. Research has demonstrated that digitalisation has the potential to augment the productivity of companies by a range of
44-52% [3]. It would be illogical to presuppose that this phenomenon, mutatis mutandis, would not be applicable to the sphere of GLAM and, more expansively, to the field of Digital Humanities as well, which is why digitisation will inevitably evolve into digitalisation.
The optimization of research is inherently a two-tier process.
Digitisation already enables academics to access metadata on primary sources as well as data from secondary sources online, instead of having to travel to multiple locations. But they still need to collate all the separate data into significant syntheses. Digitalisation aims to facilitate that process.
The growing interest in interconnectivity is a clear indication
of the onset of digitalisation. Even in this early stage, the obstacles encountered in establishing such interconnectedness point out a transformative shift in the objectives for which data should be captured: a transition from an accurate rendition of primary objects/texts and their metadata to the linkage and collation of data into coherent knowledge. This alteration in goal orientation and the correlated requirements are the basis of five major challenges that must be addressed. These challenges, in as far as not already recognised in the field, became apparent in the course of the TAPITTA project that will be discussed further on.
3 Challenges for digitisation
3.1 Standards that are not standards
Various standards, including but not limited to MARC and
IFLA, are extensively utilized by cataloguers to encode book metadata. However, these standards predominantly pertain to book properties (e.g., genre, type) and are often ambiguous, or at best, imprecise with respect to broader data such as producers and the place and date of publication.
It is therefore a prevalent practice in book cataloguing to
complement existing standards such as MARC with selfdesigned proprietary ontologies (e.g. of professions) and standards (e.g. for orthographic conventions).
These bespoke standards, in process theory referred to as
"local optima" [4], are a contradictio in terminis and poison pills in any consolidation project.
Indeed, while they may optimize within their specific domain
(catalogue), different bespoke standards will often prove to be incompatible with one another when integrated into a broader ecosystem.
3.2 Garbage in, garbage out
Another common practice is the processing of original data,
such as printers’ names, within -and limited to- the scope of “local” goals. An impressum might state e.g. Balthazar Moerentorff as the printer but the cataloguer might actualise and uniformise it to “Balthasar Moretus” to be consistent with, and internally linkable to other works by the same producer. This, however, prohibits any lexicographic and related research into imprints within this catalogue. It will also introduce incorrect and possibly incongruent data into a linked network of
catalogues. In Systems Theory and in Database Management this phenomenon is known as GIGO: “garbage in, garbage out”[5].
3.3 Incurable incompleteness
Books (and manuscripts) and their metadata are digitised in
roughly two different ways.
Metadata are stored (manually) into datasets, facilitating a
comprehensive spectrum of data management procedures, inclusive of both collation and aggregation.
Primary sources, on the other hand, are scanned and made
available as PDFs. Contemporary secondary sources are conventionally published in PDF format, while the majority of pre-existing secondary sources are (all but systematically) scanned, for obvious economic reasons. The inherent limitation of PDF files, permitting little beyond the mere reading of content, consequently renders substantial quantities of data inaccessible for integration within connectivity projects. Although emergent artificial intelligence text recognition technologies, exemplified by tools such as Transcribus, are projected to routinely transmute both paper and PDF documents into analysable datasets, the comprehensive inclusion of the entirety of accessible material remains a hope rather than a prospect.
Moreover, as far as secondary literature is concerned,
document conversion is a textbook case of mopping with the tap open: all research results, even concerning digitised data, is published in paper and/or pdf form and not as datasets. The reason: academic output ratings [6].
This has two negative consequences for the digitisation of
book history research processes. Firstly, datasets will remain conspicuously incomplete for an indeterminate future period. Secondly, the reconciliation of qualitative and quantitative research findings is subject to challenges that are similar to those encountered in the merging of qualitative and quantitative research results, known as Mixed Research Synthesis [7]. Within this context, manual labour emerges as the sole viable recourse.
So, paradoxically, the quality of digital knowledge ecosystems
will -at least provisionally- be contingent upon manual data entry.
3.4 Unverified and unverifiable
Metadata, when archived in catalogues, are consistently and
securely linked to their corresponding object, allowing for prompt verification. But data in other datasets are outlawed.
Within the academic domain, the omission of a reference in a
scholarly publication is viewed as a deadly sin. In stark contrast, datasets unrelated to metadata characteristically lack, almost axiomatically, any indication of the provenance of their data.
Consider Wikidata, the ultimate data repository and leading
supplier of Unique Resource Locators, that has a button “add reference” next to every entry. It is hardly ever used. An example: the entry “Balthasar Moretus” (see 3.2 above) lists four variants (of the twenty-two listed in TAPITTA, see below) with no reference to their origin. Consequently, the identification of these variants as either impressum data, actualizations, or
translations, is impossible. This is a severe case of data impoverishment.
Similarly, the CERL Thesaurus, “accessing the record of
Europe's book heritage” and widely used by libraries, lists six Moretus variants. While it associates the preferred "heading" with a source (in this instance, the Polish National Library, an entity scarcely regarded as authoritative on Antwerp printers), it fails to establish connections between sources and individual variants.
Even ISNI, the strictly curated International Standard Name
Identifier database, is inaccurate in this respect: it also lists sources without linking them to individual data.
Conclusion: adding these data as such is a GIGO operation.
3.5 Myopia
Digitisation is currently pre-eminently focussed on creating
digital catalogues and cataloguing can be defined as the maintenance of bibliographic records [10].
A rather important question in this respect is: does an
accurate entry of all metadata produce an accurate description of the object that goes beyond the registration of the features of, and entries in the book? On a philosophical level, some authors will posit that it doesn’t because one needs to take aspects like its creation into account [9]. But even on a elementary level it doesn’t. An example: printers in the 16th and 17th centuries invested a lot in obtaining privileges to print certain books because, once obtained, they yielded a fairly guaranteed income. When the printer died, his widow and or son(s) often continued the shop, still using the name of their spouse/father to safeguard the privilege. As a consequence, the fact that the printer is mentioned in the impressum -the only valid way of recording this metadata- does not mean that he did indeed printed the book. But standard recording practice does record it as such.
In order to verify this, one needs biographical information,
which can be found in church records and alike.
In order to digitalise this process, catalogues and
demographic databases need to be linked with some sort of ontology of names, by means of Unique Reference Indicators (URIs).
Much work has been done in this respect in terms of
Wikidata-URIs and alike. But what about other aspects, such as geographical metadata which would enable research on the exact location where the book was printed? None of the databases, repositories and catalogues pertaining to the printing industry in Antwerp, Belgium which will be discussed further on, has made provisions for this type of links. This type of myopia will severely restrict the viability of extensive ecosystems.
4 Superconnectors
Digitisation has increased the efficiency of book history
research considerably by diminishing/obliterating travel time between libraries (while increasing the risk of alienation from primary objects, but that is another discussion). Further
digitalisation of data gathering and analysis, which is the foreseeable next step, will have to significantly enhance the useability of the current body of digital material by collating data from different sources, many of them outside the “bookish” realm.
This, however, requires a very broad linkability of these
databases, the data of which should be readily verifiable. In short: a solution needs to be found for the challenges described in section 3.
Any viable solution would have to provide for the need of
standard, authoritative vocabularies, while allowing for elaborate linking of datasets and manual additions of non-digital data (print and PDF).
None of the existing datasets is suited for the task of
connecting different types of data in a more efficient and meaningful way nor is the manpower for elaborate curation available. Library catalogues aren’t, mainly because they have another, more restricted goal. Repositories such as Wikidata aren’t either, mainly because they lack curation and source information.
Fortunately, there is nihil novi sub sole: such entities have
already been developed in the context of the Semantic Web and in social media: ontologies and superconnectors.
4.1 Ontologies
In data science and the Semantic Web, ontologies are
structured representations of knowledge in a specific domain, allowing for interoperability among disparate systems and helping to overcome semantic obstacles when integrating data from different domains [11]. They typically provide controlled vocabularies to which data can be linked. An example on a very general level is the Resource Description Framework (RDF), developed by W3C, to provide a standard model for data interchange on the Web and facilitate data merging even when the underlying schemas differ [12].
However, such ontologies meet only one of the challenges:
they provide authoritative vocabularies, but they do not (necessarily) allow for the structured addition of varied data on their lemmas. However, such “content-rich ontologies” play an important role in Network Theory when applied to social networks,.
4.2 Hubs and superconnectors
Network Theory recognises that nodes with a significantly
high number of connections are vital to network dynamics [13]. A distinction is made between two types of such nodes: hubs that don’t add their own content, and superconnectors that do. An example of the former is Google while social media influencers belong to the second category: they constantly share their own information while linking to thousands or even millions other users.
The significance of "superconnectors" in social media was
further developed by Malcolm Gladwell [14].
In the realm of book historical research both hubs and
superconnectors could be created.
In areas where all metadata can be drawn from existing
datasets, “bookish” hubs could be constructed, e.g. for the books themselves: one single dataset containing all books, listing all libraries that keep them, shops that sell them, translations, summaries, etcetera.
In those areas where more data management is needed,
superconnectors could be created: authors, producers, subjects, printers’ addresses and alike.
One important requirement: such superconnectors would
have to be hybrid in nature: linking existing datasets while allowing for manual input of analog and scanned data.
4.3 Curation
Another but equally essential characteristic of these
superconnectors would be their curation facility. Indeed, when collating data large numbers of ambiguities, contradictions, errors and inconsistencies will appear: different birthdates, conflicting data on succession, fictitious addresses and may more.
Of course, systems are being developed to complete, debug
and repair ontologies [15], But, given the fact that “knowledge cleaning”[16] tasks (disambiguation, evaluation, completion and alike) require elaborate factual knowledge and historical research efforts, it is highly unlikely that they can be performed by machines unless, obviously, enabled by future AI advancements. At least for now, digitalisation will simplify but not completely replace manual research…
Moreover, a superconnector should respect the existence of
all (secondary) data, however incomplete or even incorrect they are: irrelevance cannot undo existence. Instead, all unusable data should be listed with their provenance and, whenever discovered, marked as such by the expert user.
In conclusion, the three characteristics of a functioning
superconnector are: linking, manual entry and curation.
5 TAPITTA: a sample superconnector
In order to prove the viability and the benefits of such
superconnectors, a first instance was created in the form of “The Antwerp Printing Industry Through The Ages”(TAPITTA) in the form of a Knowledge Graph (KG). A KG mainly:
describes real world entities and their interrelations, organized in a graph,
defines possible classes and relations of entities in a schema,
allows for potentially interrelating arbitrary entities with each other
and covers various topical domains.
It is therefore the most suitable form for the purpose at hand [17].
The TAPITTA graph was populated with data from an
OpenOffice relational database of manually entered data on producers (printers, publishers, illuminators, booksellers, etc.).
5.1 Technical aspects
Graph databases are designed to render relationships
between data in the most efficient way. Their use of triplets (node-relationship-node) instead of records in relational databases provides this flexibility [18].
Neo4j and its SQL-like programming language Cypher were
used as a platform. It is the technology underlying the work of the investigative journalists of the Panama Papers.
The back-office component of the project is Neo4j Workspace,
in which a schema was populated with all different type of data and their sources [see figure 1].
Figure 1: The TAPITTA schema in Neo4j’s Workspace (work in progress, status Aug 1st, 2023)
The UI of the front-office components is a Wordpress website
which also provides documentation, training and a membership function to distinguish between three roles: administrator with full access, editor with access to the read-only and the read/write component, and user with access to the read-only interface.
The latter interface enables consultation of all available data
of every producer individually, as well as of synthesised reports: activity in a certain year, all producers at any address or with a certain profession, etcetera. It was created in NeoDash, an opensource, low-code dashboard builder provided by Neo4j. Users can send additions and corrections to the administrator by email.
The read/write module is accessible to registered editors only
and provides full editorial access to all data and a data entry function. By lack of a NeoDash-like tool, is was developed from scratch in React, a JavaScript library for building UIs.
5.2 Nature and structure of the data
In the first phase of the project, data on roughly 3.700 producers in the printing industry of Antwerp, Belgium (historically in the Southern Netherlands) from 1481 up to today were imported from the original OpenOffice relational database. Every data element is linked to a full source reference.
The main challenge in this phase was the conversion of
relational database records into triplets, a task that took 1,000+ hours of spreadsheet manipulation.
The schema is divided into twelve data subsets: name
variants, professional activity, life events, career steps, addresses, family and succession data, schooling, production, additional literature, marks and examples. [see Figure 1.]. Additionally, separate ontologies have been created for - addresses: comprising 130,000+ houses, based upon the
authoritative database of the Flemish government, thus enabling future linking to any knowledge graph that uses the same basis: GIStorical maps, architectural heritage
descriptions, catalogues of archaeological finds and any data on inhabitants,
professions: a multilingual graph of 550 variants of 80+
professions and growing, facilitating even lexicographical research, e.g.: can a structure be found in the 14+ spelling variants of the Dutch "boekbinder"?
5.3 Digitalising research
From the onset, user-friendliness has been the primary focus of the KG design, not unlike the main driver behind the Supply Chain Management phase in Logistics. The UI renders data on different tabs and by means of a wide variety of visualisation tools [see Figure 2].
Figure 2: Sample screen print of TAPITTA’s r/o UI (work in progress, status Aug 1st, 2023)
The power of a KG in digitalising research processes can
further be demonstrated by answering very diverse questions by one buflon click. This includes questions as the following, that would otherwise require many hours, if not days of research:
what is the line of succession of a specific printer; when did his
heir(s) start using his name, and possibly also his privilege?
can a printer’s widow, through her maiden name, be linked to
other printers?
what did the genealogy of a producer look like? - which spelling variants of a name can be found in which
source?
In addition, users can request analyses that haven’t been produced yet.
5.4 Have the challenges been overcome?
Superconnectors have been proposed to meet the challenges arising from current digitising practices. Is TAPITTA capable of overcoming them? The answer is affirmative.
Standards that are not standards: ontologies have been set
up for all relevant components. They are authoritative by nature since they are (potentially) made up of all accessible datasets, including existing standards and local optima.
GIGO: all data are included and, when applicable, are or can
be qualified as incorrect, incomplete or uncertain. This disqualifies them as “garbage”.
Incompleteness: exhaustiveness is inversely proportionate
to abundance: profiles of thoroughly researched producers will never be complete. On the other hand, a KG-based ecosystem can theoretically expand forever (or at least until the original concept of the Semantic Web has been reached). However, by design, superconnectors like TAPITTA have the potential to be best in class.
Unverifiability: since a full reference to the source of all
data is mandatory, every component of the KG can be verified. Myopia: the additional ontologies of professions and addresses are the first instances of areas where the boundaries of book history have been crossed. Furthermore, the graph structure underlying TAPITTA is extremely flexible and allows for virtually any type of expansion.
6 Next steps
6.1 Linked Open Data and persistent URIs
As TAPITTA presently only consists of manually entered data, the next phase of the project, planned in 2024, will address Linked Open Data to connect to other datasets, such as heritage library catalogues.
This will further address the challenge of persistent URIs.
Indeed, all linkable data must labelled with a URI, as laid down in repositories such as CERL or Wikidata.
The GIGO-risk resides in the fact that, in the present state,
many of the URIs are not unique because they originate in non- or poorly curated ontologies. E.g. “cnp01951488” in CERL states that Hieronymus III Verdussen is a variant of Hieronymus II Verdussen, who is his father.
Secondly, “properties” (place and date of publication, printer,
collocation, etc.) need to be uniformly coded in order to make reconciliation possible. To this end, the International Federation of Library Associations (IFLA) has laid out a number of reference models. They haven’t been widely adopted yet since Linked Open Data has only relatively recently become a hot topic: most heritage libraries and catalogue managers are pondering their position and ambitions in this novel playing field. Additionally, IFLA models object-oriented: books, manuscripts, prints. Data and metadata about their producers remain very
much scaflered around in non-standardised fields in online catalogues, isolated repositories and, otien not (yet) digitalised, primary and secondary sources. Meeting this challenge will require desk research as well as field interviews with library cataloguers.
6.2 Scalability
TAPITTA has an open architecture, currently populated with data on producers from Antwerp. It is, however, designed to allow for expansion into other (heritage) areas and other territories.
There are no technical obstacles: the two UIs can easily be
litied from the proprietary website and the whole set-up, including the schema can even be copied while the existing data are erased to make room for other data.
Expanding the region to include the whole of Flanders,
Belgium, The Netherlands or even larger areas -within or outside the current dataset- would require no more work than what has been done for Antwerp: finding the right data and convert or otherwise enter them.
Contingent upon an adequate solution to the URI-problem,
data on e.g. other producers of heritage objects could be added, house-nodes could be complemented with data from Geographic Information Systems, etcetera.
On an even broader scale, this concept of superconnectors,
mutatis mutandis, can be applied to any area where KGs can create more value than the addition of their components. The cloud is the limit.
ACKNOWLEDGMENTS TAPITTA would not have been possible without the advice and contributions of dozens of people, including members of the Antwerp Bibliophile Society, the University of Antwerp, the Royal Library of Belgium, Memoo and many other institutions. I would like to extend a special token of gratitude to the kind folks at NeoDash and of Neo4j in general for allowing the free use of their technology and for their support “beyond the call of duty”.
REFERENCES
[1] Klaus, H., Rosemann, M., & Gable, G. G. (2000). What is ERP? Information
Systems Frontiers, 2(2), 141-162. doi:10.1023/A:1026543906354
[2] Brennen, J. S., & Kreiss, D. (2016). Digitalization. In K. B. Jensen, R. T. Craig, J.
D. Pooley, & E. W. Rothenbuhler (Eds.), The International Encyclopedia of Communication Theory and Philosophy (1st ed). John Wiley & Sons. doi:10.1002/9781118766804.wbiect165
[3] Zarnigor Akhmadalieva and Zulfizar Akhmadalieva. 2022. Impact of
digitalization on firms’ productivity. In The 6th International Conference on Future Networks & Distributed Systems (ICFNDS ’22),December 15, 2022, Tashkent, TAS, Uzbekistan. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3584202.3584254
[4] Goldratt, E. M., & Cox, J. (2004) The Goal: A Process of Ongoing Improvement
(third revised edition), North River Press, ISBN 978-0884271789.
[5] Redman, T. C. (1996). "Data Quality for the Information Age." Artech House,
Inc. ISBN 0-89006-883-1
[6] Terras, M. (2011). "Present, Not Voting: Digital Humanities in the Panopticon
Closing Plenary Speech, Digital Humanities 2010". Literary and Linguistic Computing, 26(3): 257-269.
[7] Sandelowski, M., Voils, C. I., & Barroso, J. (2006). Defining and Designing
Mixed Research Synthesis Studies. Research in the Schools, 13(1), 29.
[8] Chan, L. M., & Hodges, T. (2007). Cataloging and Classification: An
Introduction (3rd ed.). Scarecrow Press
[9] Coflrell, Barry (2013). Origins as Ontology in Printmaking (Paper submifled to
the Journal of Aesthetics and Art Criticism special Printmaking issue December 2013)
[10] Chan, L. M., & Hodges, T. (2007). Cataloging and Classification: An
[11] Smith, B., & Welty, C. (2001). "Ontology: Towards a New Synthesis". In FOIS
'01: Proceedings of the international conference on Formal Ontology in Information Systems (Vol. 2001, pp. 3-9). ACM.
[12] World Wide Web Consortium (W3C). (2014). RDF 1.1 Concepts and Abstract
Syntax. https://www.w3.org/TR/2014/REC-rdf11-concepts-20140225/
[13] Barabási, A.-L., & Albert, R. (1999). Emergence of scaling in random networks.
Science, 286(5439), 509-512.
[14] Gladwell, M. (2000). The Tipping Point: How Little Things Can Make a Big
Difference. Little, Brown.
[15] Lambrix, Patrick (2023). Completing and Debugging Ontologies: State of the
Art and Challenges in Repairing Ontologies. In: Journal of Data and Information Quality. https://doi.org/10.1145/3597304
[16] H. Paulheim (2017). Knowledge Graph Refinement: A Survey of Approaches
and Evaluation Methods. Semantic Web Volume 8 Issue 3 2017 pp 489-508 https://doi.org/10.3233/SW-160218
[17] Elwin Huaman and Dieter Fensel. 2021. Knowledge Graph Curation: A
Practical Framework. In The 10th International Joint Conference on Knowledge Graphs (IJCKG’21), December 6–8, 2021, Virtual Event, Thailand. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3502223.3502247
[18] Angles, R., & Gutierrez, C. (2008). Survey of graph database models. ACM
Computing Surveys (CSUR), 40(1), 1-39. doi:10.1145/1322432.1322433
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime