HyperBrain: Human-inspired Hypermedia Guidance using a Large Language Model
Authors: Danai Vachtsevanou, Jérémy Lemee, Raffael Rot, Simon Mayer, Andrei Ciortea, Ganesh Ramanathan
Published in: HT '23: 34th ACM Conference on Hypertext and Social Media · DOI: 10.1145/3603163.3609077 · License: © Copyright held by the owner/author(s).
HyperBrain: Human-inspired Hypermedia
Guidance using a Large Language Model
HyperBrain: Human-inspired Hypermedia
Guidance using a Large Language Model
Danai Vachtsevanou
University of St. Gallen
St. Gallen, Switzerland danai.vachtsevanou@unisg.ch
Jérémy Lemee University of St. Gallen
St. Gallen, Switzerland jeremy.lemee@unisg.ch
Raffael Rot University of St. Gallen
St. Gallen, Switzerland raffael.rot@student.unisg.ch
Simon Mayer University of St. Gallen
St. Gallen, Switzerland simon.mayer@unisg.ch
Andrei Ciortea University of St. Gallen
St. Gallen, Switzerland andrei.ciortea@unisg.ch
Ganesh Ramanathan
University of St. Gallen
St. Gallen, Switzerland ganesh.ramanathan@student.unisg.ch
Abstract
We present HyperBrain, a hypermedia client that autonomously navigates hypermedia environments to achieve user goals specified in natural language. To achieve this, the client makes use of a large language model to decide which of the available hypermedia con-trols should be used within a given application context. In a demon-strative scenario, we show the client’s ability to autonomously select and follow simple hyperlinks towards a high-level goal, suc-cessfully traversing the hypermedia structure of Wikipedia given only the markup of the respective resources. We show that hy-permedia navigation based on language models is effective, and propose that this should be considered as a step to create hyperme-dia environments that are used by autonomous clients alongside people.
Ccs Concepts
• Information systems →World Wide Web; • Human-centered computing →Interactive systems and tools; • Computing methodologies →Natural language processing.
Keywords
Hypermedia, Large Language Model, HATEOAS, Web of Things
Acm Reference Format:
Danai Vachtsevanou, Jérémy Lemee, Raffael Rot, Simon Mayer, Andrei Ciortea, and Ganesh Ramanathan. 2023. HyperBrain: Human-inspired Hy-permedia Guidance using a Large Language Model. In 34th ACM Conference on Hypertext and Social Media (HT ’23), September 4–8, 2023, Rome, Italy. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3603163.3609077
1 Introduction
In his article The Computer for the 21st Century, Mark Weiser opens with the statement that “The most profound technologies are those that disappear” [8]. Weiser conveys how truly impactful technolo-gies seamlessly integrate into our lives to the extent that we no
longer consciously recognize them as technology at all. This con-cept forms the foundation of ubiquitous computing, which proposes that computing devices could be present anywhere in the environ-ment and that humans could interact with these devices in a more natural fashion — famously, this might “make using a computer as refreshing as taking a walk in the woods.” [8].
Hypermedia systems can significantly contribute to achieving the vision of ubiquitous computing by providing users with in-tuitive, and flexible nonlinear access to an extensive range of in-formation and services distributed throughout the environment. However, as hypermedia systems such as the Web grow in complex-ity (even extending to the physical world 1), it becomes challenging for human users to navigate them efficiently without additional support. For this, tools and services such as search engines were developed to facilitate exploration by reducing the number of hy-perlinks users need to navigate, ideally providing the most relevant hyperlink based on a user goal, often expressed as a textual query. Nowadays, search engines, such as Google, use Large Language Models (LLMs) to better understand user queries and increase the relevance of the search results [1]. Search engines significantly assist users in discovering a relevant entry point upon starting their navigation. Yet, such guidance focuses on providing initial search results and does not seamlessly integrate throughout the remaining navigation process towards users’ goals.
Our guiding idea for this paper is that an LLM might be used to provide hypermedia guidance, that is, to indicate to a user which hy-permedia control to use during navigation and given the user’s goal when several controls are available. We argue that LLMs, with their capability to process text, such as user queries or hypertext, offer a promising solution to bridge the gap between users’ goals and the run-time state of hypermedia environments, providing valuable guidance for goal-driven hypermedia navigation. We propose the use of LLMs in hypermedia systems for automating the matching of application goals with dynamically discovered hypermedia in-terfaces. This proposal draws inspiration from how human users navigate hypermedia environments, performing a similar process to align their objectives with available interfaces [5].
1The Web of Things initiative, driven by the World Wide Web Consortium, defines a set of standards for the integration of physical devices in the Web: https://www.w3. org/WoT/.
Vachtsevanou et al. HT ’23, September 4–8, 2023, Rome, Italy
In the paper, following a short overview of background knowl-edge, we show that an LLM can effectively determine, when sup-plied with a set of hypermedia controls and the user’s goal, which hypermedia control is more likely to assist the user in pursuing their objective. We begin exploring the potential of our approach by demonstrating LLM-based hypermedia guidance in a well-known environment, rich in plain hypertext: Wikipedia. Here, the user goal of finding a specific Wikipedia page is represented as a single term, and we initialize our system with another Wikipedia page; the system then uses an LLM to navigate hyperlinks that are provided by Wikipedia to ultimately find the target page.
2 Background Knowledge
In the following, we offer an overview of how interaction takes place in hypermedia systems, identifying the key features required for LLM-based guidance that guides users from local hypermedia controls to global user goals2.
In hypermedia systems like the Web, users chain the usage of hypermedia controls, which can range from simple hyperlinks that are followed by clicking to more complex hypermedia forms that may require (sometimes validated) input. This chaining of inter-actions is rarely arbitrary or uninformed; instead, users typically navigate hypermedia with a clear sense of purpose, guided by the overarching direction — or global guidance — that is provided by their goal. For example, to buy a book online, a user first visits the Web page of a store where the book is sold, then adds the book to a (virtual) cart, possibly along with additional items, and finally proceeds to checkout. Inspired by how humans navigate hyperme-dia, we explore how LLMs can perform this goal-driven chaining of hypermedia controls. This involves interpreting the goal of a user and then predicting how the available hypermedia at an application state can help reach the state desired by the user, even in cases where multiple state transitions are needed in between.
Despite being driven by global user goals, interactions in hyper-media systems are also facilitated between state transitions. This more fine-grained local guidance is prevalent on the Web, where users encounter only a subset of hypermedia controls at a given ap-plication state, and this subset is updated after each interaction. For instance, when a user first visits the Web page of a bookstore, they may find interaction options to search for books or add books to the cart, while the option to checkout may not be available at that state. However, once the user interacts to add a book to the cart, a new set of hypermedia will be provided, including options to remove the book from the cart or proceed to checkout. Concretely, this local guidance is realized within the Web architecture through a principle that is part of the Uniform Interface constraint in the Representa-tional State Transfer architectural style [2], namely “Hypermedia as the Engine of Application State" (HATEOAS). HATEOAS constrains the design of Web servers, where a successful request should result in a response containing a new set of hypermedia that indicates the available interactions in the updated application state. Web clients, in turn, should effectively interpret this hypermedia and make requests to the server by exploiting the hypermedia controls and creating suitable input values in the correct format.
2We adopt the terms local and global guidance from [4] to explain how users’ interac-tions are guided in hypermedia systems.
Although HATEOAS effectively reduces the knowledge required on the client side to interact with servers, the induced local guidance is often independent of the user’s specific goal or based solely on an estimate of the user’s goal (e.g., in the case of Web services that link to videos or pictures that are recommended for a specific user). As a result, the user remains responsible for bridging the gap between HATEOAS-enabled local guidance and their global goal’s guidance upon making their next interaction. Here, our aim is to leverage LLMs to, not only fulfill the responsibilities of typical Web clients, but also mimic human user behavior, resulting in an enhanced form of local guidance that takes into consideration user goals. For the former, LLM-based clients should interpret and exploit hypermedia controls based on content and formats delivered by servers (such as in HTML documents [3], OpenAPI specifications [7], or W3C Web of Things Thing Descriptions [6]). For the latter, they should compute the best link option based on the available hypermedia (as provided through HATEOAS) and the estimated relevance to a goal (as provided by a user). Finally, keeping a record of past interactions may aid in avoiding loops in navigation, and constructing more contextualized input for interactions (e.g. so that the checkout process includes only the ID of the product that the user is currently interested in buying).
3 Hypermedia Guidance With Large Language Models
We propose to use LLMs for providing local guidance to machine clients in hypermedia environments. We argue that LLMs are well-suited to extract the implications of following a hypermedia control when given an appropriate context description (e.g., the label of a hyperlink in HTML).
3.1 HyperBrain: An LLM-Based Hypermedia Client
The main component of the system we present in this paper is HyperBrain3 — an LLM-based hypermedia client that is capable of making decisions about how to navigate hypermedia environments to reach a specified user goal.
A session with HyperBrain is initialized with a keyword in a textual format that represents a user goal (e.g. "information about Rome"), and an entry-point URI that identifies the current applica-tion context, i.e. the application state that determines the current availability of hypermedia (e.g. a Wikipedia page). With these two parameters, HyperBrain starts its navigation through the hyperme-dia environment, aiming to transition from the initial application state to a state that satisfies the user’s goal (e.g. the Wikipedia page about Rome). During navigation, HyperBrain can generate and submit queries to an LLM instance to support its decision-making process within specific application contexts. These queries serve the purpose of examining the hypermedia within the application context that is most likely related to the user’s global goal, and determining whether the application context satisfies the goal.
We evaluate our approach using the Wikipedia environment, demonstrating how our LLM-based hypermedia client effectively
3HyperBrain is implemented in Python and its code is available online: https://github. com/Interactions-HSG/hyperbrain
query = f " A v a i l a b l e l i n k : ' { entrypoint } ' " f " Guess i f the a v a i l a b l e l i n k i s the hypermedia r e f e r e n c e " " f o r { keyword } . " f " Return with TRUE/ FALSE . "
Code Listing 1: The client queries whether the current appli-cation state corresponds to the goal application state.
query = f " L i s t of keywords : ' { context } ' . " f " Guess which of these keywords i s most l i k e l y r e l a t e d " " to ' { keyword } ' . " f " Provide a keyword and one−sentence e xplanation . "
Code Listing 2: The client queries which of the keywords in the current context is likely to help transition towards the user’s goal.
Figure 1: Illustration of how the LLM-based hypermedia client, HyperBrain, navigates the Wikipedia environment.
navigates across Wikipedia pages, assisting users in finding in-formation about their topics of interest. Our system is not lim-ited to a specific LLM; however, during evaluation, we used the gpt-3.5-turbo and GPT-4 models (with the temperature hyperpa-rameter set to 0.9), as these models are capable of processing the large number of tokens commonly found in Wikipedia pages.
3.2 Scenario: Navigating Wikipedia
Wikipedia is a great testing environment for HyperBrain, since its entries commonly offer many hyperlinks accompanied by highly expressive context in natural language. At the same time, the diver-sity of topics in a Wikipedia entry poses an interesting challenge in
deciding on the most suitable hyperlink among all the hyperlinks available.
In the Wikipedia environment, each session with HyperBrain starts with the initialization of a keyword representing a topic of interest. This keyword indicates the Wikipedia entry to which the client needs to navigate. Additionally, an entry point, specified as the URI of the current Wikipedia entry, is provided. For our experiments, we selected “Lion” and “Genus” as keywords, and the locations of the Wikipedia entries for “Zurich” and “Rome” in the English Wikipedia4 as entry points.
Initially, the client performs a first query (as shown in Lst. 1), prompting the LLM to guess whether the entry point corresponds to the hypermedia reference that identifies the Wikipedia entry related to the given keyword. If the LLM considers that the entry point does not correspond to the keyword, the client proceeds by obtaining the context (i.e., the current page). This includes executing an HTTP GET request to retrieve the current Wikipedia entry in HTML format, and then extracting the hypermedia references (e.g. https://en.wikipedia.org/wiki/Italy) and their associated titles (e.g. “Italy”). Here, we limit the search to the first 200 available links within paragraphs, and ignore certain hypermedia references that do not relate to the topic documented in the Wikipedia entry (such as the hypermedia reference “/wiki/Wikipedia:About”). This decision derives from the structure and content of the Wikipedia environment (but it possibly limits the client’s effectiveness in other environments).
After obtaining the application context, another query is sub-mitted (as shown in Lst. 2), prompting the LLM to guess the most suitable hypertext for transitioning the application state closer (in terms of the Wikipedia hypergraph) to the Wikipedia entry associated with the keyword. The suggestions are accompanied by a textual explanation, providing insights into why a particu-lar hypertext is being recommended. HyperBrain then follows the recommended hyperlink. These steps (querying about goal achieve-ment and hyperlink recommendations, gathering the context, and following links) are repeated until the first query returns a result indicating that the visited Wikipedia entry likely corresponds to the given keyword. The behavior of the client is depicted in Fig. 1.
Given this algorithm and queries, HyperBrain is able to navigate the Wikipedia environment. Our results provide interesting insights: First, the path identified from “Zurich” to “Genus” is not straight-forward. The client navigates through the entries for “Turicum settlement”, “La Tène culture”, “Celts language”, and eventually, reaches “Genus” through “taxonomy biology”. In total, HyperBrain required five steps, which is not optimal since a shorter route could have been to traverse to “Zurich Zoo" and then, via any animal, arrive at “Genus”. However, the provided steps and explanations are reasonable. For instance, in the case of “Turicum settlement”, HyperBrain assumes that the word “Turicum” might refer to a plant or animal genus found in the area during Roman times (see Lst. 3). It hence mimics understanding of complex relationships and can search for different meanings of a word. Second, when attempting to reach “Lion” from “Rome”, the client follows an infinite navi-gation loop. It, first, traverses to “Roman mythology” because, as
4That is, https://en.wikipedia.org/wiki/Zurich and https://en.wikipedia.org/wiki/Rome
response = {
" answer " : False , " keyword " : " Turicum " , " explan ation " : " I t i s the Roman name f o r the c i t y of Zurich " " and could p o t e n t i a l l y r e f e r to a genus of " " p l a n t s or animals found in the area during " " Roman times . " } response = {
" answer " : False , " keyword " : " La Tene c u l t u r e " , " explan ation " : " The La Tene c u l t u r e i s a pre −Roman Iron Age " " c u l t u r e t h a t e x i s t e d from the 5 th to the 1 s t " " century BCE and i s known f o r i t s d i s t i n c t i v e " " a r t s t y l e , which included i n t r i c a t e geometric " " designs and s t y l i z e d animal motifs , as well " " as i t s advanced metalworking techniques . " } response = {
" answer " : False , " keyword " : " Continental C e l t i c " , " explan ation " : " I t r e f e r s to a group of an cien t C e l t i c " " languages and c u l t u r e s t h a t shared a common " " ancesto r . " } response = {
" answer " : False , " keyword " : " C e l t i c languages " , " explan ation " : " I t r e f e r s to a group of r e l a t e d languages " " belonging to the same language family or " " genus . " } response = {
" answer " : False , " keyword " : " Language family " , " explan ation " : "A language family i s a group of languages " " t h a t are r e l a t e d through a common anc est or " " language or proto −language , much l i k e how " " b i o l o g i c a l organisms are c l a s s i f i e d i n t o " " genera based on t h e i r e v o l u t i o n a r y h i s t o r y . " } response = {
" answer " : False , " keyword " : " Taxonomy ( biology ) " , " explan ation " : " Re fe rs to the branch of biology t h a t d e a l s " " with the c l a s s i f i c a t i o n of organisms i n t o " " groups or taxa , such as genus and s p e c i e s . " } response = {
" answer " : True , " keyword " : " Genus " }
Code Listing 3: HyperBrain output when tasked to navigate from “Zurich” to “Genus” on Wikipedia.
reported by the client, lions were often depicted in Roman mythol-ogy. Next, HyperBrain repeatedly transitions among “Hercules”, “Heracles”, “Greek mythology”, and “Hera”. Finally, HyperBrain fol-lowed the optimal path to navigate from “Zurich” to “Lion”, initially traversing to “Coat of Arms” and then to “Lion”.
3.3 Discussion and Limitations
Our demonstrator shows that LLMs can be valuable tools to as-sist goal-driven exploration of hypermedia environments. Through our experiments, we gained interesting insight on how to form queries that can effectively prompt the LLM to assist in hyperme-dia navigation. First, concealing the user’s exact intention may be beneficial, as when the intention is too strictly expressed, the LLM behaves as if the means to satisfy it are too limited, potentially leading to fewer suggestions. In our demonstrator, we preferred receiving a low-confidence suggestion to none at all. We tested different query forms with verbs like “return”, “provide”, “deliver”, and “get”, and found that “guess” yielded the best results when determining the goal application state and querying for hyperlink suggestions. The open-ended nature of guessing, combined with the phrasing used to ask for the "most likely related" keyword, led to the most suitable responses. Second, our experiments revealed that we can adequately restrict query responses to the keywords TRUE and FALSE (see Lst. 1), which can be used directly to evaluate
the condition for terminating navigation. Finally, requesting an explanation for the suggested keywords from HyperBrain (see Lst. 2) can improve its reliability and help the user to reach their goal application state more effectively.
Our hypermedia client has demonstrated its flexibility by provid-ing useful suggestions for navigation steps within the Wikipedia environment. However, certain implementation choices are tai-lored to the characteristics of this environment, which may limit the client’s effectiveness in more general hypermedia environments. Such choices concern processing HTML documents for ignoring a subset of the available hypermedia references, and the ability of the client to only execute HTTP GET requests. Additionally, the presented version of HyperBrain lacks support to dynamically con-struct input during interactions with a hypermedia environment5, and does not maintain a record of interactions for effectively avoid-ing loops in navigation. Finally, the effectiveness of the client can be hindered in cases where query results are too random or deviate from the expected structure. For example, if the LLM is unable to provide a suggestion because it cannot determine the relevance of the given keyword to the available hypertext, the client is unable to proceed with its navigation.
Further investigation is needed to validate LLM-based hyperme-dia guidance in various environments, including plain hypertext and more generic hypermedia settings, and for complex goals. We argue that LLMs effectively assisting navigation based on the appli-cation context and users’ goals would represent an advancement beyond typical search engine pages that require users to restart nav-igation each time they seek suggestions. Additionally, considering the abundance of models and formats for hypermedia-based inter-action descriptions, we argue that LLM-based hypermedia guidance can be effective in hybrid virtual and physical hypermedia environ-ments, where users may benefit from guidance upon interacting in more heterogeneous and large-scale settings. For instance, an LLM-based hypermedia client capable of interpreting W3C Web of Things Thing Descriptions [6] describing hypermedia interfaces of physical devices could potentially facilitate the more seamless navigation across Web-based physical environments.
4 Conclusion
In this paper, we introduced HyperBrain, an LLM-based hyperme-dia client that navigates hypermedia environments to achieve user goals, where an LLM provides local guidance considering users’ global goals. We demonstrated the effectiveness of our system in a scenario for exploring plain hypertext in the Wikipedia environ-ment. We discussed our observations on LLM-based hypermedia navigation, as well as the potential of this approach to facilitate interaction in hypermedia environments with resources that rep-resent both, virtual and physical resources. Overall, we conclude that the usage of an LLM to predict which hypermedia controls are most likely to support goal-driven hypermedia exploration is feasible and effective, and we propose that this might provide an effective interface between humans and software agents who use the Web together.
5A version of the LLM-based hypermedia client that supports the contextual generation and usage of input entries is available online: https://github.com/Interactions-HSG/ hyperbrain/blob/main/hyperbrain/objects/hyperbraincherrybot.py.
ACKNOWLEDGMENTS
This article emerged as one of the consequences of joint discussion by two working groups at the Dagstuhl Seminar on “Agents on the Web”6 held in February 2023, and from further explorations across seven research groups in different fields of Computer Science. We would like to thank the organizers of the Dagstuhl Seminar (Olivier Boissier, Andrei Ciortea, Andreas Harth, Alessandro Ricci) as well as the members of those working groups: Samuele Burattini, Brian Logan, Matthias Kovatsch, Sebastian Schmid, Andreas Harth, Daniel Schraudner, Rem Collier, Mahda Noura, Cleber Jorge Amaral, and Jean-Paul Calbimonte; it is a highly enjoyable, and a privilege, to work with you! This research was furthermore supported by the Swiss National Science Fund (HyperAgents, Project #189474), and the EU Horizon 2020 program (IntellIoT, Grant #957218).
References
[1] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Pra-fulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901. [2] R. T. Fielding and R. N. Taylor. 2002. Principled Design of the Modern Web
Architecture. ACM Transactions on Internet Technology 2 (2002), 115–150. Issue 2. [3] Larry M. Masinter and Daniel W. Connolly. 2000. The “text/html” Media Type.
RFC 2854. https://doi.org/10.17487/RFC2854 [4] Simon Mayer. 2014. Interacting with the Web of Things. Ph. D. Dissertation. ETH
Zurich. [5] Simon Mayer and David S. Karam. 2012. A Computational Space for the Web of
Things. In Proceedings of the 3rd International Workshop on the Web of Things (WoT 2012). Newcastle, UK. [6] Michael McCool, Ege Korkan, and Sebastian Käbisch. 2023. Web of Things (WoT) Thing Description 1.1. W3C Proposed Recommendation. W3C. https://www.w3.org/TR/2023/PR-wot-thing-description11-20230711/. [7] Darrel Miller, Jeremy Whitlock, Marsh Gardiner, Mike Ralphson, Ron Ratovsky,
and Uri Sarid. 2021. OpenAPI Specification v3.1.0. Technical Report. OpenAPI Initiative. https://spec.openapis.org/oas/v3.1.0. [8] Mark Weiser. 1999. The computer for the 21st century. ACM SIGMOBILE mobile
computing and communications review 3, 3 (1999), 3–11.
6https://www.dagstuhl.de/en/seminars/seminar-calendar/seminar-details/23081
Figures and table reproductions
Figure 1: source-region reproduction from the supplied PDF.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime