Publicité

Information retrieval: From the traditional way to the Web

4 avril 2005, 00:00

Par

Partager cet article

Facebook X WhatsApp

lexpress.mu | Toute l'actualité de l'île Maurice en temps réel.

<B>(Part I) </B>

As noted by Martzoukou (2005), ?the Web has grown into a vital channel of communication and an important vehicle for information dissemination and retrieval.? Whilst more than one would be persuaded that information retrieval on the Web is easy and at hand with a simple click of the mouse, still, the situation is not as bright. Infor-mation retrieval on the Web is far more difficult and different from other forms of information retrieval than one would believe since Web pages are voluminous, heterogeneous and dynamic with a lot of hyperlinks including useless information such as spams, have more than a dozen search engines, make use of free text searching, are mostly without metadata and standards, and contain multimedia elements.

The principles of searching a catalogue are different from that of a WWW and the WWW is devoid of cataloguing and classification as opposed to a library system. Even the evaluation of information retrieval on the Web is different. All these points contribute to make information retrieval on the Web different from say, a library catalogue or even an electronic database.

?Information retrieval deals with the representation, storage, organization of and access to information items. The representation and organization of the information items should provide the user with easy access to the information in which he is interested? (Baeza-Yates and Ribeiro-Neto, 1999). Infor-mation retrieval on the Web differs greatly from other forms of information retrieval and this because of the very nature of the Web itself and its documents. Indeed, ?the WWW has created a revolution in the accessibility of information? (NISO, 2004). It goes without saying that the Internet is growing very rapidly and it will be difficult to search for the required information in this gigantic digital library.

● <B> Information overload</B>

Information overload can be defined as ?the inability to extract needed knowledge from an immense quantity of information for one of many reasons? (Nelson, 1997). One of the main points that would differentiate information retrieval on the Web from other forms of information retrieval is information overload since the Web would represent items throughout the universe. ?The volume of information on the Internet creates more problems than just trying to search an immense collection of data for a small and specific set of knowledge? (Nelson, 1997). An OPAC will only represent the items of a specific library or a union catalogue will only represent items of participating libraries within a region or country.

<B><I>?The volume of information on the Internet creates more problems than just trying to search an immense collection of data for a small and specific set of knowledge.?</I></B>

It is known that large volumes of data, especially uncontrolled data, are full of errors, inconsistencies and uselessness. As said by Nelson (1997), ?when we try to retrieve or search for information, we often get conflicting information or information which we do not want?. Also, ?finding authoritative information on the Web is a challenging problem? (Savoy in Baeza-Yates and Schauble, 2002) as opposed to a library where we would get mostly authoritative information.

Information overload on the Web makes retrieval become a more laborious task for the user. The latter will either have to do several searches or refine searches before arriving at the relevant and accurate document. Therefore, as opposed to a search in a catalogue, with the Web, there is no instantaneous response as such. There is no instantaneous response in the sense that there are no instant relevant results and this is not to be mixed up with instantaneous response in the sense that it is true that whilst we tap in keywords, we do get a thousand of hits. With the advent of the Web, users will need easier access to the thousands of resources that are available but yet hard to find.

● <B> Hyperlinks</B>

?With the advent of the Web new sources of information became available, one of them being hyperlinks between documents?? (Henzinger, 2000). In an OPAC for example, one would not find hyperlinks. The Web would be the only place to find hyperlinks and this brings a new dimension to information retrieval. With hyperlinks, one information leads to another. If the user were viewing a library catalogue, he/she would be directed to let?s say a book and that?s the end to it.

But if that same user does a search on the Web, he/she would be presented with a thousand of hits, he/she would then need to choose what to view and then the Web pages viewed would refer the user to yet other pages and so on. As pointed out by Henzinger (2000), ?hyperlinks provide a valuable source of information for Web information retrieval?? Hyperlinks are thus providing us with more or even newer information retrieval capabilities.

● <B> Heterogeneous nature of the Web</B>

?A Web page typically contains various types of materials that are not related to the topic of the Web page? (Yu et. al., 2003). As such, the heterogeneous nature of the Web affects information retrieval. Most of the Web pages would consist of multiple topics and parts such as pictures, animations, logos, advertisements and other such links. ?Although traditional documents also often have multiple topics, they are less diverse so that the impact on retrieval performance is smaller? (Yu et. al., 2003). For instance, whilst searching in an OPAC, one won?t find any animations or pictures interfering with the search.

Documents on the Web are presented in a variety of formats as opposed to catalogues and databases. We have HTML, pdf, MP3, text formats, etc. The latter can be barriers to information retrieval on the Web. For one to retrieve a pdf document on the Web, one must have Acrobat Reader software installed and enough space on the computer?s hard disk to install it.

<B>Tara HÉLÉNA LAM

Publicité