Publicité
Metadata and standards
(Part IV)
Metadata is ?structured information that describes, explains, locates, or otherwise makes it easier to retrieve, use, or manage an information resource? (National Information Standards Organisation, 2004). More simply put, metadata is ?information about data? (Ianella and Waught, 1997). One common example of metadata is our traditional cataloguing records. Metadata is so much an important factor of information retrieval because it helps in discovery, identification and retrieval of information.
As said by the National Information Standards Organisation (2004), metadata ?is key to ensuring that resources will survive and continue to be accessible into the future?. Metadata can improve retrieval in the sense that ?appropriate metadata tags around the different data elements allow search engines to seek information in a more discriminating way? (Haynes, 2004).
As said by Milstead and Feldman (1999),?Used effectively, it makes information accessible by labelling its contents consistently. Metadata leaves a pathway for users to follow to find the information they need all in one place. In invisible cyberspace, this is even more important than in a library where desperate users at least have shelves to browse.?
Unlike metadata in traditional information retrieval tools, ?metadata for the Internet is an extremely complex issue? (Thomas and Griffin, 1998). As we are all aware of, data in library catalogues have for long consisted of highly standardised descriptions based upon set down rules but in the networked environment, it remains a challenge.
Metadata would not exist without standards, that is, the actual rules themselves. ?Metadata cannot fully serve its purpose unless it is subjected to a certain amount of standardisation? (Milstead and Feldman, 1999). With the advent of the World Wide Web, it seems that standards are being to some extent neglected. Standards provide for consistency in order to produce effective information retrieval tools for users. Well-known traditional standards include the International Standard Bibliographic Description and the Anglo-American Cataloguing Rules that have been used for decades in libraries. Standards in an online environment include the Z39.50 and Dublin Core amongst others. As it can be seen, we are not short of standards but rather there is some reluctance in using standards on the part of search engines. As said by Gill (n.d), ?the Lawrence and Giles 1999 survey, for example, found that only 0.3 % of Web sites contained Dublin Core metadata? and as said by Brin and Page (1998), ??it is interesting to note that metadata efforts have largely failed with Web search engines?. In the future, with the growth of the Web, if metadata and standards are not being used, information retrieval will become extremely difficult.
With the advent of the Web, more and more documents are including multimedia elements and to cater for these more and more sophisticated tools are required. As said by Angelides, Dustdar and Lesk (1997) in Jansen, Goodrum and Spink (2000), ?the World Wide Web is an immense repository of multimedia information.? In the traditional world of library catalogues, we did not need to retrieve multimedia elements like music or live interviews and therefore, this is a new feature requiring new information retrieval techniques.
Multimedia information retrieval
Multimedia includes audio, video and image and therefore, means and ways should be sought in order to effectively retrieve them. In fact, ?Web users must search for multimedia information as they would search for textual information? (Schauble 1997 in Jansen, Goodrum and Spink, 2000) and most Web information retrieval systems, Yahoo being an example, would not provide for any other special mechanisms for multimedia searching. Fortunately, some like AltaVista or Lycos provide for special features for those searching for multimedia. Such features would include radio boxes and media specific search syntax. Others yet provide for multimedia searching in MP3 format.
Most literatures on multimedia information retrieval refer to Content-based image retrieval. This relies upon the ?characterization of primitive features such as color, shape and texture that can be automatically extracted from the images themselves? (Goodrum, 2000).
Here, queries can be submitted by simply either drawing a sketch or clicking on a texture palette or by selecting a particular shape of interest. However, the problem does not end here. The main problem lies in the areas of indexing, assignment of terms, vocabulary control and user needs amongst others. As pointed out by Ruthven (1999), ?the diversity and complexity of interlinked multimedia data offered in many applications required new indexing methods and an enhanced retrieval functionality as well as new interaction pardons and metaphors as compared to traditional IR systems.?
Taking the film Titanic by James Cameron for example. If a user wants all the live interviews of James Cameron on the film plus the pictures, what keywords should the user enter? Would it be ?Titanic?, ?James Cameron?, ?live interviews? or ?pictures? or a combination of those. ?Pictures? have no words and we wonder how the keywords would do the match. As clearly pointed out by Goodrum (2000),?Assignment of terms to describe images is not solved entirely by the use of controlled vocabularies or classification schemes however. The textual representation of images is problematic because images convey information relating to what is actually depicted in the image as well as what the image is about.?
Moreover, it has been noted by Foote (1997) that, ?Though powerful IR algorithms are available for text, it is clear that for audio, or multimedia in general, common term matching approaches are useless due to the simple lack of identifiable words (or comparable entities) in audio documents. The problem becomes even more open-ended when one considers audio, such as music, which may have no speech.?
These however remain questions that are still to be answered and solutions found. As we can see, information retrieval on the Web is not as easy and straightforward as retrieving information from a catalogue.
Tara Héléna LAM
Publicité
Publicité
Les plus récents