Showing posts with label information_retrieval. Show all posts
Showing posts with label information_retrieval. Show all posts

Wednesday, 28 November 2012

Using Wikipedia for extracting hierarchy and building geo-ontology

an article by Quoc-Hung Ngo (University of Information Technology, HoChiMinh City, Vietnam), Son Doan (University of California, San Diego, USA) and Werner Winiwarter (University of Vienna, Austria) published in International Journal of Web Information Systems, Volume 8 Issue 4 (2012)

Abstract

Purpose
This paper aims to serves two main purposes: First, it seeks to provide an overview of the location hierarchy from the highest divisions (continents) to the lowest divisions (wards, villages) in reality and in the Wikipedia pages. Secondly, it aims to introduce an approach to building a geographical ontology from Wikipedia.
Design/methodology/approach
The paper first reviews existing applications which extract information from Wikipedia and use it as a data resource to develop natural language processing tools. The paper also reviews the structure of Wikipedia pages which show the location’s information. Based on the analysis, the paper then proposes an approach to extract location hierarchy as well as geographical characteristics for the geo-ontology. The approach also rebuilds the relations between locations in the ontology.
Findings
Existing location name systems are mainly based on probabilistic locations, which are mined from the data and they lack the administrative relations between locations for full levels and all countries and territories. The literature review in geographical hierarchy and using Wikipedia for natural language processing tasks offers an approach to build a geographical ontology from Wikipedia pages. The proposed approach is believed to be the first which provides a full geo-ontology for all countries.
Practical implications
The paper builds a geo-ontology with full levels for all countries and territories. The administrative relations between locations are needed for real-world applications.
Originality/value
The comprehensive overview on existing work on geo-ontology provides a valuable reference for researchers and system developers in related research communities. The proposed approach to build a geographical ontology by using the Wikipedia offers a promising alternative to build a knowledge system from free online multi-language encyclopedia.

Hazel’s comment:
It sounds as though this would be a valuable tool in the hands of labour market researchers and also for guidance practitioners unsure of geography outside their own area.


Monday, 20 August 2012

A social inverted index for social-tagging-based information retrieval

an article by Kang-Pyo Lee, Hong-Gee Kim and Hyoung-Joo Kim (Seoul National University, South Korea) published in Journal of Information Science Volume 38 Number 4 (August 2012)

Abstract

Keywords have played an important role not only for searchers who formulate a query, but also for search engines that index documents and evaluate the query.

Recently, tags chosen by users to annotate web resources are gaining significance for improving information retrieval (IR) tasks, in that they can act as meaningful keywords bridging the gap between humans and machines.

One critical aspect of tagging (besides the tag and the resource) is the user (or tagger); there exists a ternary relationship among the tag, resource, and user. The traditional inverted index, however, does not consider the user aspect, and is based on the binary relationship between term and document.

In this paper we propose a social inverted index – a novel inverted index extended for social-tagging-based IR – that maintains a separate user sublist for each resource in a resource-posting list to contain each user’s various features as weights.

The social inverted index is different from the normal inverted index in that it regards each user as a unique person, rather than simply count the number of users, and highlights the value of a user who has participated in tagging. This extended structure facilitates the use of dynamic resource weights, which are expected to be more meaningful than simple user-frequency-based weights.

It also allows a flexible response to the conditional queries that are increasingly required in tag-based IR. Our experiments have shown that this user-considering indexing performs better in IR tasks than a normal inverted index with no user sublists.

The time and space overhead required for index construction and maintenance was also acceptable.


Wednesday, 11 April 2012

A new term-weighting scheme for naïve Bayes text categorization

an article by Marcelo Mendoza (Universidad Técnica Federico Santa María, Santiago, Chile) published in International Journal of Web Information Systems Volume 8 Issue 1 (2012)

Abstract

Purpose
Automatic text categorization has applications in several domains, for example e-mail spam detection, sexual content filtering, directory maintenance, and focused crawling, among others. Most information retrieval systems contain several components which use text categorization methods. One of the first text categorization methods was designed using a naïve Bayes representation of the text. Currently, a number of variations of naïve Bayes have been discussed. The purpose of this paper is to evaluate naïve Bayes approaches on text categorization introducing new competitive extensions to previous approaches.
Design/methodology/approach
The paper focuses on introducing a new Bayesian text categorization method based on an extension of the naïve Bayes approach. Some modifications to document representations are introduced based on the well-known BM25 text information retrieval method. The performance of the method is compared to several extensions of naïve Bayes using benchmark datasets designed for this purpose. The method is compared also to training-based methods such as support vector machines and logistic regression.
Findings
The proposed text categorizer outperforms state-of-the-art methods without introducing new computational costs. It also achieves performance results very similar to more complex methods based on criterion function optimization as support vector machines or logistic regression.
Practical implications
The proposed method scales well regarding the size of the collection involved. The presented results demonstrate the efficiency and effectiveness of the approach.
Originality/value
The paper introduces a novel naïve Bayes text categorization approach based on the well-known BM25 information retrieval model, which offers a set of good properties for this problem.


Monday, 7 November 2011

A Method to Assess Search Engine Results

an article by Judit Bar-Ilan (Bar-Ilan University) and Mark Levene (Birkbeck University of London) published in Online Information Review Volume 35 Issue 6

Abstract

Purpose
To develop a methodology for assessing search results retrieved from different sources.
Design/methodology/approach
This is a two-phase method, where in the first stage users select and rank the ten best search results from a randomly ordered set. In the second stage they are asked to choose the best pre-ranked result from a set of possibilities. This two-stage method allows the users to consider each search results separately (in the first stage) and to express their views on the rankings as a whole, as they were retrieved by the search provider. The method was tested in a user study that compared different country-specific search results of Google and Live search (now Bing). The users were Israelis and the search results came from six sources: Google Israel, Google.com, Google UK, Live Search Israel, Live Search US and Live Search UK. The users evaluated the results of nine pre-selected queries, created their own preferred ranking and picked the best ranking from the six sources.
Findings
The results indicate that the group of users in this study liked most the local Google interface, i.e. Google succeeded in it its country-specific customisation of search results. Live.com was much less successful in this aspect.
Research limitations/implications
Search engines are highly dynamic, thus the findings of the case study have to be viewed cautiously.
Originality/value
The main contribution of the paper is a two-phase methodology for comparing and evaluating search results from different sources.


Monday, 26 September 2011

You can’t build a car with just one wheel …

 (why duplication may not be such a bad thing), and some limitations of Internet search/retrieval

an article by David Zeitlyn published in First Monday Volume 16 Number 9 (September 2011)

Abstract

In this article I survey different approaches to the indexing of time based media (sound and video recordings) in response to two articles published in December 2010. Issues of overlap and duplication are discussed as a positive boon. They enable real comparison to be made and different styles may suit different categories of user. The lack of synoptic overview of different applications approaching the same topic is noted. Suggestions are made as to why search engines are not picking up on this. The failure leaves continuing role for human agents such as subject specialists and librarians to see the connections and make complete comparison lists of what is available for end-users.

Full Text: HTML


Thursday, 4 August 2011

A combined measure for representative information retrieval in enterprise information systems

an article by Baojun Ma, Qiang Wei and Guoqing Chen (Tsinghua University, Beijing) published in Journal of Enterprise Information Management (Volume 24 Issue 4 (2011))

Abstract

Purpose
The purpose of this paper is to propose a framework for describing and evaluating the representativeness of a small set of search results extracted from the original results: this is deemed desirable in information retrieval in enterprise information systems.
Design/methodology/approach
The paper proposes a combined measure, namely RFß, to evaluate the extracted small set in terms of the notions of coverage and redundancy. Data experiments were conducted on three different extraction strategies to evaluate the representativeness, i.e. coverage and redundancy.
Findings
Both from intuitive and experimental perspectives, the proposed coverage measure, redundancy measure and RFß measure could effectively evaluate the representativeness.
Research limitations/implications
The search results, e.g. in the form of documents and texts, are modeled using a vector space model and cosine similarity. Semantic models and linguistic models could be further introduced into this research to improve the proposed measures.
Practical implications
With the rapidly growing need for information retrieval in enterprise information systems, the representativeness of search results become more desirable and important for search engine users. The well-designed representativeness measures will help them achieve satisfactory results.
Originality/value
The originality of the paper lies in the definition of representativeness of a small set of search results extracted from the original results. This focuses on the two aspects of coverage rate and redundancy rate both from intuitive and experimental perspectives.


Saturday, 19 December 2009

Web searching by the “general public”: ...

an individual differences perspective

an article by Nigel Ford, Barry Eaglestone, Andrew Madden and Martin Whittle published in Journal of Documentation Volume 65 Issue 4 (2009)

Abstract

Purpose
The purpose of this paper is to explore the impact of a number of human individual differences on the web searching of a sample of the general public.
Design/methodology/approach
In total, 91 members of the general public performed 195 controlled searches. Search activity and ratings of search difficulty and success were recorded and statistically analysed. The study was exploratory, and sought to establish whether there is a prima facie case for further systematic investigation of the selection and combination of variables studied here.
Findings
Results revealed a number of interactions between individual differences, the use of different search strategies, and levels of perceived search difficulty and success. The findings also suggest that the open and closed nature of searches may affect these interactions. A conceptual model of these relationships is presented.
Practical implications
Better understanding of factors affecting searching may help one to develop more effective search support, whether in the form of personalised search interfaces and mechanisms, adaptive systems, training or help systems. However, the findings reveal a complexity and variability suggesting that there is little immediate prospect of developing any simple model capable of driving such systems.
Originality/value
There are several areas of this research that make it unique: the study’s focus on a sample of the general public; its use of search logs linked to personal data; its development of a novel search strategy classifier; its temporal modelling of how searches are transformed over time; and its illumination of four different types of experienced searcher, linked to different search behaviours and outcomes.


Saturday, 8 November 2008

Human information behaviour and design, development and evaluation of information retrieval systems

an article by Hamid Keshavarz in Program: electronic library and information systems Volume 42 Issue 4 (2008)


Abstract
Purpose
The purpose of this paper is to introduce the concept of human information behaviour and to explore the relationship between information behaviour of users and the existing approaches dominating design and evaluation of information retrieval (IR) systems and also to describe briefly new design and evaluation methods in which extensive attention is dedicated to the users and their behaviours and conditions.
Design/methodology/approach
The paper takes the form of a literature review with particular concentration on the efforts made by information science researchers.
Findings
The paper finds that there are four classic approaches to IR systems design: system-centred, user-centred, interactive and cognitive. Not enough research has been carried out to explore the relationship between information behaviour and information systems design to date. Contextual design and participatory design are among the new methods where users' behaviour, factors and contexts are considered more proactively than previously when designing information systems.
Originality/value
The paper introduces new methods and research frameworks being investigated currently in the area of information systems design and evaluation in which considerable attention is given to the users' information behaviour and situation. The paper is also useful in gaining a broad understanding about issues explored that have not previously been presented in one publication.

Thursday, 6 November 2008

An integrative model of “information visibility” and “information seeking” on the web

an article by Yazdan Mansourian, Nigel Ford, Sheila Webber and Andrew Madden in Program: electronic library and information systems Volume 42 Issue 4 (2008)

Abstract
Purpose
This paper aims to encapsulate the main procedure and key findings of a qualitative research on end-users' interactions with web-based search tools in order to demonstrate how the concept of “information visibility” emerged and how an integrative model of information visibility and information seeking on the web was constructed.
Design/methodology/approach
The study was formed of three parts. The first looked at conceptions of the Invisible Web; the second explored conceptualisations of the causes of search success/failure; the third organised the findings of parts 1 and 2 into a series of theoretical frameworks. Data collection was carried out in three phases based on interviews with a sample of biologists.
Findings
The first part led to the development of a model of information visibility which suggests a complementary definition for the Invisible Web. The results also showed the participants were aware of the possibility that they had missed some relevant information in their searches. However, perceptions of the importance and the volume of missed information varied, so users reacted differently to the possibility that they were missing information. The third part indicated the “Locus of Control” and “Attribution Theory” that can help us to better understand web-based information seeking patterns. Moreover, “Bounded Rationality” and “Satisficing Theory” supported the inductive findings and showed that users' estimates of the likely volume and importance of missed information affect their decision to persist in searching.
Research limitations/implications
The study creates new understanding of web users' information seeking behaviour which contributes to the theoretical basis of web search research. It also raises various questions within the context of library and information science practice to know whether, and if so how, we can assist end-users to develop more efficient search strategies and satisfactory approaches.
Originality/value
The research adopted a combination of inductive-deductive methods with a qualitative approach in the area of information seeking on the web which is mainly dominated by quantitative studies.


Hazel's comment:
I don't know about the other authors but you can keep up with Sheila's doings from her blogs on information literacy and working, teaching and living in Second Life.

Friday, 27 June 2008

To Search or to Research?

via Librarian of the Internet by findingDulcinea Staff on 28 May

The term "search" has taken on a different kind of life in recent years, thanks to our companion, the World Wide Web. While so many of us want to believe that Google is not a verb, we all know better. It's too easy; all you have to do is submit a query into the little box in and the world of information is opened up—it's instant gratification at its finest! But, does instant gratification come with a price? In the context of searching on the WWW, it certainly does.

Read the full article

Hazel's comment:
I was going to profusely apologise for the irregularity of posting here which resulted in this not being brought to your attention for a whole month. Gasp! Shock, horror!
But then I thought, what the heck, I've done it now and it's a good article and it's worth reading so maybe, just maybe, it's worth waiting for.
And then I realised that the theme here links in with my having been to a seminar yesterday organised by the International Society for Knowledge Organization on the "Agenda for Information Retrieval". And what a wonderful afternoon it was. Starting with Brian Vickery, who at 90 years of age has seen it all, on the "Issues in Information Retrieval" which was potted history going back some 70 years. Brian's very interesting approach was followed by Stephen Robertson on "The State of Information Retrieval: a researcher's view" which I must admit lost me a bit what with statistical probability theory and the rather stuffy room at UCL. A swift coffee woke me up enough to appreciate Ian Rowlands' talk on "The Google Generation" which the subsequent discussion decided was a myth.
It was great to catch up with Karen Blakeman and when Karen posts about passing a CV through a tag cloud generator I'll link you up to it. Fascinating discussion over the nibbles afterwards.
And as for sitting with Graham Robertson of Bracken Associates -- he thought it had been fifteen years since we'd last met at an ADSET seminar on information auditing. Although I think it was nearer ten years it was still a long time -- and we both still miss Peter Gillman who has left the information scene completely to concentrate on other work.