an article by Nassim Abdeldjallal Otmani and Malik Si-Mohammed (Mouloud Mammeri University of Tizi Ouzou, Algeria) and Catherine Comparot and Pierre-Jean Charrel (University Toulouse – Jean Jaurès, Toulouse, France) International Journal of Web Information Systems Volume 15 Issue 3 (2019)
Abstract
Purpose
The purpose of this study is to propose a framework for extracting medical information from the Web using domain ontologies. Patient–Doctor conversations have become prevalent on the Web. For instance, solutions like HealthTap or AskTheDoctors allow patients to ask doctors health-related questions. However, most online health-care consumers still struggle to express their questions efficiently due mainly to the expert/layman language and knowledge discrepancy. Extracting information from these layman descriptions, which typically lack expert terminology, is challenging. This hinders the efficiency of the underlying applications such as information retrieval. Herein, an ontology-driven approach is proposed, which aims at extracting information from such sparse descriptions using a meta-model.
Design/methodology/approach
A meta-model is designed to bridge the gap between the vocabulary of the medical experts and the consumers of the health services. The meta-model is mapped with SNOMED-CT to access the comprehensive medical vocabulary, as well as with WordNet to improve the coverage of layman terms during information extraction. To assess the potential of the approach, an information extraction prototype based on syntactical patterns is implemented.
Findings
The evaluation of the approach on the gold standard corpus defined in Task1 of ShARe CLEF 2013 showed promising results, an F-score of 0.79 for recognizing medical concepts in real-life medical documents.
Originality/value
The originality of the proposed approach lies in the way information is extracted. The context defined through a meta-model proved to be efficient for the task of information extraction, especially from layman descriptions.
Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts
Friday, 30 August 2019
Monday, 30 July 2018
Twitter's vast metadata haul is a privacy nightmare for users
via ResearchBuzz Firehose
Metadata is everywhere.
Everything you tweet, every picture you take, and every status update you post on Facebook. It’s used by police and security forces to identify people who try to hide their identities and locations, while associated metadata in selfies can inadvertently ensnare criminals unaware that the data can destroy their alibi.
And metadata on Twitter can also be used in extremely precise identification each and every one of us – according to a new paper by researchers at University College London and the Alan Turing Institute.
Original article by Chris Stokel-Walker published in WIRED
Working with publicly available metadata from Twitter, a machine learning algorithm was able to identify users with 96.7 per cent accuracy
Conference paper (PDF 10pp)
Abstract
Metadata are associated to most of the information we produce in our daily interactions and communication in the digital world. Yet, surprisingly, metadata are often still categorized as non-sensitive. Indeed, in the past, researchers and practitioners have mainly focused on the problem of the identification of a user from the content of a message.
In this paper, we use Twitter as a case study to quantify the uniqueness of the association between metadata and user identity and to understand the effectiveness of potential obfuscation strategies. More specifically, we analyze atomic fields in the metadata and systematically combine them in an effort to classify new tweets as belonging to an account using different machine learning algorithms of increasing complexity.
We demonstrate that through the application of a supervised learning algorithm, we are able to identify any user in a group of 10,000 with approximately 96.7% accuracy. Moreover, if we broaden the scope of our search and consider the 10 most likely candidates we increase the accuracy of the model to 99.22%.
We also found that data obfuscation is hard and ineffective for this type of data: even after perturbing 60% of the training data, it is still possible to classify users with an accuracy higher than 95%.
These results have strong implications in terms of the design of metadata obfuscation strategies, for example for data set release, not only for Twitter, but, more generally, for most social media platforms.
Metadata is everywhere.
Everything you tweet, every picture you take, and every status update you post on Facebook. It’s used by police and security forces to identify people who try to hide their identities and locations, while associated metadata in selfies can inadvertently ensnare criminals unaware that the data can destroy their alibi.
And metadata on Twitter can also be used in extremely precise identification each and every one of us – according to a new paper by researchers at University College London and the Alan Turing Institute.
Original article by Chris Stokel-Walker published in WIRED
Working with publicly available metadata from Twitter, a machine learning algorithm was able to identify users with 96.7 per cent accuracy
Conference paper (PDF 10pp)
Abstract
Metadata are associated to most of the information we produce in our daily interactions and communication in the digital world. Yet, surprisingly, metadata are often still categorized as non-sensitive. Indeed, in the past, researchers and practitioners have mainly focused on the problem of the identification of a user from the content of a message.
In this paper, we use Twitter as a case study to quantify the uniqueness of the association between metadata and user identity and to understand the effectiveness of potential obfuscation strategies. More specifically, we analyze atomic fields in the metadata and systematically combine them in an effort to classify new tweets as belonging to an account using different machine learning algorithms of increasing complexity.
We demonstrate that through the application of a supervised learning algorithm, we are able to identify any user in a group of 10,000 with approximately 96.7% accuracy. Moreover, if we broaden the scope of our search and consider the 10 most likely candidates we increase the accuracy of the model to 99.22%.
We also found that data obfuscation is hard and ineffective for this type of data: even after perturbing 60% of the training data, it is still possible to classify users with an accuracy higher than 95%.
These results have strong implications in terms of the design of metadata obfuscation strategies, for example for data set release, not only for Twitter, but, more generally, for most social media platforms.
Monday, 22 July 2013
A Primer on Metadata: Separating Fact from Fiction
Ann Cavoukian, Information & Privacy Commissioner, Ontario, Canada
Since the recent revelations of the NSA’s sweeping surveillance of the public’s metadata, the term “metadata” has been regularly used in the media, frequently without any explanation of its meaning. Metadata’s reach can be extensive – including information that reveals the time and duration of a communication, the particular devices used, email addresses, or numbers contacted, which kinds of communications services were used, and at what geolocations. And since virtually every device we use has a unique identifying number, our communications and Internet activities may be linked and traced with relative ease – ultimately back to the individuals involved.
All this metadata is collected and retained by communications service providers for varying periods of time and, for legitimate business purposes. Key questions arise, however, including who else has access to all this information, and for what purposes? Senior U.S. government officials have been defending their sweeping and systemic seizure of the public’s personal communications on the basis that it is “only metadata.” They say it is neither sensitive nor privacy-invasive since it does not access any of the content contained in the associated communications.
A Primer on Metadata: Separating Fact from Fiction, explains that metadata can actually be more revealing than accessing the content of our communications. The paper aims to provide a clear understanding of metadata and disputes popular claims that the information being captured is neither sensitive, nor privacy-invasive, since it does not access any content. Given the implications for privacy and freedom, it is critical that we all question the dated, but ever-so prevalent either/or, zero-sum mindset to privacy vs. security. Instead, what is needed are proactive measures designed to provide for both security and privacy, in an accountable and transparent manner.
Full text (PDF 18pp)
Since the recent revelations of the NSA’s sweeping surveillance of the public’s metadata, the term “metadata” has been regularly used in the media, frequently without any explanation of its meaning. Metadata’s reach can be extensive – including information that reveals the time and duration of a communication, the particular devices used, email addresses, or numbers contacted, which kinds of communications services were used, and at what geolocations. And since virtually every device we use has a unique identifying number, our communications and Internet activities may be linked and traced with relative ease – ultimately back to the individuals involved.
All this metadata is collected and retained by communications service providers for varying periods of time and, for legitimate business purposes. Key questions arise, however, including who else has access to all this information, and for what purposes? Senior U.S. government officials have been defending their sweeping and systemic seizure of the public’s personal communications on the basis that it is “only metadata.” They say it is neither sensitive nor privacy-invasive since it does not access any of the content contained in the associated communications.
A Primer on Metadata: Separating Fact from Fiction, explains that metadata can actually be more revealing than accessing the content of our communications. The paper aims to provide a clear understanding of metadata and disputes popular claims that the information being captured is neither sensitive, nor privacy-invasive, since it does not access any content. Given the implications for privacy and freedom, it is critical that we all question the dated, but ever-so prevalent either/or, zero-sum mindset to privacy vs. security. Instead, what is needed are proactive measures designed to provide for both security and privacy, in an accountable and transparent manner.
Full text (PDF 18pp)
Friday, 16 November 2012
Using social bookmarks and tags as alternative indicators of journal content description
an article by Stefanie Haustein (Central Library at Forschungszentrum Jülich, Germany, and Heinrich Heine University Düsseldorf, Germany) and Isabella Peters (Heinrich Heine University Düsseldorf, Germany) published in First Monday Volume 17 Number 11 (November 2012)
Abstract
Qualitative journal evaluation cumulates content descriptions of single articles. Articles are either represented by author-generated keywords, professionally indexed subject headings, automatically extracted terms or, as recently introduced, by reader–generated tags as used in social bookmarking systems.
The study presented here shows that different types of keywords each reflect a different perspective on documents and that tags can be used in journal evaluation to represent a reader-specific view.
After providing a broad theoretical background and literature review, methods for extensive automatic term cleaning and calculation of term overlaps are introduced. The efficiency of tags and other metadata for journal content description is illustrated for one particular journal.
Full text (HTML)
Abstract
Qualitative journal evaluation cumulates content descriptions of single articles. Articles are either represented by author-generated keywords, professionally indexed subject headings, automatically extracted terms or, as recently introduced, by reader–generated tags as used in social bookmarking systems.
The study presented here shows that different types of keywords each reflect a different perspective on documents and that tags can be used in journal evaluation to represent a reader-specific view.
After providing a broad theoretical background and literature review, methods for extensive automatic term cleaning and calculation of term overlaps are introduced. The efficiency of tags and other metadata for journal content description is illustrated for one particular journal.
Full text (HTML)
Monday, 3 August 2009
Metadata interoperability in public sector information
an article by Lina Bountouri and Christos Papatheodorou (Ionian University, Corfu), Vasilis Soulikias (State General Archives of Greece) and Mathios Stratis (National Library of Greece) in Journal of Information Science Volume 35 Number 2 (2009)
Abstract
Over recent years, there has been a worldwide growing need for interoperability among the systems that manage and reuse public sector information. This paper explores the documentation needs for public sector information and focuses on metadata interoperability issues. The research work studies a variety of public sector information metadata standards and guidelines internationally accepted and presents two methodologies to obtain interoperability. The first develops an application profile, while the second is based on the semantic integration approach and results in the creation of an ontology. The outcomes of the two approaches are compared under the prism of their scope and usage in terms of interoperability during the metadata integration process.
Abstract
Over recent years, there has been a worldwide growing need for interoperability among the systems that manage and reuse public sector information. This paper explores the documentation needs for public sector information and focuses on metadata interoperability issues. The research work studies a variety of public sector information metadata standards and guidelines internationally accepted and presents two methodologies to obtain interoperability. The first develops an application profile, while the second is based on the semantic integration approach and results in the creation of an ontology. The outcomes of the two approaches are compared under the prism of their scope and usage in terms of interoperability during the metadata integration process.
Subscribe to:
Posts (Atom)