Showing posts with label data_mining. Show all posts
Showing posts with label data_mining. Show all posts

Tuesday, 7 August 2018

International comparison of active citizenship by using Twitter data, the case of England and the Netherlands

an article by Cristina Rosales Sánchez (European Commission, Italy; Wageningen University and Research Centre, The Netherlands) published in First Monday Volume 23 Number 6 (June 2018)

Abstract

Can social media become the new data source for certain social indicators?

What does social media offer in comparison with classical sources for official statistics?

This research analyzes the potential of Twitter data to obtain indicators of active citizenship. We use a methodology developed for a case study in Spain and test its replicability by applying it to England and the Netherlands.

Twitter data offers an advantage to assess change in active citizenship over time and at high spatial resolution, while survey data — traditional sources for official statistics — are more authoritative and robust. Collecting and analyzing Twitter data is also faster than traditional survey processes.

However, we found that Twitter usage specificities in each country ease or hinder the process, affecting the time to secure an indicator. Whilst future research will focus on other social indicators and sources of data, this research is an important step in developing an evidence-based understanding of the strengths and weakness of social media to inform policies.

Full text (HTML)


Tuesday, 17 January 2017

An efficient method for discovery of large item sets

an article by Deepa S. Deshpande (MGM's Jawaharlal Nehru Engineering College, Aurangabad, India) published in International Journal of Data Mining, Modelling and Management Volume 8 Number 4 (2016)

Abstract

In today's emerging field of descriptive data mining, association rule mining (ARM) has been proven helpful to describe essential characteristics of data from large databases.

Mining frequent item sets is the fundamental task of ARM. Apriori, the most influential traditional ARM algorithm, adopts iterative search strategy for frequent item set generation. But, multiple scans of database, candidate item set generation and large load of system's I/O are major abuses which degrade the mining performance of it.

Therefore, we proposed a new method for mining frequent item sets which overcomes these shortcomings.

It judges the importance of occurrence of an item set by counting present and absent count of an individual item. Performance evaluation with Apriori algorithm shows that proposed method is more efficient as it finds fewer items in frequent item set in 50% less time without backtracking.

It also reduces system I/O load by scanning the database only once.


Thursday, 28 February 2013

Which method predicts recidivism best?: a comparison of statistical, machine learning and data mining predictive models

an article by N. Tollenaar (Ministry of Security and Justice, The Hague, The Netherlands) and P. G. M. van der Heijden (University of Utrecht, The Netherlands) published in Journal of the Royal Statistical Society: Series A (Statistics in Society) Volume 176 Issue 2 (February 2013)

Summary

Using criminal population conviction histories of recent offenders, prediction models are developed that predict three types of criminal recidivism: general recidivism, violent recidivism and sexual recidivism.

The research question is whether prediction techniques from modern statistics, data mining and machine learning provide an improvement in predictive performance over classical statistical methods, namely logistic regression and linear discriminant analysis. These models are compared on a large selection of performance measures.

Results indicate that classical methods do equally well as or better than their modern counterparts.

The predictive performance of the various techniques differs only slightly for general and violent recidivism, whereas differences are larger for sexual recidivism.

For the general and violent recidivism data we present the results of logistic regression and for sexual recidivism of linear discriminant analysis.


Tuesday, 8 November 2011

Reviewing person’s value of privacy of online social networking

and article by Ulrike Hugl (University of Innsbruck, Austria) published in Internet Research Volume 21 Issue 4 (2011)

Abstract

Purpose
The paper aims at a multi-faceted review of scholarly work, analysing the current state of empirical studies dealing with privacy and online social networking (OSN) as well as the theoretical “puzzle” of privacy approaches related to OSN usage from the background of diverse disciplines. Drawing on a more pragmatic and practical level, aspects of privacy management are presented as well.
Design/methodology/approach
Based on individual privacy concerns and also publicly communicated threats, information privacy has become an important topic of public and scholarly discussion. Beside diverse positive aspects of OSN sites for users, their information is for example also being used for data mining and profiling, pre-recruiting information as well as economic espionage. This review highlights information privacy mainly from an individual point-of-view, focusing on the usage of OSN sites (OSNs).
Findings
This analysis of scholarly work shows the following findings:
  • first, adults seem to be more concerned about potential privacy threats than younger users; 
  • second, policy makers should be alarmed by a large part of users who underestimate risks of their information privacy on OSNs; 
  • third, in the case of using OSNs and its services, traditional one-dimensional privacy approaches fall short.
Hence, findings of this paper further highlight the necessity to focus on multidimensional and multidisciplinary frameworks of privacy, for example considering a so-called “privacy calculus paradigm” and rethinking “fair information practices” from a more and more ubiquitous environment of OSNs.
Originality/value
The results of the work presented in this paper give new opportunities for research as well as suggestions for privacy management issues for OSN providers and users.