Showing posts with label big_data. Show all posts
Showing posts with label big_data. Show all posts

Friday, 13 December 2019

Using history to understand hidden wealth in the UK

a column by Neil Cummins for VOX: CEPR’s Policy Portal

Sharp declines in the concentration of declared wealth occurred across Europe and the US during the 20th century. But the rich may have been hiding much of their wealth.

This column introduces a new method to measure this hidden wealth, in any form. It finds that between 1920 and 1992, English elites concealed 20-32% of their wealth. Accounting for hidden wealth eliminates one-third of the observed decline of top 10% wealth share over the past century.

Continue reading

The column contains some really useful charts/graphs but I have not reproduced any of them in this very short introduction as I think you will need to read the surrounding text for them to make a lot of sense!
Hazel
 


Friday, 27 September 2019

Big data and collective intelligence

an article by Mirjana Ivanović and Aleksandra Klašnja-Milićević (University of Novi Sad, Serbia) published in International Journal of Embedded Systems Volume 11 Number 5 (2019)

Abstract

Nowadays the creation and accumulation of big data is an unavoidable process in a wide range of situations and scenarios. Smart environments and diverse sources of sensors, as well as the content created by humans, contribute to the big data's enormous size and characteristics.

To make sense of the data, analyse and use these data, more and more efficient algorithms are being developed constantly. Still, the effectiveness of these algorithms depends on the specific nature of big data: analogue, noisy, implicit, and ambiguous.

At the same time, there is the unavoidable scientific area of collective intelligence. It represents the capability of interconnected intelligences to collectively and more efficiently solve concrete problems than each individual intelligence would be able to do on its own.

The paper presents an overview of recent achievements in big data and collective intelligence research areas. At the end, the perspectives and challenges of the common directions of these two areas will be discussed.


Monday, 10 June 2019

Where fairness fails: data, algorithms, and the limits of antidiscrimination discourse

an article by Anna Lauren Hoffmann (The Information School, University of Washington, Seattle, WA, USA) published in Information, Communication & Society Volume 22 Issue 7 (2019)

Abstract

Problems of bias and fairness are central to data justice, as they speak directly to the threat that ‘big data’ and algorithmic decision-making may worsen already existing injustices. In the United States, grappling with these problems has found clearest expression through liberal discourses of rights, due process, and antidiscrimination.

Work in this area, however, has tended to overlook certain established limits of antidiscrimination discourses for bringing about the change demanded by social justice.


  1. In this paper, I engage three of these limits:an overemphasis on discrete ‘bad actors’,
  2. single-axis thinking that centers disadvantage, and
  3. an inordinate focus on a limited set of goods.

I show that, in mirroring some of antidiscrimination discourse’s most problematic tendencies, efforts to achieve fairness and combat algorithmic discrimination fail to address the very hierarchical logic that produces advantaged and disadvantaged subjects in the first place.

Finally, I conclude by sketching three paths for future work to better account for the structural conditions against which we come to understand problems of data and unjust discrimination in the first place.

Full text (PDF 17pp)


Friday, 17 May 2019

Collecting user data is a competitive disadvantage

a post by Cary Doctorow for the Boing Boing blog


Warren Buffet is famous for identifying the need for businesses to have "moats" and "walls" around their profit-centers to keep competitors out, and data-centric companies often cite their massive collections of user-data as "moats" that benefit from "network effects" to make their businesses good investments.

In a smart, eye-opening essay, Martin Casado and Peter Lauten from the VC firm Andreesen Horowitz dismantle the idea that data benefits from "network effects" and that it presents any kind of "moat" to protect businesses: instead, the VCs demonstrate how collecting data gets more expensive, and less useful, over time.

To understand why, think of Netflix's data-collection, performed in service to its famous recommendation engine, which suggests programs you might enjoy based on the preferences of people who are similar to you. When Netflix is starting out, it needs to develop a "minimum viable corpus" in order to produce recommendations, but once that data is in place, new data produces diminishing returns in recommendations. Going from 100 to 1,000,000 users allows Netflix to dramatically improve its recommendations, but going from 1,000,000 to 1,000,100 (or even 2,000,000) produces very little new benefit.

Continue reading


Friday, 10 May 2019

Ethics of identity in the time of big data

an article by James Brusseau (Pace University, New York City, USA) published in First Monday Volume 24 Number 5 (May 2019)

Abstract

Compartmentalizing our distinct personal identities is increasingly difficult in big data reality.

Pictures of the person we were on past vacations resurface in employers’ Google searches; LinkedIn which exhibits our income level is increasingly used as a dating web site. Whether on vacation, at work, or seeking romance, our digital selves stream together.

One result is that a perennial ethical question about personal identity has spilled out of philosophy departments and into the real world. Ought we possess one, unified identity that coherently integrates the various aspects of our lives, or, incarnate deeply distinct selves suited to different occasions and contexts?

At bottom, are we one, or many?

The question is not only palpable today, but also urgent because if a decision is not made by us, the forces of big data and surveillance capitalism will make it for us by compelling unity. Speaking in favor of the big data tendency, Facebook’s Mark Zuckerberg promotes the ethics of an integrated identity, a single version of selfhood maintained across diverse contexts and human relationships.

This essay goes in the other direction by sketching two ethical frameworks arranged to defend our compartmentalized identities, which amounts to promoting the dis-integration of our selves. One framework connects with natural law, the other with language, and both aim to create a sense of selfhood that breaks away from its own past, and from the unifying powers of big data technology.

Full text (HTML)


Tuesday, 19 March 2019

Rationality and politics of algorithms. Will the promise of big data survive the dynamics of public decision making?

H.G.van der Voort and A.J.Klievink (TU Delft, the Netherlands), M.Arnaboldi (Politecnico di Milano, Italy) and A.J.Meijer (Utrecht University, the Netherlands) published in Government Information Quarterly Volume 36 Issue 1 (January 2019)

Highlights

  • A framework respecting both a rational and a political view on decision-making and respecting the perspectives of both data analysts and decision makers.
  • We found that data analysts are politically significant in all stages. They have a crucial role in constructing information, facilitating framing and adjusting information to political preferences. However, powers of data analysts are limited. Data analysts do their work in the context of predefined policies. Decision makers sometimes just neglect or bypass big data analysis.
  • Mere functional support for big data to decision making or mere warnings for the politics of big data may miss an important aspect of big data. An important additional question is whether the powers of data analysts and decision makers are well-balanced and how checks and balances could be organized.

Abstract

Big data promises to transform public decision-making for the better by making it more responsive to actual needs and policy effects.

However, much recent work on big data in public decision-making assumes a rational view of decision-making, which has been much criticized in the public administration debate.

In this paper, we apply this view, and a more political one, to the context of big data and offer a qualitative study. We question the impact of big data on decision-making, realizing that big data – including its new methods and functions – must inevitably encounter existing political and managerial institutions.

By studying two illustrative cases of big data use processes, we explore how these two worlds meet. Specifically, we look at the interaction between data analysts and decision makers.

In this we distinguish between a rational view and a political view, and between an information logic and a decision logic. We find that big data provides ample opportunities for both analysts and decision makers to do a better job, but this doesn't necessarily imply better decision-making, because big data also provides opportunities for actors to pursue their own interests.

Big data enables both data analysts and decision makers to act as autonomous agents rather than as links in a functional chain. Therefore, big data's impact cannot be interpreted only in terms of its functional promise; it must also be acknowledged as a phenomenon set to impact our policy-making institutions, including their legitimacy.

Full text (HTML)


Tuesday, 12 March 2019

Equal access to online information? Google’s suicide-prevention disparities may amplify a global digital divide

an article by Sebastian Scherr (University of Leuven, Belgium) and Mario Haim and Florian Arendt (University of Munich (LMU), Germany) published in New Media & Society Volume 21 Issue 3 (March 2019)

Abstract

Worldwide, people profit from equally accessible online health information via search engines. Therefore, equal access to health information is a global imperative.

We studied one specific scenario, in which Google functions as a gatekeeper when people seek suicide-related information using both helpful and harmful suicide-related search terms.

To help prevent suicides, Google implemented a “suicide-prevention result” (SPR) at the very top of such search results. While this effort deserves credit, the present investigation compiled evidence that the SPR is not equally displayed to all users.

Using a virtual agent-based testing methodology, a set of 3 studies in 11 countries found that the presentation of the SPR varies depending on where people search for suicide-related information.

Language is a key factor explaining these differences. Google’s algorithms thereby contribute to a global digital divide in online health-information access with possibly lethal consequences. Higher and globally balanced display frequencies are desirable.


Wednesday, 16 January 2019

Big Data's "theory-free" analysis is a statistical malpractice

a post by Cory Doctorow for the Boing Boing blog



One of the premises of Big Data is that it can be "theory free": rather than starting with a hypothesis ("men at buffets eat more when women are present," "more people will click this button if I move it here," etc) and then gathering data to validate your guess, you just gather a ton of data and look for patterns in it.

The thing is, patterns emerge in every large dataset, without necessarily being representative of a wider statistical truth. Think of the celebrated rise and fall of Google Flu: researchers examined the 45 search terms that were most prevalent where the flu had spread and concluded that these were predictors of flu, but the predictive power turned out to be an illusion. Every place has 45 top search terms, all the time, and some of them will coincide with flu outbreaks, but without a causal theory that you can test, all you know for sure is that you've found an incident of correlation, and no way to know whether the correlation is coincidence or a newly discovered iron law.

Continue reading

I have only thought to add. “Eight out of ten cats …”.


Thursday, 25 October 2018

The Wild West of information markets: What we need to know before law and order can rule

a column by Dirk Bergemann and Alessandro Bonatti for VOX: CEPR’s Policy Portal

The growth of social media over the past decade has brought a parallel explosion in the size and value of information markets.

This column presents the findings from a comprehensive model of data trading and brokerage. The model identifies three aspects of information markets – the value of information, the nature of competition in these markets, and consumers’ incentives – which are in particular need of further research and understanding.

Continue reading


Saturday, 14 April 2018

Prediction, pre-emption and limits to dissent: Social media and big data uses for policing protests in the United Kingdom

an article by Lina Dencik and Arne Hintz (Cardiff University, UK) and Zoe Carey (The New School, USA) published in New Media & Society Volume 20 Issue 4 (April 2018)

Abstract

Social media and big data uses form part of a broader shift from ‘reactive’ to ‘proactive’ forms of governance in which state bodies engage in analysis to predict, pre-empt and respond in real time to a range of social problems.

Drawing on research with British police, we contextualize these algorithmic processes within actual police practices, focusing on protest policing.

Although aspects of algorithmic decision-making have become prominent in police practice, our research shows that they are embedded within a continuous human–computer negotiation that incorporates a rooted claim to ‘professional judgement’, an integrated intelligence context and a significant level of discretion. This context, we argue, transforms conceptions of threats.

We focus particularly on three challenges:

  • the inclusion of pre-existing biases and agendas,
  • the prominence of marketing-driven software, and
  • the interpretation of unpredictability.

Such a contextualized analysis of data uses provides important insights for the shifting terrain of possibilities for dissent.

Full text (PDF 18pp)


Friday, 13 April 2018

What Is Big Data, Why Is It Important, and How Dangerous Is It?

a post by Bertel King, Jr. for the MakeUseOf blog

Data is information, but that’s only part of the story. One detail about an event or a factoid about human health isn’t much data to work with. It’s the collection, organization, and storage of information that we think of when we talk about data.

In the internet age, companies and organizations around the world have collected so much data that we’re now talking about matters on an exponentially larger scale. Now there’s big data, and it’s having a huge impact on all of our lives.

What Is Big Data?

Big data is a data set so large that our traditional means of managing information isn’t up for the job. This collection can take many forms.

Examples of Big Data
  • The tweets stored on Twitter’s servers
  • The information Google gets from tracking car rides
  • A country’s full set of local and national election results, going back as far as records have been kept
  • What health insurance companies know about who gets what treatments done at which hospitals
  • The types of purchases and places that appear on credit cards
  • What people watch on Netflix, when, where, and how long
Continue reading


Thursday, 22 March 2018

Digital revolutions in public finance

a column by Sanjeev Gupta, Michael Keen, Alpa Shah and Geneviève Verdier for VOX: CEPR’s Policy Portal

Digitalisation has vastly increased our ability to collect and exploit the information that governments use to implement macroeconomic policy. The column argues that the ability of governments to use the vast amounts of information held in the private sector on financial transactions are already making fiscal policy more efficient and effective. Problems of access to digital technology, cybersecurity risks, and the difficulty of organisational change in the public sector may slow the pace at which these opportunities are exploited.

Continue reading


Thursday, 8 February 2018

Skillset and match (Cedefop’s magazine promoting learning for work)

The January 2018 issue of Skillset and match, Cedefop’s magazine promoting learning for work, is now available to read and download.

In this issue:

  • A feature on the Second European vocational skills week and the #CedefopPhotoAward 2017;
  • Lifelong guidance in the digital age;
  • Hacking big data for better labour market policies;
    I have recently seen a VOX item on this topic.
    Access it here. https://voxeu.org/article/economic-predictions-big-data-illusion-sparsity
  • An interview with Martina Dlabajová MEP, who hosted a Cedefop working dinner on digitalisation and new forms of work, their opportunities and challenges;
  • Insights on what European citizens think about vocational education and training;
  • and Europass passes the 100 million CV milestone!

The Member State contribution comes from current EU Presidency holder, Bulgaria.

And, as usual, you can browse through the latest Cedefop publications and upcoming events.


Skillset and match – January 2018 issue 12


Wednesday, 15 November 2017

Enterprise data breach: causes, challenges, prevention, and future directions

an article by Long Cheng,Fang Liu and Danfeng (Daphne) Yao (Virginia Tech, USA) published in WIREs Data Mining and Knowledge Discovery Volume 7 Issue 5 (September/October 2017)

Abstract

A data breach is the intentional or inadvertent exposure of confidential information to unauthorized parties.

In the digital era, data has become one of the most critical components of an enterprise. Data leakage poses serious threats to organizations, including significant reputational damage and financial losses. As the volume of data is growing exponentially and data breaches are happening more frequently than ever before, detecting and preventing data loss has become one of the most pressing security concerns for enterprises.

Despite a plethora of research efforts on safeguarding sensitive information from being leaked, it remains an active research problem.

This review helps interested readers to learn about enterprise data leak threats, recent data leak incidents, various state-of-the-art prevention and detection techniques, new challenges, and promising solutions and exciting opportunities.

Full text (PDF 14pp)

I did not understand much of the text but the images provided made this whole area of unseen attacks clearer.


100,000 false positives for every real terrorist: Why anti-terror algorithms don't work

an article by Timme Bisgaard Munk (University of Copenhagen, Denmark) published in First Monday Volume 22 Number 9 (September 2017)

Abstract

Can terrorist attacks be predicted and prevented using classification algorithms? Can predictive analytics see the hidden patterns and data tracks in the planning of terrorist acts?

According to a number of IT firms that now offer programs to predict terrorism using predictive analytics, the answer is yes. According to scientific and application-oriented literature, however, these programs raise a number of practical, statistical and recursive problems. In a literature review and discussion, this paper examines specific problems involved in predicting terrorism.

The problems include the opportunity cost of false positives/false negatives, the statistical quality of the prediction and the self-reinforcing, corrupting recursive effects of predictive analytics, since the method lacks an inner meta-model for its own learning- and pattern-dependent adaptation.

The conclusion is algorithms don’t work for detecting terrorism and is ineffective, risky and inappropriate, with potentially 100,000 false positives for every real terrorist that the algorithm finds.

Full Text (HTML)


Saturday, 5 November 2016

Privacy concerns in smart cities

an article by Liesbet van Zoonen (Erasmus University Rotterdam, Netherlands) published in Government Information Quarterly Volume 33 Issue 3 (July 2016)

Highlights

• Discussion of arguments for including people’s privacy concerns in research, policy and design of smart cities.
• Thorough review of current research about people’s privacy concerns and the paradoxes that typify them.
• People’s concerns are shown as structured by how they perceive city data, and for which purpose they feel this data is used.
• Framework to assess if and how specific technologies and data-usage in smart cities will evoke people’s privacy concerns.
• Clear directions for further academic research about people’s privacy concerns in smart cities.
• Sensitizing instrument for policymakers and operational managers about privacy concerns among their citizens.

Abstract

In this paper a framework is constructed to hypothesize if and how smart city technologies and urban big data produce privacy concerns among the people in these cities (as inhabitants, workers, visitors, and otherwise). The framework is built on the basis of two recurring dimensions in research about people's concerns about privacy: one dimensions represents that people perceive particular data as more personal and sensitive than others, the other dimension represents that people's privacy concerns differ according to the purpose for which data is collected, with the contrast between service and surveillance purposes most paramount.

These two dimensions produce a 2 × 2 framework that hypothesizes which technologies and data-applications in smart cities are likely to raise people's privacy concerns, distinguishing between raising hardly any concern (impersonal data, service purpose), to raising controversy (personal data, surveillance purpose).

Specific examples from the city of Rotterdam are used to further explore and illustrate the academic and practical usefulness of the framework. It is argued that the general hypothesis of the framework offers clear directions for further empirical research and theory building about privacy concerns in smart cities, and that it provides a sensitizing instrument for local governments to identify the absence, presence, or emergence of privacy concerns among their citizens.

Full text


Tuesday, 20 August 2013

Measuring the UK's digital economy with big data

a research report by Max Nathan and Anna Rosso with Tom Gatten, Prash Majmudar and Alex Mitchell for the National Institute of Economic and Social Research

Key Findings
  • The digital economy is poorly served by conventional definitions and datasets. Big data methods can provide richer, more informative and more up to date analysis.
  • Using Growth Intelligence data on a benchmarking sample, we find that the digital economy is substantially larger than conventional estimates suggest. On our preferred measure, it comprises almost 270,000 active companies in the UK (14.4% of all companies as of August 2012). This compares to 167,000 companies (10.0%) when the Government’s conventional SIC-based definitions are used.
  • SIC-based definitions of the digital economy miss out a large number of companies in business and domestic software, architectural activities, engineering, and engineering-related scientific and technical consulting, among other sectors.
  • Companies in the digital economy have a similar average age to those outside it. Shares of start-ups (companies up to three years old) are very similar. Given the popular image of the digital economy as start-up dominated, this may be surprising to some. As digital platforms and tools spread out into the wider economy, and become pervasive in a greater number of sectors, so the set of ‘digital’ companies widens.
  • Inflows of digital companies into the economy have always been relatively small, given its sectoral share. However, using our new definitions of the digital economy, inflow levels are substantially higher.
  • As far as we can tell, digital economy companies have lower average revenues than the rest of the economy, but the median digital company has higher revenues than the median company elsewhere in the economy. Revenue growth rates are also higher for digital companies. However, these results come from a sub-sample of older, likely stronger-performing companies, so there is some positive selection at work.
  • Switching from SIC-based to Growth Intelligence derived measures substantially increases the digital economy’s share of employment, from around 5% to 11% of jobs. Digital economy companies also show higher average employment than companies in the rest of the economy (this reverses when we use conventional SIC-based measures of the digital economy). Looking at median employees per firm, the digital/non-digital differences are always a lot smaller. Our employment results should also be treated with some care, as not all companies report their workforce information.
  • The digital economy is highly concentrated in a few locations around the UK: Growth Intelligence software provides a fresh look at these patterns. In terms of raw firm counts, London dominates the pictures, but Manchester, Birmingham, Brighton and locations in the Greater South East (such as Reading and Crawley) also feature in the top 10. Location quotients show the extent of local clustering, which for the UK’s digital economy is highest for areas in the Western arc around London, such as Basingstoke, Newbury and Milton Keynes. Areas like Aberdeen and Middlesbrough also show high concentrations of digital economy activity.
Full text (PDF 43pp)