SlideShare a Scribd company logo
The web of data:
how are we
doing so far?
E L E N A S I M P E R L
K I N G ’ S C O L L E G E LO N D O N
@ E S I M P E R L
THE WEB CONFERENCE, APRIL 2021
The web of data: how are we doing so far?
The web has shaped our understanding and
interactions with data
Answering
factual
questions
Sharing
data
online
Publishing
data for
others to
use
Creating
datasets in
collaboration
Creating
digital
traces
Labelling
data for
algorithms
to use
(Source: Fensel, 2013)
The theory and practice of the web of data
are different
We are living through a crucial moment in
how data is published and used on the web
(Source: Hitzler, 2021)
European Data Portal
Technology, resources and support to increase the value of European open government data
Highlights of our work
Supporting the entire data value chain from publishing to reuse
Low uptake of linked data, limited vocabulary
reuse, proprietary, non-dereferenceable
vocabularies, reasonable metadata quality
Content metadata published as linked data,
joint data model, data sharing framework,
Europeana identifiers
25 million datasets (DCAT, schema.org) in
summer 2020
There is a lot of annotated data online,
especially about products, people and
businesses
Making portals more user-centric
Walker & Simperl, 2017
The ten guidelines
Organise for use of the datasets - rather than simply for publication
Promote use through data storytelling and community building, borrowing from open-source communities and other
peer-production systems
Invest in discoverability best practices, borrowing from e-commerce and web search
Publish good quality metadata - to enhance reuse
Adopt standards to ensure interoperability
Co-locate tools so that a wider range of users can be engaged with
Link datasets to enhance value
Be accessible by offering options for from APIs to CSV downloads
Co-locate documentation - users should not need to be domain experts to understand the data;
Be measurable - as a way to assess how well they are meeting users’ needs.
Operationalising the guidance
Literature review to develop 5* schemes to operationalise indicators.
Application of the schemes on 10 open data portals at different maturity level.
(Walker & Simperl, 2017)
Example: Organise for use
Each dataset is accompanied by a comprehensive descriptive
record (going beyond a collection of structured metadata)
An extract of the data can be previewed (for sense making)
The portal provides recommendations for related datasets
The portal enables users to review/rate the datasets
Keywords from datasets are linked to other published datasets
Example: Promote for use
The portal is connected with social media to create a social distribution channel
for open data.
The portal provides users with online support for feedback, to request/suggest
the publication of new datasets, and when problems arise during use (e.g.
contact form, discussion forum, FAQs, helpdesk, search tips, tutorials, demos).
The portal provides a way for users to keep informed of updates to the data (e.g.
news feed).
Datasets are accompanied by links or resources that provide user guidance and
support.
Examples of reuse (fictitious or real) are provided (e.g. information contributed
by other users, last reuse, best reuse, data stories).
Example: Co-locate documentation
Supporting documentation does not exist.
Supporting documentation exists, but as a document found separately from the data.
Supporting documentation is found at the same time as the data (e.g. the link to the document is
next to the link to the data in the search).
Supporting documentation can be immediately accessed from within the dataset but it is not context
sensitive (e.g. a link to the documentation or text contained within the dataset).
Supporting documentation can be immediately accessed from within the dataset and it is context
sensitive so that users can immediately access information about a specific item of concern (e.g. a
link to a specific point in the documentation or the text contained within the dataset).
Varying open data maturity levels
Be discoverable , co-locate documentation, be
measurable are universally challenging
A lot of guidance available already
Is there any evidence that it works?
Be
measurable:
GitHub as a
data platform
~1.4 million datasets (e.g. CSV,
excel) from ~65K repos
Map literature features to both
dataset and repository features
Use engagement metrics as
proxies for data reuse
Train a predictive model to see
what publishing guidance leads to
higher engagement values
Size Attributes
Age
Quality
Documentation
Reviews
Recommendations for publishers
Co-locate documentation:
◦ Informative, short text about the dataset
◦ Comprehensive README file in a structured form,
links to further information
Co-locate tools:
◦ Standard processable file sizes for dataset
distributions
◦ Openable with a standard configuration of a
common library (such as Pandas)
Can people find the data they need?
Analysis of logs and data requests
(2018)
• Four national open government data portals, 2.2 million queries (2013 – 2016), 1500 data
requests.
• Shorter queries, include temporal and location information.
• Explorative search.
• Native and external queries topically different.
• Data requests offer more context to user intent.
Analysis of logs (2020 - 21)
844k sessions from 04/2018 to 06/2020, web search as well as
native search sessions from the European Data Portal
Location, provenance, format, licence, time frame and date, publishing
date, location of publication and data schema
Mostly web search, web search and native search users have different
information needs and different success rates
Dataset preview page is important in web search
Linking to stories and other content helps with traffic
Recommendations for publishers
Two types of
users
Spatial and
temporal
queries
Result
presentation
Quality
reviews
Data stories
More logs
needed!
Data documentation and sensemaking
practices
(Source: Gregory et al., 2020)
Data work is teamwork
Open approaches and standards work best when solving
actual problems. These problems are rarely about a set of
technologies.
Conclusions
We are at a crucial moment in data availability and use, online and elsewhere
There is an increasing body of evidence about what people’s data needs and about how data is
published on the web
We don’t have links and we don’t always have great business cases for creating and maintaining
them on the open, decentralised web. In fact, we need better models to resource data publishing
all together
There are other data modalities e.g. charts which web technologies can help share responsibly
Metadata vocabularies used where there is a clear
business case
More documentation needed to make data useful
for others
Some data is missing, with serious consequences
Charts as alternative to ‘raw’ data. Where are the
links to the data?
Thank you
Talking Datasets: understanding data sensemaking behaviours. L Koesten, K Gregory, P Groth, E Simperl.
International Journal of Human-Computer Studies. 146:102562. 2021
Everything You Always Wanted to Know about a Dataset: Studies in Data Summarisation. L Koesten, E Simperl,
E Kacprzak, T Blount, J Tennison. International Journal of Human-Computer Studies. 2019
Collaborative Practices with Structured Data: Do Tools Support what Users Need? L Koesten, E Kacprzak, E
Simperl, J Tennison; ACM CHI Conference on Human Factors in Computing Systems, CHI 2019.
Dataset search: a survey. A Chapman, E Simperl, L Koesten, G Konstantinidis, LD Ibáñez, E Kacprzak, P Groth.
The International Journal on Very Large Data Bases, 2019.
Characterising dataset search — An analysis of search logs and data requests. E Kacprzak, L Koesten, LD
Ibáñez, T Blount, J Tennison, E Simperl; Journal of Web Semantics, 2018
Making sense of numerical data-semantic labelling of web tables. Kacprzak, E., Giménez-García, J.M., Piscopo,
A., Koesten, L., Ibáñez, L.D., Tennison, J. and Simperl, E. In European Knowledge Acquisition Workshop (pp. 163-
178). Springer, 2018
The Trials and Tribulations of Working with Structured Data - a Study on Information Seeking Behaviour. L
Koesten, E Kacprzak, J Tennison, E Simperl. Proceedings of ACM CHI Conference on Human Factors in
Computing Systems, CHI 2017
Dataset Reuse: Toward Translating Principles to Practice. L Koesten, P Vougiouklis, E Simperl, P Groth - Patterns,
2020
Characterising Dataset Search on the European Data Portal . L Ibáñez, L Koesten, E Kacprzak, E Simperl.
European Data Portal Analytical Report 18, 2020
Understanding Supply and Demand on the European Data Portal. L Ibáñez, E Simperl. European Data Portal
Analytical Report 19, 2020
The Future of Open Data Portals. J Walker, E Simperl. European Data Portal Analytical Report 8, 2017
Smart Rural: The Open Data Gap. J Walker, G Thuermer, E Simperl, L Carr. Data for Policy, 2020

More Related Content

PDF
Are our knowledge graphs trustworthy?
Elena Simperl
 
PDF
The data we want
Elena Simperl
 
PDF
Building better knowledge graphs through social computing
Elena Simperl
 
PDF
Data stories
Elena Simperl
 
PDF
Loops of humans and bots in Wikidata
Elena Simperl
 
PDF
Pie chart or pizza: identifying chart types and their virality on Twitter
Elena Simperl
 
PDF
One does not simply crowdsource the Semantic Web: 10 years with people, URIs,...
Elena Simperl
 
PDF
The human face of AI: how collective and augmented intelligence can help sol...
Elena Simperl
 
Are our knowledge graphs trustworthy?
Elena Simperl
 
The data we want
Elena Simperl
 
Building better knowledge graphs through social computing
Elena Simperl
 
Data stories
Elena Simperl
 
Loops of humans and bots in Wikidata
Elena Simperl
 
Pie chart or pizza: identifying chart types and their virality on Twitter
Elena Simperl
 
One does not simply crowdsource the Semantic Web: 10 years with people, URIs,...
Elena Simperl
 
The human face of AI: how collective and augmented intelligence can help sol...
Elena Simperl
 

What's hot (20)

PDF
Introduction to data science
Tharushi Ruwandika
 
PPTX
Data science
SouravSadhukhan6
 
PPTX
Data Discovery and Visualization
Dr. Neil Brittliff
 
PDF
Isolating values from big data with the help of four v’s
eSAT Journals
 
PPTX
Noshir Contractor's view on the future of Linked Data
Carlos Pedrinaci
 
PDF
Data Science and its impact on society
Vienna Data Science Group
 
PDF
Crowdsourcing and citizen engagement for people-centric smart cities
Elena Simperl
 
PDF
Data science and visualization lab presentation
iHub Research
 
PDF
Graphs in Government
Neo4j
 
PPSX
USING BIGDATA WITH ACADEMIC LIBRARY SERVICES: A VIEW
Nellore Harilakshmi
 
PPTX
The profile of the management (data) scientist: Potential scenarios and skill...
Juan Mateos-Garcia
 
PDF
EDF2013: Invited Talk Julie Marguerite: Big data: a new world of opportunitie...
European Data Forum
 
PPTX
EDF2013: Invited talk Florian Bauer: Unleashing climate and energy knowledge ...
European Data Forum
 
PPTX
SMART Seminar Series: "From Big Data to Smart data"
SMART Infrastructure Facility
 
PPTX
State of Florida Neo4J Graph Briefing - Keynote
Neo4j
 
PPTX
Big data divided (24 march2014)
Han Woo PARK
 
PPT
Quality, Relevance and Importance in Information Retrieval with Fuzzy Semanti...
tmra
 
PPTX
Bigdatacooltools
suresh sood
 
PPTX
More ways of symbol grounding for knowledge graphs?
Paul Groth
 
PPT
Data Processing and Semantics for Advanced Internet of Things (IoT) Applicati...
Artificial Intelligence Institute at UofSC
 
Introduction to data science
Tharushi Ruwandika
 
Data science
SouravSadhukhan6
 
Data Discovery and Visualization
Dr. Neil Brittliff
 
Isolating values from big data with the help of four v’s
eSAT Journals
 
Noshir Contractor's view on the future of Linked Data
Carlos Pedrinaci
 
Data Science and its impact on society
Vienna Data Science Group
 
Crowdsourcing and citizen engagement for people-centric smart cities
Elena Simperl
 
Data science and visualization lab presentation
iHub Research
 
Graphs in Government
Neo4j
 
USING BIGDATA WITH ACADEMIC LIBRARY SERVICES: A VIEW
Nellore Harilakshmi
 
The profile of the management (data) scientist: Potential scenarios and skill...
Juan Mateos-Garcia
 
EDF2013: Invited Talk Julie Marguerite: Big data: a new world of opportunitie...
European Data Forum
 
EDF2013: Invited talk Florian Bauer: Unleashing climate and energy knowledge ...
European Data Forum
 
SMART Seminar Series: "From Big Data to Smart data"
SMART Infrastructure Facility
 
State of Florida Neo4J Graph Briefing - Keynote
Neo4j
 
Big data divided (24 march2014)
Han Woo PARK
 
Quality, Relevance and Importance in Information Retrieval with Fuzzy Semanti...
tmra
 
Bigdatacooltools
suresh sood
 
More ways of symbol grounding for knowledge graphs?
Paul Groth
 
Data Processing and Semantics for Advanced Internet of Things (IoT) Applicati...
Artificial Intelligence Institute at UofSC
 
Ad

Similar to The web of data: how are we doing so far? (20)

PDF
The web of data: how are we doing so far
Elena Simperl
 
PDF
Open government data portals: from publishing to use and impact
Elena Simperl
 
PDF
High-value datasets: from publication to impact
Elena Simperl
 
PDF
General Presentation European Data Portal
EuropeanDataPortal
 
PDF
Data management plans – EUDAT Best practices and case study | www.eudat.eu
EUDAT
 
PPTX
fosscomm2013_ENGAGE_workshop_on_open_public_data
Charalampos Alexopoulos
 
PPT
H2020 data pilot openaire
Sarah Jones
 
PPT
The Horizon 2020 Open Data Pilot - OpenAIRE webinar (Oct. 21 2014) by Sarah J...
OpenAIRE
 
PDF
The linked open government data and metadata lifecycle
Open Data Support
 
PPTX
Open Access Week 2017: Introduction to Open Data Policies in H2020
OpenAIRE
 
PPTX
General introduction to Open Data Policies H2020, influence of OD policies on...
Nancy Pontika
 
PPTX
Bosman and Kramer Open Research: A 2024 NISO Training Series, Session Four: O...
National Information Standards Organization (NISO)
 
PPTX
Open data pilot
Sarah Jones
 
PDF
US EPA OSWER Linked Data Workshop 1-Feb-2013
3 Round Stones
 
PDF
Online promises beyond the policies: what's under the skin
Nicolaie Constantinescu
 
PPTX
A coordinated framework for open data open science in Botswana/Simon Hodson
African Open Science Platform
 
PPT
Open Data Publication - Requirements, Good practices, and Benefits
ariadnenetwork
 
PDF
Open Science - Global Perspectives/Simon Hodson
Academy of Science of South Africa (ASSAf)
 
PDF
How to overcome obstacles to data publication: Issues, requirements, and good...
ariadnenetwork
 
PDF
GBIF and reuse of research data, Bergen (2016-12-14)
Dag Endresen
 
The web of data: how are we doing so far
Elena Simperl
 
Open government data portals: from publishing to use and impact
Elena Simperl
 
High-value datasets: from publication to impact
Elena Simperl
 
General Presentation European Data Portal
EuropeanDataPortal
 
Data management plans – EUDAT Best practices and case study | www.eudat.eu
EUDAT
 
fosscomm2013_ENGAGE_workshop_on_open_public_data
Charalampos Alexopoulos
 
H2020 data pilot openaire
Sarah Jones
 
The Horizon 2020 Open Data Pilot - OpenAIRE webinar (Oct. 21 2014) by Sarah J...
OpenAIRE
 
The linked open government data and metadata lifecycle
Open Data Support
 
Open Access Week 2017: Introduction to Open Data Policies in H2020
OpenAIRE
 
General introduction to Open Data Policies H2020, influence of OD policies on...
Nancy Pontika
 
Bosman and Kramer Open Research: A 2024 NISO Training Series, Session Four: O...
National Information Standards Organization (NISO)
 
Open data pilot
Sarah Jones
 
US EPA OSWER Linked Data Workshop 1-Feb-2013
3 Round Stones
 
Online promises beyond the policies: what's under the skin
Nicolaie Constantinescu
 
A coordinated framework for open data open science in Botswana/Simon Hodson
African Open Science Platform
 
Open Data Publication - Requirements, Good practices, and Benefits
ariadnenetwork
 
Open Science - Global Perspectives/Simon Hodson
Academy of Science of South Africa (ASSAf)
 
How to overcome obstacles to data publication: Issues, requirements, and good...
ariadnenetwork
 
GBIF and reuse of research data, Bergen (2016-12-14)
Dag Endresen
 
Ad

More from Elena Simperl (18)

PDF
When stars align: studies in data quality, knowledge graphs, and machine lear...
Elena Simperl
 
PDF
Knowledge engineering: from people to machines and back
Elena Simperl
 
PDF
This talk was not generated with ChatGPT: how AI is changing science
Elena Simperl
 
PDF
Knowledge graph use cases in natural language generation
Elena Simperl
 
PDF
Knowledge engineering: from people to machines and back
Elena Simperl
 
PDF
What Wikidata teaches us about knowledge engineering
Elena Simperl
 
PDF
Ten myths about knowledge graphs.pdf
Elena Simperl
 
PDF
What Wikidata teaches us about knowledge engineering
Elena Simperl
 
PDF
Data commons and their role in fighting misinformation.pdf
Elena Simperl
 
PDF
The story of Data Stories
Elena Simperl
 
PDF
Qrowd and the city: designing people-centric smart cities
Elena Simperl
 
PDF
Qrowd and the city
Elena Simperl
 
PDF
Inclusive cities: a crowdsourcing approach
Elena Simperl
 
PDF
Making transport smarter, leveraging the human factor
Elena Simperl
 
PDF
Data storytelling
Elena Simperl
 
PDF
Quality and collaboration in Wikidata
Elena Simperl
 
PDF
Beyond monetary incentives: experiments with paid microtasks
Elena Simperl
 
PDF
The Data Pitch call
Elena Simperl
 
When stars align: studies in data quality, knowledge graphs, and machine lear...
Elena Simperl
 
Knowledge engineering: from people to machines and back
Elena Simperl
 
This talk was not generated with ChatGPT: how AI is changing science
Elena Simperl
 
Knowledge graph use cases in natural language generation
Elena Simperl
 
Knowledge engineering: from people to machines and back
Elena Simperl
 
What Wikidata teaches us about knowledge engineering
Elena Simperl
 
Ten myths about knowledge graphs.pdf
Elena Simperl
 
What Wikidata teaches us about knowledge engineering
Elena Simperl
 
Data commons and their role in fighting misinformation.pdf
Elena Simperl
 
The story of Data Stories
Elena Simperl
 
Qrowd and the city: designing people-centric smart cities
Elena Simperl
 
Qrowd and the city
Elena Simperl
 
Inclusive cities: a crowdsourcing approach
Elena Simperl
 
Making transport smarter, leveraging the human factor
Elena Simperl
 
Data storytelling
Elena Simperl
 
Quality and collaboration in Wikidata
Elena Simperl
 
Beyond monetary incentives: experiments with paid microtasks
Elena Simperl
 
The Data Pitch call
Elena Simperl
 

Recently uploaded (20)

PDF
The_Future_of_Data_Analytics_by_CA_Suvidha_Chaplot_UPDATED.pdf
CA Suvidha Chaplot
 
PDF
202501214233242351219 QASS Session 2.pdf
lauramejiamillan
 
PPTX
short term project on AI Driven Data Analytics
JMJCollegeComputerde
 
PPTX
Fuzzy_Membership_Functions_Presentation.pptx
pythoncrazy2024
 
PPTX
Multiscale Segmentation of Survey Respondents: Seeing the Trees and the Fores...
Sione Palu
 
PPTX
Presentation on animal welfare a good topic
kidscream385
 
PDF
Technical Writing Module-I Complete Notes.pdf
VedprakashArya13
 
PDF
717629748-Databricks-Certified-Data-Engineer-Professional-Dumps-by-Ball-21-03...
pedelli41
 
PDF
An Uncut Conversation With Grok | PDF Document
Mike Hydes
 
PDF
Practical Measurement Systems Analysis (Gage R&R) for design
Rob Schubert
 
PPTX
White Blue Simple Modern Enhancing Sales Strategy Presentation_20250724_21093...
RamNeymarjr
 
PPTX
Presentation (1) (1).pptx k8hhfftuiiigff
karthikjagath2005
 
PPTX
Pipeline Automatic Leak Detection for Water Distribution Systems
Sione Palu
 
PPTX
short term internship project on Data visualization
JMJCollegeComputerde
 
PPTX
Introduction to Data Analytics and Data Science
KavithaCIT
 
PDF
202501214233242351219 QASS Session 2.pdf
lauramejiamillan
 
PPTX
Introduction to Biostatistics Presentation.pptx
AtemJoshua
 
PPTX
IP_Journal_Articles_2025IP_Journal_Articles_2025
mishell212144
 
PDF
D9110.pdfdsfvsdfvsdfvsdfvfvfsvfsvffsdfvsdfvsd
minhn6673
 
PDF
Fundamentals and Techniques of Biophysics and Molecular Biology (Pranav Kumar...
RohitKumar868624
 
The_Future_of_Data_Analytics_by_CA_Suvidha_Chaplot_UPDATED.pdf
CA Suvidha Chaplot
 
202501214233242351219 QASS Session 2.pdf
lauramejiamillan
 
short term project on AI Driven Data Analytics
JMJCollegeComputerde
 
Fuzzy_Membership_Functions_Presentation.pptx
pythoncrazy2024
 
Multiscale Segmentation of Survey Respondents: Seeing the Trees and the Fores...
Sione Palu
 
Presentation on animal welfare a good topic
kidscream385
 
Technical Writing Module-I Complete Notes.pdf
VedprakashArya13
 
717629748-Databricks-Certified-Data-Engineer-Professional-Dumps-by-Ball-21-03...
pedelli41
 
An Uncut Conversation With Grok | PDF Document
Mike Hydes
 
Practical Measurement Systems Analysis (Gage R&R) for design
Rob Schubert
 
White Blue Simple Modern Enhancing Sales Strategy Presentation_20250724_21093...
RamNeymarjr
 
Presentation (1) (1).pptx k8hhfftuiiigff
karthikjagath2005
 
Pipeline Automatic Leak Detection for Water Distribution Systems
Sione Palu
 
short term internship project on Data visualization
JMJCollegeComputerde
 
Introduction to Data Analytics and Data Science
KavithaCIT
 
202501214233242351219 QASS Session 2.pdf
lauramejiamillan
 
Introduction to Biostatistics Presentation.pptx
AtemJoshua
 
IP_Journal_Articles_2025IP_Journal_Articles_2025
mishell212144
 
D9110.pdfdsfvsdfvsdfvsdfvfvfsvfsvffsdfvsdfvsd
minhn6673
 
Fundamentals and Techniques of Biophysics and Molecular Biology (Pranav Kumar...
RohitKumar868624
 

The web of data: how are we doing so far?

  • 1. The web of data: how are we doing so far? E L E N A S I M P E R L K I N G ’ S C O L L E G E LO N D O N @ E S I M P E R L THE WEB CONFERENCE, APRIL 2021
  • 3. The web has shaped our understanding and interactions with data
  • 11. The theory and practice of the web of data are different We are living through a crucial moment in how data is published and used on the web
  • 13. European Data Portal Technology, resources and support to increase the value of European open government data
  • 14. Highlights of our work Supporting the entire data value chain from publishing to reuse
  • 15. Low uptake of linked data, limited vocabulary reuse, proprietary, non-dereferenceable vocabularies, reasonable metadata quality
  • 16. Content metadata published as linked data, joint data model, data sharing framework, Europeana identifiers
  • 17. 25 million datasets (DCAT, schema.org) in summer 2020
  • 18. There is a lot of annotated data online, especially about products, people and businesses
  • 19. Making portals more user-centric Walker & Simperl, 2017
  • 20. The ten guidelines Organise for use of the datasets - rather than simply for publication Promote use through data storytelling and community building, borrowing from open-source communities and other peer-production systems Invest in discoverability best practices, borrowing from e-commerce and web search Publish good quality metadata - to enhance reuse Adopt standards to ensure interoperability Co-locate tools so that a wider range of users can be engaged with Link datasets to enhance value Be accessible by offering options for from APIs to CSV downloads Co-locate documentation - users should not need to be domain experts to understand the data; Be measurable - as a way to assess how well they are meeting users’ needs.
  • 21. Operationalising the guidance Literature review to develop 5* schemes to operationalise indicators. Application of the schemes on 10 open data portals at different maturity level. (Walker & Simperl, 2017)
  • 22. Example: Organise for use Each dataset is accompanied by a comprehensive descriptive record (going beyond a collection of structured metadata) An extract of the data can be previewed (for sense making) The portal provides recommendations for related datasets The portal enables users to review/rate the datasets Keywords from datasets are linked to other published datasets
  • 23. Example: Promote for use The portal is connected with social media to create a social distribution channel for open data. The portal provides users with online support for feedback, to request/suggest the publication of new datasets, and when problems arise during use (e.g. contact form, discussion forum, FAQs, helpdesk, search tips, tutorials, demos). The portal provides a way for users to keep informed of updates to the data (e.g. news feed). Datasets are accompanied by links or resources that provide user guidance and support. Examples of reuse (fictitious or real) are provided (e.g. information contributed by other users, last reuse, best reuse, data stories).
  • 24. Example: Co-locate documentation Supporting documentation does not exist. Supporting documentation exists, but as a document found separately from the data. Supporting documentation is found at the same time as the data (e.g. the link to the document is next to the link to the data in the search). Supporting documentation can be immediately accessed from within the dataset but it is not context sensitive (e.g. a link to the documentation or text contained within the dataset). Supporting documentation can be immediately accessed from within the dataset and it is context sensitive so that users can immediately access information about a specific item of concern (e.g. a link to a specific point in the documentation or the text contained within the dataset).
  • 25. Varying open data maturity levels Be discoverable , co-locate documentation, be measurable are universally challenging
  • 26. A lot of guidance available already
  • 27. Is there any evidence that it works?
  • 28. Be measurable: GitHub as a data platform ~1.4 million datasets (e.g. CSV, excel) from ~65K repos Map literature features to both dataset and repository features Use engagement metrics as proxies for data reuse Train a predictive model to see what publishing guidance leads to higher engagement values Size Attributes Age Quality Documentation Reviews
  • 29. Recommendations for publishers Co-locate documentation: ◦ Informative, short text about the dataset ◦ Comprehensive README file in a structured form, links to further information Co-locate tools: ◦ Standard processable file sizes for dataset distributions ◦ Openable with a standard configuration of a common library (such as Pandas)
  • 30. Can people find the data they need?
  • 31. Analysis of logs and data requests (2018) • Four national open government data portals, 2.2 million queries (2013 – 2016), 1500 data requests. • Shorter queries, include temporal and location information. • Explorative search. • Native and external queries topically different. • Data requests offer more context to user intent.
  • 32. Analysis of logs (2020 - 21) 844k sessions from 04/2018 to 06/2020, web search as well as native search sessions from the European Data Portal Location, provenance, format, licence, time frame and date, publishing date, location of publication and data schema Mostly web search, web search and native search users have different information needs and different success rates Dataset preview page is important in web search Linking to stories and other content helps with traffic
  • 33. Recommendations for publishers Two types of users Spatial and temporal queries Result presentation Quality reviews Data stories More logs needed!
  • 34. Data documentation and sensemaking practices
  • 35. (Source: Gregory et al., 2020) Data work is teamwork
  • 36. Open approaches and standards work best when solving actual problems. These problems are rarely about a set of technologies.
  • 37. Conclusions We are at a crucial moment in data availability and use, online and elsewhere There is an increasing body of evidence about what people’s data needs and about how data is published on the web We don’t have links and we don’t always have great business cases for creating and maintaining them on the open, decentralised web. In fact, we need better models to resource data publishing all together There are other data modalities e.g. charts which web technologies can help share responsibly
  • 38. Metadata vocabularies used where there is a clear business case More documentation needed to make data useful for others
  • 39. Some data is missing, with serious consequences
  • 40. Charts as alternative to ‘raw’ data. Where are the links to the data?
  • 41. Thank you Talking Datasets: understanding data sensemaking behaviours. L Koesten, K Gregory, P Groth, E Simperl. International Journal of Human-Computer Studies. 146:102562. 2021 Everything You Always Wanted to Know about a Dataset: Studies in Data Summarisation. L Koesten, E Simperl, E Kacprzak, T Blount, J Tennison. International Journal of Human-Computer Studies. 2019 Collaborative Practices with Structured Data: Do Tools Support what Users Need? L Koesten, E Kacprzak, E Simperl, J Tennison; ACM CHI Conference on Human Factors in Computing Systems, CHI 2019. Dataset search: a survey. A Chapman, E Simperl, L Koesten, G Konstantinidis, LD Ibáñez, E Kacprzak, P Groth. The International Journal on Very Large Data Bases, 2019. Characterising dataset search — An analysis of search logs and data requests. E Kacprzak, L Koesten, LD Ibáñez, T Blount, J Tennison, E Simperl; Journal of Web Semantics, 2018 Making sense of numerical data-semantic labelling of web tables. Kacprzak, E., Giménez-García, J.M., Piscopo, A., Koesten, L., Ibáñez, L.D., Tennison, J. and Simperl, E. In European Knowledge Acquisition Workshop (pp. 163- 178). Springer, 2018 The Trials and Tribulations of Working with Structured Data - a Study on Information Seeking Behaviour. L Koesten, E Kacprzak, J Tennison, E Simperl. Proceedings of ACM CHI Conference on Human Factors in Computing Systems, CHI 2017 Dataset Reuse: Toward Translating Principles to Practice. L Koesten, P Vougiouklis, E Simperl, P Groth - Patterns, 2020 Characterising Dataset Search on the European Data Portal . L Ibáñez, L Koesten, E Kacprzak, E Simperl. European Data Portal Analytical Report 18, 2020 Understanding Supply and Demand on the European Data Portal. L Ibáñez, E Simperl. European Data Portal Analytical Report 19, 2020 The Future of Open Data Portals. J Walker, E Simperl. European Data Portal Analytical Report 8, 2017 Smart Rural: The Open Data Gap. J Walker, G Thuermer, E Simperl, L Carr. Data for Policy, 2020