Rogue Scholar

WikidataWikipediaComputer and Information Sciences

Wikidata, Wikipedia, and #wikisci

Published September 7, 2015

Last week I attended the Wikipedia Science Conference (hashtag: #wikisci) at the Wellcome Trust in London. it was an interesting two days of talks and discussion. Below are a few random notes on topics that caught my eye. What is Wikidata? A recurring theme was the emergence of Wikidata, although it never really seemed clear what role Wikidata saw for itself.

AnnotationBHLBioStorHypothes.isComputer and Information Sciences

Hypothes.is revisited: annotating articles in BioStor

https://doi.org/10.59350/5d5x5-qzk46

Published September 2, 2015

Author Roderic Page

Over the weekend, out of the blue, Dan Whaley commented on an earlier blog post of mine (Altmetrics, Disqus, GBIF, JSTOR, and annotating biodiversity data. Dan is the project lead for hypothes.is, a tool to annotate web pages.

DNA BarcodingEnvironmental DNAComputer and Information Sciences

Dark taxa, drones, and Dan Janzen: 6th International Barcode of Life Conference

https://doi.org/10.59350/5xxw2-yk380

Published September 1, 2015

Author Roderic Page

A little over a week ago I was at the 6th International Barcode of Life Conference, held at Guelph, Canada. It was my first barcoding conference, and was quite an experience. Here are a few random thoughts. Attendees It was striking how diverse the conference crowd was. Apart from a few ageing systematists (including veterans of the cladistics wars), most people were young(ish), and from all over the world.

NamestreamPossible ProjectTaxonomic NamesComputer and Information Sciences

Possible project: NameStream - a stream of new taxonomic names

https://doi.org/10.59350/68pt9-7zv22

Published August 14, 2015

Author Roderic Page

Yet another barely thought out project, although this one has some crude code. If some 16,000 new taxonomic names are published each year, then that is roughly 40 per day. We don't have a single place that aggregates these, so any major biodiversity projects is by definition out of date. GBIF itself hasn't had an update list of fungi or plant names for several years, and at present doesn't have an up to date list of animal names.

Possible ProjectPubMed CentralComputer and Information Sciences

Possible project: A PubMed Central for taxonomy

https://doi.org/10.59350/837e3-9k809

Published August 14, 2015

Author Roderic Page

I need more time to sketch this out fully, but I think a case can be made for a taxonomy-centric (or, perhaps more usefully, a biodiversity-centric) clone of PubMed Central. Why? We already have PubMed Central, and a European version Europe PubMed Central, and the content of Open Access journals such as ZooKeys appears in both, so, again, why?

BHLCloudantCouchDBDjVuSearchComputer and Information Sciences

Demo of full-text indexing of BHL using CouchDB hosted by Cloudant

https://doi.org/10.59350/4crdc-fm682

Published August 10, 2015

Author Roderic Page

One of the limitations of the Biodiversity Heritage Library (BHL) is that, unlike say Google Books, its search functions are limited to searching metadata (e.g., book and article titles) and taxonomic names. It doesn't support full-text search, by which I mean you can't just type in the name of a locality, specimen code, or a phrase and expect to get back much in the way of results.

ISNIORCIDPossible ProjectWikipediaComputer and Information Sciences

Possible project: mapping authors to Wikipedia entries using lists of published works

https://doi.org/10.59350/25rrm-gaj56

Published August 10, 2015

Author Roderic Page

One of the less glamorous but necessary tasks of data cleaning is mapping "strings to things", that is, taking strings such as "George A. Boulenger" and mapping them to identifiers, such as ISNI: 0000 0001 0888 841X. In case of authors such as George Boulenger, one way to do this would be through Wikipedia, which has entries for many scientists, often linked to identifiers for those people (see the bottom of the Wikipedia page for George A.

GBIFIPNINeo4JComputer and Information Sciences

More Neo4J tests of GBIF taxonomy: Using IPNI to find objective synonyms

https://doi.org/10.59350/3r86d-ctj80

Published August 9, 2015

Author Roderic Page

Following on from Testing the GBIF taxonomy with the graph database Neo4J I've added a more complex test that relies on linking taxa to names. In this case I've picked some legume genera ( Coursetia and Poissonia ) where there have been frequent changes of name.

Note To SelfPossible ProjectComputer and Information Sciences

Possible project: #itaxonomist, combining taxonomic names, DOIs, and ORCID to measure taxonomic impact

https://doi.org/10.59350/xe6ry-h7t08

Published August 9, 2015

Author Roderic Page

Imagine a web site where researchers can go, log in (easily) and get a list of all the species they have described (with pretty pictures and, say, GBIF map), and a list of all DNA sequences/barcodes (if any) that they've published. Imagine that this is displayed in a colourful way (e.g., badges), and the results tweeted with the hastag #itaxonomist.

GBIFGraph DatabaseNeo4JRDFTaxonomyComputer and Information Sciences

Testing the GBIF taxonomy with the graph database Neo4J

https://doi.org/10.59350/tedm1-m9476

Published August 7, 2015

Author Roderic Page

I've been playing with the graph database Neo4J to investigate aspects of the classification of taxa in GBIF's backbone classification. Neo4J is a graph database, and a number of people in biodiversity informatics have been playing with it. Nicky Nicolson at Kew has a nice presentation using graph databases to handle names Building a names backbone, and the Open Tree of Life project use it in their tree machine.

FolksonomyMachine LearningNote To SelfPossible ProjectTagsComputer and Information Sciences

Possible project: extract taxonomic classification from tags (folksonomy)

https://doi.org/10.59350/n20vx-a4930

Published August 4, 2015

Author Roderic Page

Note to self about a possible project. This PLoS ONE paper: describes a method for inferring a hierarchy from a set of tags (and cites related work that is of interest). I've grabbed the code and data from http://hiertags-beta.elte.hu/home/ and put it on GitHub. Possible project Use Tibély et al. method (or others) on taxonomic names extracted from BHL text (or other) and see if we can reconstruct taxonomic classifications.

iPhylo

Wikidata, Wikipedia, and #wikisci

Hypothes.is revisited: annotating articles in BioStor

Dark taxa, drones, and Dan Janzen: 6th International Barcode of Life Conference

Possible project: NameStream - a stream of new taxonomic names

Possible project: A PubMed Central for taxonomy

Demo of full-text indexing of BHL using CouchDB hosted by Cloudant

Possible project: mapping authors to Wikipedia entries using lists of published works

More Neo4J tests of GBIF taxonomy: Using IPNI to find objective synonyms

Possible project: #itaxonomist, combining taxonomic names, DOIs, and ORCID to measure taxonomic impact

Testing the GBIF taxonomy with the graph database Neo4J

Possible project: extract taxonomic classification from tags (folksonomy)