Skip to content

Wikidata Tools

Scholarly linked open data management

Wikidata Tools is a collection of Jupyter notebooks for building, enriching, and maintaining scholarly records in Wikidata. Originally developed to improve coverage of anthropology journals, the tools are applicable to any academic discipline. All notebooks run in Google Colab without local setup.

The toolkit covers the full Wikidata scholarly pipeline: querying CrossRef and Wikidata to assess journal coverage, generating QuickStatements for bulk article import, verifying and repairing post-import records, detecting and resolving duplicates, and enriching author records with ORCID identifiers and institutional affiliations.

What’s included

Query & Coverage

Query anthropologists in Wikidata, count journal articles on the scholarly endpoint, track subject tagging, and rank researchers by composite prominence scores.

Import & Cleanup

Transform bibliographic CSV data into QuickStatements for batch import. Verify imports, detect duplicates by DOI and title matching, and generate cleanup commands.

Enrichment & Reconciliation

Add citations, external identifiers, and author data from Semantic Scholar, ORCID, and CrossRef. Reconcile author name strings to person items via ORCID matching.

Scholarly data workflows

Wikidata Anthropologist List

Query all anthropologists in Wikidata with metadata (birth/death, gender, citizenship, ORCID)

Wikidata Journal Article Counter

Count scholarly articles per journal on the Wikidata scholarly endpoint

Wikidata Anthropology Main Subject Counter

Count articles tagged with “anthropology” per journal, identify journals lacking subject tags

Wikidata Prominent Anthropologists

Rank ~13K anthropologists by composite prominence score (sitelinks, doctoral students, awards, authored works)

Wikidata Article Importer

Transform bibliographic CSV data into QuickStatements for batch importing scholarly articles

Post-Import Verification and Repair

Verify QuickStatements imports against source CSV, identify missing articles, generate repair statements

Wikidata Duplicate Detector

Detect duplicate scholarly articles by DOI matching and normalized title comparison

Wikidata Non-Article Cleanup

Identify non-articles in CrossRef CSV exports, look up QIDs, and generate deletion commands

Wikidata Author Name String Enrichment

Add qualifiers (series ordinal, affiliation) and CrossRef references to existing author name string statements

Wikidata ORCID Author Reconciliation

Match author name strings to person items using ORCID, generate QuickStatements for reconciliation

Wikidata Citation Enrichment

Enrich articles with citation statements using Semantic Scholar API with batch processing and checkpoint/resume

Wikidata Article Identifier Enrichment

Add external identifiers (OpenAlex, Semantic Scholar, Fatcat, PubMed) to scholarly articles

ORCID Author Data Query

Query the ORCID Public API for comprehensive author data (names, affiliations, education, employment)

ORCID Works Query

Retrieve complete publications from ORCID profiles, including non-CrossRef sources

CrossRef Author Works Query

Query CrossRef API for all works by an author via ORCID or DOI list with full metadata extraction