Wikidata Tools is a collection of Jupyter notebooks for building, enriching, and maintaining scholarly records in Wikidata. Originally developed to improve coverage of anthropology journals, the tools are applicable to any academic discipline. All notebooks run in Google Colab without local setup.
The toolkit covers the full Wikidata scholarly pipeline: querying CrossRef and Wikidata to assess journal coverage, generating QuickStatements for bulk article import, verifying and repairing post-import records, detecting and resolving duplicates, and enriching author records with ORCID identifiers and institutional affiliations.
What’s included
Query & Coverage
Query anthropologists in Wikidata, count journal articles on the scholarly endpoint, track subject tagging, and rank researchers by composite prominence scores.
Import & Cleanup
Transform bibliographic CSV data into QuickStatements for batch import. Verify imports, detect duplicates by DOI and title matching, and generate cleanup commands.
Enrichment & Reconciliation
Add citations, external identifiers, and author data from Semantic Scholar, ORCID, and CrossRef. Reconcile author name strings to person items via ORCID matching.
Scholarly data workflows
Wikidata Anthropologist List
Query all anthropologists in Wikidata with metadata (birth/death, gender, citizenship, ORCID)
Wikidata Journal Article Counter
Count scholarly articles per journal on the Wikidata scholarly endpoint
Wikidata Anthropology Main Subject Counter
Count articles tagged with “anthropology” per journal, identify journals lacking subject tags
Wikidata Prominent Anthropologists
Rank ~13K anthropologists by composite prominence score (sitelinks, doctoral students, awards, authored works)
Wikidata Article Importer
Transform bibliographic CSV data into QuickStatements for batch importing scholarly articles
Post-Import Verification and Repair
Verify QuickStatements imports against source CSV, identify missing articles, generate repair statements
Wikidata Duplicate Detector
Detect duplicate scholarly articles by DOI matching and normalized title comparison
Wikidata Non-Article Cleanup
Identify non-articles in CrossRef CSV exports, look up QIDs, and generate deletion commands
Wikidata Author Name String Enrichment
Add qualifiers (series ordinal, affiliation) and CrossRef references to existing author name string statements
Wikidata ORCID Author Reconciliation
Match author name strings to person items using ORCID, generate QuickStatements for reconciliation
Wikidata Citation Enrichment
Enrich articles with citation statements using Semantic Scholar API with batch processing and checkpoint/resume
Wikidata Article Identifier Enrichment
Add external identifiers (OpenAlex, Semantic Scholar, Fatcat, PubMed) to scholarly articles
ORCID Author Data Query
Query the ORCID Public API for comprehensive author data (names, affiliations, education, employment)
ORCID Works Query
Retrieve complete publications from ORCID profiles, including non-CrossRef sources
CrossRef Author Works Query
Query CrossRef API for all works by an author via ORCID or DOI list with full metadata extraction