In November 2025, I founded the milestone-based Global AI Dataset (GAID) Project. I set up the site AI in Society to release subsequent GAID Project outputs. I carried out data-engineering work to harmonise, compile, curate and document the GAID w1 (v1 and v2) datasets (deposited on Harvard Dataverse; which GAID dataset is updated annually at the end of each year). GAID w1 v2 dataset contains over 24,000 AI indicators, covering 227 countries/territories across 20 AI domains from 1998 to 2025. Recently, I finished building Phases 3 and 4 of the GAID Project upon my GAID w1 v2 dataset. Phase 3 is to programmatically present the GAID AI Development Index ranking global countries’ AI capacity that covers multiple thematic pillars. Phase 4 is a model-evaluation benchmark that scores 11 frontier LLMs’ (both open-weight and proprietary) responses against GAID ground-truth data. The goal is to evaluate which LLMs are subject to higher fabrication rates, and also whether and which LLMs are more likely to fabricate in the contexts of lower-income countries relative to their upper-income counterparts (so as to address AI-driven digital colonisation concerns). I built Phase 4 upon two peer-reviewed pilots, one published on Apart Research and the other accepted by IEEE IES for conference presentation at the IEEE IRAI 2026 in Melbourne this September.
I would like to apply for 5-month funding to scale up both Phases 3 and 4:
-
Phase 3: Harmonise and map GAID with global panel socioeconomic data (such as World Bank Open Data), and publish the compiled and documented dataset on Harvard Dataverse for public download and/or API extraction, so as to facilitate worldwide research on how AI capacity, adoption, safety, readiness and beyond relate to socioeconomic development in global contexts
-
Phase 4: Publish the GAID w2 dataset in late 2026 (expected to cover data across 30+ AI domains); extract more indicators from GAID w2 dataset across the thematic pillars (AI accountability, adoption, ethics, fairness, regulation, safety, security, transparency); stress-test a wider set of Western and Chinese frontier LLMs (both open-weight and proprietary)
-
Additional Outputs: Currently in discussion about collaboration with senior Cambridge academics within my academic network to co-author two research papers, one introducing the open-access GAID methodology and AI Development Index and the other presenting the open-source model evaluation methodology and empirical findings
Note: I only cover publicly accessible, non-paywalled data in this GAID Project. As always, all code is and will be open-sourced.