While most research in the field of AI focuses on creating benchmarks, testing models, alignment, foresight, policy, and governance, there remains a clear gap in the empirical collection of information regarding the physical location of AI infrastructure and its interaction with the surrounding world—not only institutionally but also geographically. We lack sufficient data on how AI infrastructure is geographically distributed, how it affects local communities and the environment, where it is concentrated, and what the picture looks like at different scales—country, state, and county. These were precisely the questions I asked myself while working on my independent project—a geographical study of AI infrastructure. For example, The Epoch AI maintains datasets on GPU clusters and Frontier data centers, but these do not address their interaction with the outside world; they are simply published as-is. Semianalysis develops its own corporate tools, but they are very expensive and designed to solve business problems rather than serve the public good. Moreover, they view the issue at the national level, whereas increasingly heated debates and political battles are taking place at the subnational level (states, counties). The deployment of AI infrastructure and its geographical concentration are becoming important factors in politics and the economy. And I have seen firsthand just how little accessible and systematic data there is on this issue.
Even in the U.S., which has perhaps the most robust system of government statistics in the world, information has to be gathered bit by bit from hundreds of different websites, reports, and datasets. So I came up with the idea of creating this AI observatory, which would look at data at the state and county levels. Right now, I’m running this project on my own, in my spare time, and on a very limited budget. I actively use AI assistants to search for and process data, write scripts, and work on the web platform’s technical stack. I’ve developed a methodology for the Grid Constraint Index (GCI), a framework for data integration, and a web tool for visualizing and making the results publicly available. Currently, a minimum viable product (MVP) based on Virginia’s counties is ready. 133 Virginia counties and its equivivalents ingested into Supabase backend, methodology page published, validation experiment against verified Prince William County data center sites. At this stage, before scaling up further, I am focused on calibrating the methodology and testing the hypotheses that emerged during the research. To continue this work, I will need significant time and resources that will allow me to focus on this project.
Expected outputs. A public tracker spanning PJM and MISO (another ISO and RTO can be added in case of full-time work on the project); an open county-level compute-concentration dataset with a web tool; a published, citable methodology; and a research roadmap for the observatory's next stage.
Link to project MVP:
https://ai-buildout-frontier.vercel.app/
Link to GitHub:
https://github.com/Geofuturist/ai-buildout-frontier