Anthropology of Machines: AI Interpretability for Social Sciences
A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.
A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.
Project Details
Updated 07/26/26 · Edited by orgSocial scientists mostly study latent constructs we cannot see directly. Institutions, norms, beliefs, ideology. Most of our training is about that problem: how to pin one down, how to build a measure for it, and how to check whether the measure means what we say it means. Now many of us work with language models. They label our texts, build our software, and are sometimes used as synthetic respondents. Most of the time we have no idea what happens between the prompt and the output — or whether the output means what we think it does. To the best of my knowledge, there is no accessible, hands-on book teaching mechanistic interpretability, with its premises and perils, to social scientists.
I have started that book to teach a social scientist what they can do to a model. I envision it as a short book, about 30,000 words, and I aim to publish it in an open-access way within a year with a publisher. The working draft is already public at github.com/tapanyemre/anthropology-of-machines, licensed CC BY-NC-SA / MIT.
I divided the content into three sections: observation, intervention, and validation. Observation tells you what you can find. Intervention tells you what the model actually uses. Validation tells you how you know either one showed you what you think it did. Alongside intuitive explanations and engagement with state-of-the-art work, I planned that every methods chapter would end the same way: what does this technique let you claim, and what does it not? The hands-on parts, wherever possible, do not require a personal GPU. That is required if we want the tools accessible to the ones who have the theory and the interest but not the resources.
Chapters will appear when they become readable, not when they are finished, so readers can follow the history of my writing and keep me honest. Readers can also contribute to the review process by pointing out missing or mistaken parts.
I did not start with those questions. I came to this by a side door. In summer 2025 I led the programming sessions of the AIDE summer program at Northeastern's Ethics Institute. One week, a guest lecturer from the NDIF team taught the logit lens — how a language model builds its prediction of the next word, layer by layer — through a library called nnsight, which let us reach inside a model running on their GPUs. Nobody in the room needed a GPU. That week changed what I thought these systems were. A model is not a black box anymore. It is just unread. Learning to read it is the hard part. In November 2025, David Bau emailed me about Neural Mechanics, a full semester of mechanistic interpretability. As the only political scientist invited, I proposed a project on the political ideology of models; it is a working paper now, and the book will contain this study as a worked case, from question to what the result licenses. I took notes the whole time, mostly about what I had to translate. This book is growing out of those notes. Though I am the solo author, I will make it open to everyone who would like to review and check the accuracy of claims.
A book without an audience will be only self-reflections of the author. When I was starting out in computational social science, I co-founded SICSS Istanbul, a two-week summer institute that trained around twenty early-career social scientists a year from 2019 to 2023, and I went back to teach there as a guest lecturer in 2026. With the same motivation, last fall I ran two afternoon workshops on the mechanics of LLMs and a hands-on guide to using them, funded by a small grant, and it brought an interdisciplinary team of social scientists together to learn the material I gathered from computer scientists. So when the manuscript and its hands-on tutorials are done and ready to teach — I expect by summer 2027 — I plan to organize a week-long bootcamp, open to both social scientists and computer scientists, to sit and work on their own projects while learning the tools and the intuition behind them.
Manifund application: https://manifund.org/projects/anthropology-of-machines-mechanistic-interpretability-for-social-sciences
Theory of Impact
Updated 07/26/26 · By grantmaking.aiI am a social scientist, not an alignment researcher, so let me put this in the terms I actually know.
Interpretability results are starting to carry weight in safety arguments. Whether a model is deceptive, whether a goal is represented inside it, whether a behaviour was removed or only hidden and these get argued from probes, features, and steering results. Every one of those is a measurement claim. And measurement claims fail in ways my field has spent a century cataloguing: the instrument decodes something, but not the thing you named it after; it works on your prompts and on nobody else's; a feature is there, and you report that the model uses it.
That is where the risk enters, and it enters quietly. A safety case built on an unvalidated instrument can be confidently wrong, and nothing in the result tells you so.
I am not going to claim that a book prevents a catastrophe. The chain is longer than that: the book teaches people, those people bring construct validity to evaluation and interpretability work, and that work gets harder to fool. What I will claim is that it is cheap, it stays free, and it is aimed at a large group who already have the training and are missing only the way in.
People
Updated 07/26/26 · By grantmaking.aiTeam Member
Funding Details
- Sep 1, 2026
- May 1, 2027
- 8 months
- -
- -
- -
- -
- -
- -
- -
Discussion
No comments yet. Be the first to share your thoughts.