grantmaking.ai Launch Round
At scry.io, I am giving people raw readonly Structured Query Language access to hundreds of terabytes of well-prepared internet data. I am ingesting much of the internet, over 1M records per second, and having them end up cleaned, dynamically embedded, and well-indexed in performant analytical databases that I give people safe, fair, ergonomic access to. I own significant hardware for these purposes, including over 400 TB of NVMe storage, so I am not under shallower corporate survival constraints like I'd be if renting a $20,000/month (and worse) server from cloud providers, but instead can design my APIs towards longer-time horizon, rarer, more differentiated purposes.
I have had over $1000 in revenue, and after a relaunch soon since significantly updating my hardware and database technology and completely rewriting my software stack, I expect this project to become profitable in the next couple months, while also becoming a public-benefit project backbone to many EA and AI-safety adjacent initiatives.
An important thing to note is the flexibility I have when working with data of this scale. When someone working in AI safety needs 5B quality text-embedding pairs for their research, it becomes trivial for them to explore exactly what they want and for me to upload it to their S3 bucket. When someone wants to mine for phenomenology reports on reddit historical archives, or basically do a vast assortment of internet text traversal operations, it becomes something people can just do by talking to their agent, and that I can just grant them credits, or they can pay 100-1000x cheaper pricing than Amazon Athena over Common Crawl.
Ultimately I intend to ship a robust enough API and service than many tools and startups can simply rely upon Scry as a research substrate. Epistemic infra operations requires agents, skills, and big data, and non-enterprise, solo people have struggled to have genuinely ergonomic, affordable access to big data affordances until now.
min: $50,000 will cover
- $12,000 HDD SAS storage backup for the 420 TB of NVMe storage I already have, as I've been in a precarious position pushing my hardware without local backups. B2 buckets just aren't viable at the scale I am already operating at.
- $4000 - cover internet proxies for 4 months while I continue building up a differentiated web crawl index.
- $10,000 - colocation cabinet expenses for 4 months, with 40 Gbps internet, so I can start reasonable exporting big datasets for vetted people who need them.
- $12,000 SF living for 4 months while I rush doing the deep things well but not accepting arbitrary VC money that isn't a good fit.
- $12,000 cover my AI subs for the next 6 months, which has risen to about 10 Max/Pro subs, using them in advanced workflows that allows me to deliver critical infrastructure as a solo operator.
$8M is the scale to really fill the OpenScry-shaped hole in the world, that existential philanthropy and intelligence explosion alignment ops really needs. Data and traversability of data has no substitute for many world-modeling and responding purposes. The world has been so GPU obsessed, it's essentially not woken up to intelligence paradigm of supporting optimal single-threaded programs over the data we care about, such as agents writing really smart SQL queries over PBs of cleaned subsets of Common Crawl. A very healthy position would involve about 20 PB of NVME storage via 336x 60 TB U.2 drives (sub $8000 per drive, $2.7M total), 20 storage nodes (EPYC 9755 CPUs with 24x128 GB RAM, about $100K each), 5 PB HDD backup, and 400 Gbps fabric. It's important to start preparing deep query substrates early as well, as data movement and cleaning at this scale can take many months. If we want our brightest researchers leveraging Fable 6.5 in November to execute the smartest queries over the internet to unearth hidden insights about the shape of the world, we have to be preparing now.
I just want to add that ergonomic access to big data genuinely has gravity. People just have to extrapolate and understand that IT IS meaningful to have hundreds of terabytes of the most thoughtful writing and work and metadata about it, in a database that agents have low-friction access to. And that people just have to understand that they will experience the benefits of scry and scry links to query results sooner than later. I'm doing the deep things well, this tool already trivially allows people to research countless hypotheses ergonomically that they never could, like if people who mention certain health interventions have more diverse writing as calculated by embedding distances.