Dataset capturing teacher–AI co-design of STEAM projects and student use in bilingual K-12 classrooms. Includes: prompts, outputs, multimodal artifacts, and metadata: language use, automation levels, linked to learning outcomes.
Dataset capturing teacher–AI co-design of STEAM projects and student use in bilingual K-12 classrooms. Includes: prompts, outputs, multimodal artifacts, and metadata: language use, automation levels, linked to learning outcomes.
Project Details
Updated 07/06/26 · Provided via application · VerifiedWe are building the Erandi Aprende Learning Dataset, a multimodal bilingual dataset of PBL/STEAM learning across Mexico, the U.S., and Latin America, scaling from existing pilot data to 150,000+ complete teacher–AI co-design interactions, 1,000+ teachers, 50,000+ student artifacts, 10,000+ audio narrations, and 20,000+ timestamped decision sequences.
Beyond raw logs, we extract linguistic patterns, interaction dynamics, prompt strategy classifications, reliance trajectories, and outcome indicators, processed features that enable models to understand why interactions succeed and adapt scaffolding to teacher expertise, language profile, and context. The dataset spans Spanish-dominant, English-dominant, bilingual, and code-switching classrooms across grades 8–16, multiple STEAM domains, and diverse school contexts (public/private, urban/rural, resource-constrained). It captures solo and collaborative planning, open and scaffolded workflows, and feedback signals including accept/edit/reject decisions and edit magnitude.
Because it preserves full iteration paths rather than final outputs only, it enables fine-grained analyses of reliance patterns, temporal plan evolution, and equity comparisons, supporting models that personalize help: simpler scaffolds for novice teachers, Spanish-first instructions for emerging bilingual learners, and resource-light alternatives for under-resourced classrooms. This supports three concrete training targets: generating classroom-feasible projects, choosing the next-best scaffolding move, and producing bilingual adaptations that improve clarity and differentiation.
The project is led by a team combining complementary expertise. Andrea Remes (CEO) leads strategy and partnerships, with experience designing education programs across 60+ countries with the European Commission, Generation Unlimited, and Aspen Institute Mexico. Miroslava Rodríguez (CTO), AI Entrepreneur of the Year by Women in AI North America, and NASA STEM Girls Mentor of the Year, leads AI systems and dataset instrumentation, having built the capture infrastructure from scratch. Dr. Lucía Cárdenas Curiel, Associate Professor of Bi/Multilingual Education at Michigan State University, serves as Research Lead, bringing nationally recognized expertise in bilingualism and biliteracy. Mariana Ballardini, Data Scientist and AI Engineer, leads annotation frameworks and AI alignment, with her background in machine learning and her experience as a former school preceptor ensuring the data reflects authentic classroom realities.
Data will be collected through classroom pilots, annotated using structured rubrics co-designed by the research and technical teams, fully anonymized, and collected with informed consent. This will be all engineered from the ground up to be a catalytic resource for researchers, developers, and educators building AI that works in the messy, multilingual reality of real classrooms.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiThe most underappreciated x-risk in AI is not the systems that fail visibly, it is the systems that work perfectly for some and silently fail everyone else.
Today, the overwhelming majority of AI training data for education is English-dominant, Western, and built on classroom contexts that look nothing like a bilingual school in Guadalajara, a public school in South Texas, or a rural classroom in Chile. When AI systems trained on that data are deployed globally, they do not break, they just work worse for the students who already have the least. Bias at training time becomes inequity at scale, and at the pace AI is being adopted in education, that inequity compounds fast.
Erandi Aprende's dataset directly addresses this by introducing what is currently missing: a large-scale, multimodal, openly available record of how bilingual teachers and students actually interact with AI in real STEAM classrooms. By capturing Spanish-dominant, English-dominant, bilingual, and code-switching interactions across diverse school contexts, we give researchers and developers the raw material to build and evaluate AI that does not collapse when it meets a learner who switches languages mid-sentence or a teacher who has never used a digital tool before.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Funding Asks
Discussion
No comments yet. Be the first to share your thoughts.