grantmaking.ai Launch Round
LexiconForge is reimagining what translation can be. Rather than a single replacement text, we have a inspectable just-in-time interface.
Imagine you want to read a piece of text in a language you are unfamiliar with. You either have to trust the translator and just consume it in English or pull out a dictionary and struggle with the nuances of the grammar and multiple meanings of the source text words.
With lexicon forge you can participate in moving along this spectrum and progressively ask for and obtain details on pronunciation, cycle through meanings, see alignment between meaning chunks.
The goal is to build a inspectable translation interface that makes provenance, omission, conflicts between scholarly interpretations all transparent to the reader.
This will be used not just for languages but even for translation between subcultures that use different jargons within the same language of English (imagine the e/acc, ai ethics, ea worldview people trying to talk)
I am starting the auditing with translation between different languages like Malayalam, Pali, Chinese which have decently good inspectable evidence to see how well current models can assemble the language pair specific interfaces that keep the alignment at different levels - sound, morphemes, word, phrase, concepts, textual witnesses. Surfacing expert disagreement and a long history of debate.
https://read.adityaarpitha.com/sutta/mn10
LexiconForge is a step towards this JIT vision for interoperation between contexts by first creating distinct interfaces for language pairs that respects the distinctive feature of each language and benchmarking all the current model's on its ability to do a good job faithfully serving as this interoperation bridge.
https://read.adityaarpitha.com/bench/sutta-studio
These models are therefore being evaluated on their capability to be reliable interface compilers. We check if every part of the source survived - was it complete? Are the proposed alignments structurally and semantically plausible? Were the factual claims adequately grounded with citations? were disagreements and cruxes preserved? and so much more.
All this work is open source here
I will be running the benchmark on most of the frontier models and paying myself for the hours I invest into improving the code
https://docs.google.com/spreadsheets/d/1v1LTaC5z5kSFRo-DkNZUMor8vDnhnHQw36l-SHyNcek/edit?usp=sharing