Implementing Hybrid Reward Architectures (HRA) to build robust internal factual grounding and quantitative variance metrics for multimodal agents, moving beyond fragile single-seed evaluations.
Implementing Hybrid Reward Architectures (HRA) to build robust internal factual grounding and quantitative variance metrics for multimodal agents, moving beyond fragile single-seed evaluations.
Project Details
Updated 07/07/26 · Provided via application · VerifiedThis project aims to systematically address the structural fragilities of current AI evaluation metrics by developing an internal Hybrid Reward Architecture for multimodal agentic systems. Standard single-seed evaluations often mask model instability, reward density collapse, and deceptive behaviors such as reward hacking or specification gaming.
I will engineer and test an HRA framework using LLaVA-1.5-7B to investigate whether internal reward structure can improve factual grounding more reliably than external patch-based methods such as standard RAG wrappers. In parallel, I will design evaluation pipelines that measure variance, instability, and reward hacking across complex agentic workflows.
The project will be led by me, Teganmosibineba Oluwatofarati Jegede. I bring hands-on experience in building autonomous agents and evaluation pipelines using PyTorch, LangChain, and Python, alongside product management experience that has strengthened my ability to work with efficiency metrics and structured project delivery.
The deliverables will include an open-source evaluation suite for hallucinatory and deceptive behavior in agentic systems, a documented HRA implementation for LLaVA-1.5-7B, and a research paper suitable for submission to a major AI safety conference.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiExistential risk from highly capable autonomous AI systems increases when we cannot reliably measure whether they are behaving in aligned ways during training and evaluation. When systems are tested with sparse, single-seed evaluations and protected mainly by brittle external guardrails, dangerous behaviors such as reward hacking, deceptive alignment, and systematic hallucination can remain hidden until deployment.
This project reduces that risk in two ways. First, by investigating Hybrid Reward Architectures, it aims to improve factual grounding and behavioral reliability from within the model’s training signal rather than relying only on external patching. Second, by publishing quantitative variance metrics and multi-seed evaluation methods, it gives researchers and regulators better tools for detecting instability and deceptive optimization before these systems are scaled into high-stakes settings.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.