Hi Loek,
Thank you for your question and for your interest in lowering the technical barriers for independent researchers.
The smallest bounded evaluation object that an outside project could bring to EvalLabs today is a single AI-assisted transformation. In practice, this can be as small as one claim or passage, or as large as an entire document, provided there is an original source against which the generated output can be evaluated.
From that evaluation, EvalLab produces an Evaluation Record - the primary research artifact of the platform. The Evaluation Record captures the complete context of a single evaluation, including the source material, prompt, AI-generated output, evaluation findings, supporting evidence, and experimental configuration. Rather than serving as a simple score report, it provides a structured record that can be shared, archived, cited, or included as supporting evidence in future research.
When human validation is performed, EvalLab can also generate a Human Review Packet for external reviewers. Reviewer feedback is then incorporated into the Evaluation Record, creating a documented history of both automated and human assessment. EvalLab also generates experiment-level analyses across multiple Evaluation Records that identify recurring failure patterns, compare prompting strategies, and measure agreement between automated and human evaluations.
The intent is that another researcher should not have to rely solely on the original evaluator's interpretation. Instead, they receive the evidence and methodological context needed to inspect how an evaluation was conducted, repeat the evaluation process using the documented methodology, or challenge individual findings based on the preserved source material, AI output, and supporting rationale. Because the Evaluation Record is designed as a durable research artifact, it can be archived, cited, incorporated into reports or white papers, and used as a foundation for future research.
Regards,
Valencia
Hi Valencia,
I really appreciate the focus on lowering the technical barrier for independent researchers while keeping evaluation work systematic and reproducible.
What is the smallest bounded evaluation object that an outside project could bring to EvalLabs today — for example one claim, one system behavior, or one small dataset — and what exact artifacts would EvalLabs return?
I am especially curious what another independent researcher would receive in order to reproduce, inspect, or contest the resulting evaluation without relying only on the original evaluator’s description.
greetings
Hi Loek,
Thank you for your question and for your interest in lowering the technical barriers for independent researchers.
The smallest bounded evaluation object that an outside project could bring to EvalLabs today is a single AI-assisted transformation. In practice, this can be as small as one claim or passage, or as large as an entire document, provided there is an original source against which the generated output can be evaluated.
From that evaluation, EvalLab produces an Evaluation Record - the primary research artifact of the platform. The Evaluation Record captures the complete context of a single evaluation, including the source material, prompt, AI-generated output, evaluation findings, supporting evidence, and experimental configuration. Rather than serving as a simple score report, it provides a structured record that can be shared, archived, cited, or included as supporting evidence in future research.
When human validation is performed, EvalLab can also generate a Human Review Packet for external reviewers. Reviewer feedback is then incorporated into the Evaluation Record, creating a documented history of both automated and human assessment. EvalLab also generates experiment-level analyses across multiple Evaluation Records that identify recurring failure patterns, compare prompting strategies, and measure agreement between automated and human evaluations.
The intent is that another researcher should not have to rely solely on the original evaluator's interpretation. Instead, they receive the evidence and methodological context needed to inspect how an evaluation was conducted, repeat the evaluation process using the documented methodology, or challenge individual findings based on the preserved source material, AI output, and supporting rationale. Because the Evaluation Record is designed as a durable research artifact, it can be archived, cited, incorporated into reports or white papers, and used as a foundation for future research.
Regards,
Valencia
Thank you, Valencia.
Your answer gave us exactly the kind of opening we were hoping for: the Evaluation Record as a durable research artifact rather than only a score or conclusion.
You have now shown us your artifact structure, so before asking anything more of you, let me first put ours on the table.
In our own source-bound transcription work, we have been separating the following layers rather than allowing them to collapse into one final “correct” output:
the original source remains unchanged and is bound by provenance, segment boundaries and a cryptographic hash;
the raw model output is preserved exactly as produced, together with its model, configuration and run context;
uncertain passages remain explicitly marked rather than being silently completed;
a human correction does not overwrite the model output, but exists as a separate review layer;
differences between versions remain visible;
and where the source does not support a reliable reading, the result remains unresolved rather than being forced into a conclusion.
From that work, our smallest candidate evaluation object would be:
one open-licensed spoken-audio fragment;
one exact ASR run over that fragment;
one source-bound claim: “this transcript faithfully represents the recording and makes its uncertainty visible”;
one independently prepared human review;
and an explicit difference record showing where the machine output, human reading and available evidence agree, disagree or remain uncertain.
We would provide the source, license, hashes, raw run data, transcript, review and difference record together. That is the evidence package we would place into the exchange—not only an input and a request for your interpretation.
What interests us is then not whether EvalLabs gives the same conclusion we do, but how your Evaluation Record carries the same object: what it preserves, what it distinguishes differently, and where either structure sees something the other does not.
Let me know if this works for you :D
But anyway, many thanks for this open dance !!
Loek
Hi Loek,**
Thank you for taking the time to share your own evaluation object. I really enjoyed reading through it.
One thing that stood out to me is that, although we're working in different domains, we're converging on a very similar philosophy. Rather than treating an AI-assisted process as the end product, we're both trying to preserve the evidence, context, and reasoning that led to it.
Your description resonated with me because it emphasizes something I think is often overlooked: the conclusion is only part of the story. Equally important is preserving the evidence and context that allow someone else to understand how that conclusion was reached.
What I find particularly interesting is shifting the comparison away from outcomes and toward the underlying structure. Comparing what each artifact chooses to preserve tells us a great deal about the questions each framework is trying to answer.
So yes, I think your proposed exchange works very well. I'd be happy to compare one of your evidence packages with an Evaluation Record from EvalLab and explore where the two approaches align, where they differ, and what each framework makes visible. I suspect there are useful ideas to learn in both directions.
Thank you again for the thoughtful exchange. I think this could be a genuinely interesting comparison, and I'm looking forward to seeing where it leads.
Best,
Valencia
@Valencia Cooper
Hi Valencia,
Thank you — this is exactly the kind of comparison I was hoping for.
In my opinion...Let’s keep the first exchange deliberately small. I will prepare one short, openly licensed audio fragment and a fixed evidence package containing the exact source and hashes, one fixed ASR run and its raw output, visible uncertainty, a separately prepared human review, and an explicit machine–human difference record.
The question will not be whether EvalLab confirms our conclusion. It will be how each artifact represents the same transformation, and what each preserves, separates, collapses, or leaves unresolved.
One suggestion for the method: perhaps we each first write down our own comparison independently, before sharing interpretations, so neither reading pulls the other toward consensus.
Is there a preferred intake format or minimum schema for creating an Evaluation Record in EvalLab? Otherwise I can prepare a small ZIP with a README and a machine-readable manifest.
Thanks again — I see and feel this can be a very clean and useful first dance comparison :D :D :D.
Loek
@Loek Verdonk
That sounds like a great approach, and I agree that keeping the first exchange intentionally small makes sense.
I also like your suggestion that we each document our observations independently before discussing them. I think that will make for a much cleaner comparison.
A ZIP package works well. I'll create a private shared folder where you can upload it once it's ready, and I'll work from that for the EvalLabs side of the comparison. Could you send me the email address you'd like me to use for sharing the folder?
Looking forward to seeing your evidence package.
Best,
Valencia
@Valencia Cooper
Hi Valencia,
Thats perfect, thank you :D
Please use: loekverdonk@live.nl
Once the private folder is shared, I’ll upload the fixed evidence package there. It will contain the openly licensed source fragment, license and hashes, the exact ASR run and raw output, visible uncertainty, the separate human review, the machine–human difference record, a README, and a machine-readable manifest.
I’ll also prepare our initial observations separately, so we can compare both structures before either interpretation influences the other.
Really looking forward to this clean first comparison.
Best,
Loek
@Loek Verdonk
Perfect, thank you! I've just shared the private folder with loekverdonk@live.nl.
Take your time putting the evidence package together. The structure you described sounds ideal for this first comparison, and I like the idea of us documenting our observations independently before sharing interpretations. I think that will make for a much cleaner and more meaningful comparison.
I'm looking forward to exploring how the two artifacts represent the same transformation and seeing what we can learn from the exercise.
Best,
Valencia
@Valencia Cooper
Hi Valencia,
The first fixed evidence package is ready and has been uploaded to the private shared folder.
One source-boundary check before you begin: I originally said I would use an openly licensed audio fragment. The source I selected is a LibriVox recording that LibriVox treats as public domain in the United States, while also advising users outside the US to verify the local legal status. I included that caveat clearly in the package and did not claim universal redistribution rights.
Would that source boundary be acceptable for this private comparison?
If you would prefer an explicitly Creative Commons–licensed recording instead, that is completely fine. I will calmly replace the audio and upload a corrected package tomorrow.
The current package keeps all layers separate and marks the blind human review as
NOT_RUNrather than fabricating it. It also contains no initial interpretation from our side, so we can still document our first observations independently.ZIP SHA-256:
0e1af2eca0b1948600a03758e501db0f8865638369c8eba35eb4e59b55189020Thanks again 😄
Loek
@Loek Verdonk
Thank you! I received the package and appreciate the care you've taken in documenting the provenance and source boundaries.
For the purposes of this private comparison, I'm comfortable proceeding with the LibriVox recording as you've provided it. I appreciate that you've documented the licensing caveat rather than overstating the redistribution rights; that level of transparency is exactly the kind of provenance we're both interested in preserving.
I also noticed that you've explicitly marked the blind human review as NOT_RUN rather than filling in information that doesn't yet exist. I think that's an excellent example of preserving the state of the evidence, and I'm glad we're approaching this comparison with the same philosophy.
I'll begin reviewing the package and will document my observations independently just as we discussed. I'm Looking forward to comparing the two artifacts.
Best,
Valencia
@Valencia Cooper
Thanks for the positive vibe for this dance :D!
Good luck and have fun :D
@Loek Verdonk
Thank you again for preparing such a thoughtful evidence package. I found the exercise genuinely valuable.
Following your suggestion, I intentionally completed an independent first read before applying EvalLab's evaluation pipeline or receiving your interpretation. I wanted to preserve an unbiased initial comparison of the methodology and the artifacts themselves.
After that first pass, I preserved your package unchanged within EvalLab and generated an independent Evaluation Record as an additional evaluation layer rather than modifying the original evidence.
I've attached three artifacts to the shared private drive:
One observation that stood out to me was how strongly our methodologies independently converged on preserving evidence rather than only preserving conclusions. Where I think they differ most is in emphasis: your package carefully separates preservation, measurement, and interpretation, while EvalLab adds a structured evaluation layer over preserved evidence. I found those differences complementary rather than competing.
I'm looking forward to seeing your own interpretation whenever you're ready. I suspect comparing the two independent readings will be every bit as interesting as comparing the resulting artifacts. Thank you again for the opportunity to participate in this exchange. I have a feeling this is the beginning of a very interesting conversation.
Best,
Valencia
@Valencia Cooper
Hi Valencia,
Thank you again for the care you brought to this comparison. Your final sentence genuinely moved me, because the feeling is the same here: this has been a beautiful and unusually thoughtful exchange.
Precisely because of that, I need to acknowledge a serious mistake on my side.
We explicitly agreed that we would each complete and freeze our own observations before seeing the other person’s interpretation. You followed that procedure carefully. I did not. I read your returned first-read observations and Evaluation Record before I had frozen my own comparison.
I am genuinely sorry.
For me, this is an uncomfortable but important lesson. For you, however, it may be a real disappointment, and it may feel as though the care you put into preserving the independent-first procedure was not met with the same discipline from my side. It may also have created uncertainty or unnecessary additional work in a comparison that we had both approached very carefully.
I want to be completely clear that this does not reduce or contaminate your work. Your independent first read remains exactly what you intended it to be. The procedural failure occurred entirely on my side.
I cannot undo the exposure, and I will not write something afterwards and present it as an independent first read. I have still completed the comparison we agreed to make, but I have labeled it explicitly as post-exposure and not independent. I also preserved the protocol failure itself as part of the record rather than quietly correcting or hiding it.
There is no need for you to repeat your work or repair anything for me. You fulfilled your side of the agreement with care.
I fully understand if this changes how you would like to continue. We can treat the current comparison as an openly labeled post-exposure structural comparison, or at some later point restart the symmetrical independent-first method with a new small object. There is no pressure in either direction.
I am sorry that I did not meet the procedural standard that you did. Thank you for taking both the evidence and the method so seriously.
The mistake is mine. The lesson is useful. And the value of what you returned remains fully intact.
Warm regards,
Loek
@Loek Verdonk
Thank you for your note. I really appreciate your honesty, and please don't worry.
I don't think this takes away from our exchange at all. If anything, I think the fact that you documented what happened and openly shared your sentiments is admirable; very methodical and thay even more clear with everything I saw in your evidence package.
My independent first read is still exactly what it was intended to be, so I'm perfectly happy to treat this as a post exposure structural comparison rather than starting over. The important thing is that it's clearly documented, and I think we've both handled it transparently.
Honestly, I got a lot out of this exercise. I was curious to see how EvalLabs would interact with an external methodology, but I came away just as interested in how our approaches independently arrived at similar ideas around preserving evidence and separating measurement from interpretation. I wasn't expecting that, and I found it really encouraging.
I've genuinely enjoyed this exercise and hope we can continue it. As EvalLab evolves, I'd love to keep exploring how independently developed evaluation methodologies can learn from one another. Thank you again, and I'm looking forward to reading your comparison.
Best,
Valencia
@Valencia Cooper
Thank you again for such a warm and generous response. It really helped me move from feeling awkward about the mistake toward seeing what the exchange was actually teaching us :D
What stayed with me most was your observation that our independently developed approaches arrived at similar ideas around preserving evidence and separating measurement from interpretation. I think that is the real value of this first dance: not that our structures match perfectly, but that each one makes a different part of the same transformation visible.
I have now uploaded the reviewed post-exposure structural comparison, together with a small Cross-Artifact Mapping Sheet candidate.
The comparison remains clearly labelled post-exposure. The mapping is not presented as a shared standard or as an authoritative description of EvalLab. It is simply my attempt to place the two artifact structures beside one another and make visible what each preserves, separates, adds, transforms, or leaves unresolved. I would genuinely appreciate any correction wherever I may have misunderstood or inaccurately represented the EvalLab side.
One small verification question also became visible during the comparison: the HTML and JSON representations show different checksum values. I have recorded that only as a question about what exact representation or field set each checksum binds—not as a conclusion that either record is invalid.
File uploaded:
VALENCIA_EVAL_LABS_POST_EXPOSURE_DELIVERY_V0_FINAL.zipSHA-256:
89706ecf86311773f39e46a4f48a93b596250968f319b2d43c6dd69cdd464b43Thank you not only for the work you put into this comparison, but also for the openness with which you received the mistake. You helped turn it into part of the method rather than something that needed to be hidden.
This really does feel like the dance of growth to me: preserving the mistakes, learning through one another’s lenses, and allowing the next artifact to become a little more honest because of what the previous one revealed :D :D :D
I have genuinely enjoyed this first exchange and would love to keep exploring where this dance leads as EvalLab evolves.
Best,
Loek
@Loek Verdonk
Thank you for putting so much care into this. I spent some time looking through the package, and I have to say I'm genuinely impressed.
What stood out to me most is that this doesn't read like a review of EvalLabs as a software tool; it reads like a careful comparison between two independently developed research methodologies. I wasn't expecting that level of depth, and I really appreciate the thoughtfulness you've brought to it.
I also appreciate how carefully you've distinguished between observations, open questions, and conclusions throughout the comparison. The boundaries you've placed around the claims make it very easy to see what the artifacts support and what remains intentionally unresolved.
Regarding the checksum question, thank you for catching that. After looking into it, you had identified a genuine ambiguity in what the checksum represented. It turned out the checksum included export-time metadata, so the HTML and JSON naturally produced different values even though they represented the same Evaluation Record. I've since updated EvalLab so the checksum now binds only to the evaluation content itself rather than export-specific metadata. If it would be helpful for your records, I'm happy to regenerate the Evaluation Record with the updated checksum behavior and send you the revised HTML and JSON exports. Just say the word and I will upload it to our shared drive.
Over the next few days I'd like to go through the Cross-Artifact Mapping Sheet carefully, row by row. I already have a few thoughts, but I'd like to give it the same level of care that you've clearly put into preparing it before offering any corrections or clarifications.
Thank you again for such a thoughtful exchange. I feel like we've both come away with a better understanding of our own methodologies, and I'm excited to continue the conversation.
Regards,
Valencia
@Valencia Cooper
@Valencia Cooper
Hi Valencia,
Thank you — this is genuinely wonderful to hear. :D
I am especially grateful that you investigated the checksum question rather than treating it as merely a difference between export formats. The fact that it revealed a real ambiguity in what the checksum bound, and that you have now changed EvalLab so it binds to the evaluation content rather than export-specific metadata, feels like a very meaningful outcome of this first comparison.
Yes please, the regenerated HTML and JSON exports would be extremely helpful for our records.
If possible, could you leave the original exports in the shared folder and add the revised versions beside them rather than replacing them? A small note stating:
that these revised exports supersede the earlier export behavior;
what exact content or field set the new checksum binds;
and whether the HTML and JSON now intentionally expose the same content checksum;
would give us a very clean predecessor–successor comparison without erasing the original ambiguity that led to the change.
There is absolutely no rush on the Cross-Artifact Mapping Sheet. Please take the time you need to go through it row by row. Corrections are much more valuable to me than quick agreement, especially wherever my mapping may describe the EvalLab side too narrowly or inaccurately.
What I find beautiful here is that neither methodology needed to become the other. Placing the artifacts beside one another simply made something visible that was difficult to see from inside either structure alone.
Thank you for receiving the question so openly and then carrying it back into EvalLab itself. That makes this feel less like a review being completed and more like the beginning of a genuine learning loop between two independently developed methods.
Very excited to continue this dance. :D :D :D
Best,
Loek
@Loek Verdonk
Thank you so much for your patience. I've uploaded the updated HTML and JSON Evaluation Records with the corrected checksum implementation, along with my review of your Cross-Artifact Mapping Sheet and my comments. Everything has been uploaded to our shared folder.
I really appreciate the time and care you put into this comparison. Your observations, particularly around the checksum behavior, helped improve EvalLab, and I found the mapping to be a thoughtful and constructive review of the methodology.
Please let me know if you have any questions or if you'd like to discuss any of the review comments further. I'd be happy to continue the conversation, and I hope we have the opportunity to collaborate again in the future.
Best,
Valencia
@Valencia Cooper
Hi Valencia,
Thank you again. Your invitation to discuss the review comments further is very much appreciated.
After receiving the revised exports, I placed the predecessor and successor records beside one another and carried out one additional, bounded inspection. The checksum repair is clearly visible: the updated HTML and JSON now expose the same checksum, while the original records remain preserved as predecessors.
Interestingly, the repair also made a deeper integrity layer inspectable. Three narrow follow-on questions became visible around:
A separate predecessor question also appeared around whether the original HTML and JSON carried fully equivalent visible narrative content.
I have kept these observations on internal hold rather than treating them as conclusions or immediately sending a large follow-up. They do not take away from the repair; quite the opposite. The fact that the successor artifacts allow the repair itself to be inspected feels like a strong property of the process.
Should it be useful, I can share a single-page note containing the exact observations, their claim boundaries, and one small verification question. There would be no expectation of additional implementation work — correction or clarification of our reading would already be valuable.
Thank you again for leaving the original and revised artifacts side by side. That preservation is what made this next layer visible.
The dance floor is open again :D
Best,
Loek
@Loek Verdonk
I'd definitely be interested in reading the note. Please feel free to share it! I appreciate that you've continued to keep the observations bounded and clearly distinguished from conclusions. One of the things I've enjoyed most about this exchange is that it's helped strengthen not only EvalLab's implementation but also the clarity of the underlying methodology and documentation. Even when no implementation changes are needed, clarification of the methodology is valuable.
Best,
Valencia
@Valencia Cooper
Thanks for the kind words :D — though it's your own honesty, shining on the exact spots that strengthen the whole, that made this work. Big thanks for that too! The dance continues then — awesome.