Tag: evaluation
-
Could agents learn to work better from their own runs?
Most agent memory research is about remembering the user. I built a browser agent that mines its own traces instead, and the lessons that transferred were not the ones I expected.
-
I benchmarked 5 embedding models across 4 datasets
I benchmarked five embedding models across four NanoBEIR datasets and found that bigger embeddings did not always produce better retrieval.