Can an AI Agent Learn Without Fine-Tuning Its LLM?
A small replication of Memento-style agent memory, testing whether a frozen research agent improves by retrieving and reusing previous task-solving experiences.

Can an AI Agent Learn Without Fine-Tuning Its LLM?
What if improving an AI agent did not require changing the model at all? No fine-tuning, no reinforcement learning over the model weights, and no new checkpoint. Instead, the agent gets better because it remembers what happened before and learns how to reuse previous experience.
That is the idea I wanted to explore after reading Memento: Fine-tuning LLM Agents without Fine-tuning LLMs.
I built a small-scale replication around the paper's core question:
Can a frozen research agent perform better simply by retrieving and reusing previous task-solving experiences?
The answer, at least in this pilot, was yes. Across 40 tasks, the baseline agent without case retrieval achieved an Exact Match score of 32.5%. With non-parametric Case-Based Reasoning, using essentially the same agent and models, Exact Match increased to 40.0%. F1 improved from 47.1% to 52.6%.
Nothing about the underlying language models changed. What changed was what the agent could remember.
The distinction that interested me
A lot of work on improving language models starts with the model itself. You collect more data, fine-tune, apply reinforcement learning, update weights, and produce another model checkpoint. That is model learning.
Once we start building agents around these models, however, there is another place where learning can happen: the agent itself. An agent has more than an LLM. It can have tools, a planner, an executor, memory, previous trajectories, feedback, retrieval, and an environment.
Conceptually, model learning looks like:
experience
↓
update weights
↓
new model
Agent learning can instead look like:
experience
↓
store memory
↓
retrieve useful experience
↓
better future behaviour
The second path is what I wanted to test.
Memento
Memento approaches this problem through Case-Based Reasoning, or CBR. The idea behind CBR is old and intuitive: when humans encounter a new problem, we often think about similar problems we have solved before. We retrieve an experience, reuse what worked, adapt it to the current situation, and then retain what happened as another experience.
The loop looks roughly like:
Retrieve
↓
Reuse
↓
Revise
↓
Retain
Memento applies this idea to LLM agents. Instead of fine-tuning the underlying LLM every time the agent gains experience, the system maintains external memory containing previous task-solving cases.
A case can include things such as:
task
plan
trajectory
final answer
reward
When a new task arrives, the system retrieves useful past cases and exposes them to the agent. The model weights remain frozen, but the memory changes.
The agent I built
I kept the implementation deliberately small. The agent had two main model-driven components: a planner and an executor.
[ Question ] → [ Planner ] → [ Executor ] → [ Tools ]
▲ │
│ ▼
[ Memory ] ← [ Evaluation ] ← [ Answer + trace ]
Tools: Search, Read, Python
Memory: task, plan, trajectory, answer, reward
The planner receives the question and creates a short research plan. The executor follows the plan and decides when to use tools. The tools available to the executor were web search, webpage reading, and a small Python environment for calculations.
For the models, I used GPT-4.1 as the planner and o3 as the executor. Both remained frozen throughout the experiment.
Building Case Memory
After completing a task, I converted the run into a case. A simplified representation looked like this:
{
"task": question,
"plan": plan,
"trajectory": trajectory,
"final_answer": answer,
"reward": reward
}
I also generated an embedding for the task. This gave each previous experience both structured information and a vector representation that could later be used for retrieval.
The Case Bank therefore grew as the agent processed more tasks:
Task 1
→ solve
→ store case
Task 2
→ solve
→ store case
Task 3
→ solve
→ store case
The difference between the experimental conditions was whether those previous cases were actually retrieved.
No CBR
The first condition was the baseline. The agent still solved every task and stored the resulting cases, but none of the cases were retrieved when solving future questions.
Conceptually:
Task
↓
Planner
↓
Executor
↓
Tools
↓
Answer
Past experience existed, but the agent could not use it. In code, this was simply:
if condition == "no_cbr":
retrieved = []
This gave me a baseline against which memory-based retrieval could be compared.
Non-parametric CBR
The second condition enabled Case-Based Reasoning. For every new task, I embedded the question and compared it against the embeddings of previous cases. The closest cases were then retrieved using cosine similarity.
Conceptually:
New task
↓
Embed task
↓
Compare with Case Bank
↓
Retrieve similar cases
↓
Planner
↓
Executor
No additional model had to be trained for this retrieval mechanism, which is why I refer to it as non-parametric retrieval. The planner could now see previous cases before creating its plan, while the rest of the agent remained as similar as possible between the two conditions.
The evaluation setup
This was a small pilot rather than a full reproduction of the paper. I used four datasets for the experience stream: Natural Questions, TriviaQA, HotpotQA, and 2WikiMultihopQA.
I sampled 10 questions from each dataset:
10 Natural Questions
10 TriviaQA
10 HotpotQA
10 2Wiki
Total = 40 tasks
The random seed was fixed at:
seed = 42
Both experimental conditions used the same tasks and the same ordering. I also kept the planner model, executor model, prompts, available tools, tool limits, evaluation functions, and dataset ordering fixed.
The major variable that changed was retrieval:
No CBR
vs
Non-parametric CBR
Measuring the answers
I tracked four values: Exact Match, F1, reward, and execution success.
Exact Match checks whether the normalized predicted answer exactly matches one of the benchmark references. For example:
prediction:
Handley Page Limited
reference:
Handley Page Limited
produces:
EM = 1
F1 gives partial credit when answer tokens overlap. This mattered in examples such as:
Prediction:
Kurt Weill
Reference:
Kurt Julian Weill
Exact Match gives this:
EM = 0
but token-level F1 gives:
F1 = 0.8
That distinction became useful because strict benchmark matching sometimes hid answers that were substantially correct.
Reward was deliberately simple. If Exact Match was 1, reward was 1. Otherwise reward was 0.
Execution success was kept separate from benchmark correctness. A task could be successfully researched and answered while still receiving reward 0, so:
success = True
reward = 0
was a perfectly valid outcome. That distinction turned out to matter.
Building the executor was harder than expected
The most interesting part of the replication was not the retrieval code. It was getting the research agent to behave reliably.
My first executor could search the web and often arrive at the right answer. But when I inspected the trajectory, I noticed something wrong: it was frequently doing this:
search
→ search
→ answer
It was not actually opening the sources. That meant an answer could look grounded even when the model was relying mostly on search snippets or its own prior knowledge.
I changed the executor flow so that after a successful search, the next useful action had to be a source read. The intended behaviour became:
search
→ read
→ reason
→ search again if necessary
→ read
→ answer
Fixing that exposed several other problems.
The forced final answer bug
At one point the agent would sometimes perform one or two searches, fail to gather enough evidence, and still return a confident answer.
The reason was a fallback function I had added earlier. When the research loop ran out of actions, the system called a separate force_final_answer() function.
That meant the architecture was effectively doing:
research fails
↓
ask model to answer anyway
That defeated the purpose of building a grounded research trajectory. I removed the fallback completely. If the agent could not gather enough evidence, it now failed cleanly instead of manufacturing an answer.
Tool calls were being counted incorrectly
Another bug appeared during the first multi-task smoke test. I had configured:
max_tool_calls = 10
I assumed this meant the agent could perform 10 actual tool calls, but the loop was counting every executor decision. That included rejected final answers, malformed actions, and actions that violated the search-to-read rule.
As a result, an agent could reach its supposed 10-tool limit after only one or two actual tool calls.
The fix was to separate executor decisions from tool executions:
tool_calls = 0
and only increment the counter after:
result = use_tool(...)
tool_calls += 1
After that change, questions that previously failed started completing successfully. A three-question smoke test went from only one successful execution to all three completing.
Then the prompts became too large
Once the agent started performing longer research trajectories, another problem appeared. The executor received the entire trajectory every time it had to choose its next action, including full webpage contents.
Eventually o3 rejected one request:
TPM limit: 30,000
Requested: 38,992 tokens
The issue was not the question itself. It was accumulated research context. Each executor call was effectively sending the question, plan, all searches, all snippets, all webpage contents, and feedback again.
I kept the complete trajectory for logging, but created a compact representation for the executor. Search results were reduced to the first few results, while webpage contents were truncated to a few thousand characters.
The full evidence still existed in the experiment logs, but the model no longer needed to repeatedly ingest all of it. That solved the context explosion.
Model output is not always clean JSON
Another assumption I had to remove was that a model instructed to return JSON would always return only JSON.
Sometimes the executor returned text like:
JSON in response should be:
{
"action": "tool",
...
}
instead of the raw JSON object.
Other times it returned:
null
or a final action containing:
{
"answer": null
}
All of these initially crashed the experiment. Eventually the executor parser became tolerant enough to extract JSON objects from surrounding text, reject non-dictionary responses, retry malformed decisions, convert None answers into empty answers, and fail an individual task instead of crashing the entire condition.
This became particularly important once each condition started taking tens of minutes to run.
Checkpointing became mandatory
The first few runs were small enough that restarting was not painful. The full conditions were different.
No CBR took more than 20 minutes, and non-parametric CBR also required a long sequential run. I therefore added checkpointing after every completed task.
The checkpoint stored the condition, number of completed tasks, results, and Case Bank. That meant a crash at task 30 no longer meant restarting from task 1.
I could reload:
30 completed
30 results
30 cases
and continue from task 31.
This ended up being necessary multiple times.
A strange benchmark example
One of the more interesting failures came from MuSiQue.
The question was:
Who is the spouse of the actor of Ethan in A Dog's Purpose?
My live-web research agent followed the chain:
adult Ethan
→ Dennis Quaid
→ current spouse
→ Laura Savoie
The benchmark answer was:
Meg Ryan
At first I thought the benchmark was simply stale. Dennis Quaid and Meg Ryan divorced years ago, while Laura Savoie is his current spouse.
But when I inspected the actual MuSiQue record, the situation was more interesting. The example contained supporting paragraphs and a predefined question decomposition:
who plays old Ethan in A Dog's Purpose
→ Dennis Quaid
Dennis Quaid → spouse
→ Meg Ryan
The benchmark was constructed to answer the question from its supplied evidence, while my agent was answering using the live web.
Those are not necessarily the same task:
benchmark-grounded QA
≠
current-world web research
I left the benchmark labels untouched, but this became a useful reminder that evaluation methodology can matter as much as agent capability.
Results
After debugging the executor and running both completed conditions, I got the following overall results.
| Metric | No CBR | Non-parametric CBR | Difference |
|---|---|---|---|
| Exact Match | 0.325 | 0.400 | +0.075 |
| F1 | 0.471 | 0.526 | +0.055 |
| Reward | 0.325 | 0.400 | +0.075 |
| Execution Success | 1.000 | 0.975 | -0.025 |
Another way to read the same result is as a comparison between what changed and what stayed fixed:
| Comparison | No CBR | Non-parametric CBR |
|---|---|---|
| Planner model | GPT-4.1 | GPT-4.1 |
| Executor model | o3 | o3 |
| Model weights | Frozen | Frozen |
| Task order | Same 40 tasks | Same 40 tasks |
| Tools | Search, read, Python | Search, read, Python |
| Memory retrieval | None | Similar previous cases retrieved |
| Main change | Solves each task without past experience | Planner sees relevant past cases before planning |
Adding similarity-based Case Memory improved Exact Match from 32.5% to 40.0%, an improvement of 7.5 percentage points.
F1 increased from 47.1% to 52.6%, or approximately 5.5 percentage points.
The underlying planner and executor models were unchanged.
Results by dataset
The gains were not evenly distributed.
| Dataset | No CBR EM | CBR EM | EM Difference | No CBR F1 | CBR F1 | F1 Difference |
|---|---|---|---|---|---|---|
| 2WikiMultihopQA | 0.10 | 0.10 | +0.00 | 0.168 | 0.221 | +0.053 |
| HotpotQA | 0.30 | 0.40 | +0.10 | 0.480 | 0.558 | +0.078 |
| Natural Questions | 0.10 | 0.20 | +0.10 | 0.266 | 0.390 | +0.124 |
| TriviaQA | 0.80 | 0.90 | +0.10 | 0.971 | 0.936 | -0.035 |
2WikiMultihopQA improved on F1 but not Exact Match. HotpotQA and Natural Questions both improved on Exact Match and F1. TriviaQA already performed strongly without memory, so the Exact Match improvement came with a small F1 drop.
With only 10 examples from each dataset, I would not read too much into any individual dataset-level movement. The overall direction is more interesting than any single subset.
What does this actually show?
The result does not prove that Case-Based Reasoning universally improves agents. Forty tasks are nowhere near enough for that claim.
What it does show is that the mechanism worked in a real agent implementation. The same frozen models, operating over the same task stream, performed better when the planner could access similar previous experiences.
That matters because the additional capability did not come from changing model weights. It came from changing context, specifically by making structured previous experience available to the agent.
Memory is not the same as a longer prompt
You could describe retrieval as simply adding more tokens to a prompt, and technically that is part of what happens. But the interesting part is the selection mechanism.
A memory system is not useful merely because information exists somewhere. The agent still has to answer a harder question:
Which past experience matters now?
That makes retrieval itself part of the learning problem. If every past task is dumped into context, memory quickly becomes noise. If the wrong cases are retrieved, previous experience can actively hurt performance.
So the problem starts to move away from only asking:
How do we update the model?
and towards asking:
What should the agent remember?
and:
When should it remember it?
Those are different problems, and they may require different system designs.
Where I stopped
Memento also explores a parametric version of retrieval. Instead of retrieving cases only through embedding similarity, the idea is to learn which cases are useful.
I started implementing this. My first attempt immediately demonstrated why the problem is harder than it appears.
I trained a small network using case reward as the target. The loss collapsed from:
0.6700
to:
0.0015
which initially looked excellent.
Then I tested retrieval. For a question about when Somewhere Over the Rainbow came out, the network retrieved cases about the Communist Manifesto, a Scottish football club, a Kentucky county, and a Gilbert and Sullivan operetta.
Those cases happened to have reward 1.
The network had learned:
successful case = useful case
instead of:
useful for this query = useful case
That is a very different objective.
I changed the training pairs to use actual query-case interactions from the non-parametric run, but with only 40 tasks the supervision remained weak. Rather than report an unreliable parametric result, I stopped there.
For this replication, the completed comparison is therefore:
No CBR
vs
Non-parametric CBR
not the full three-condition Memento evaluation.
What I did not test
There are several things this experiment does not establish.
I did not complete the parametric CBR condition, out-of-distribution evaluation, retrieval-depth ablation, the full task scale used in the paper, or statistical significance testing.
The original plan also included separate OOD evaluation using MuSiQue, Bamboogle, and PopQA, as well as testing multiple values of (K), the number of retrieved cases.
Those would require additional agent runs and significantly more API usage. Given the pilot nature of the experiment, I chose to stop after establishing the No-CBR versus non-parametric comparison.
What I would test next
There are several directions I would want to explore next.
The first is a better learned retriever. Instead of treating final task reward as if it tells us the utility of every retrieved case, I would want a stronger case-level learning signal. That could come from retrieval interventions, pairwise case preferences, counterfactual runs, task-level improvement attributed to specific cases, or online feedback collected over a much larger experience stream.
The second is scale. With only 40 tasks, memory is still extremely sparse. The more interesting question is what happens after 100 tasks, 1,000 tasks, or 10,000 tasks. At some point, storing more experience may stop helping. Retrieval may become noisy, old experiences may become actively harmful, and the system may need forgetting, consolidation, or memory compression.
I would also separate two different evaluation environments. One would be closed benchmark evidence, where the agent must reason only over supplied passages. The other would be live web research, where current-world answers are expected. Mixing those two can produce misleading evaluations, as the MuSiQue example demonstrated.
Another direction I want to test is whether these results depend heavily on using OpenAI models. This experiment used GPT-4.1 as the planner and o3 as the executor, but the architecture itself does not require proprietary models. I would like to repeat the same experiment with open-source or open-weight models such as Qwen and DeepSeek, keeping the memory and tool system as similar as possible. That would make it possible to ask whether memory-based agent improvement survives when the underlying planner and executor are cheaper, locally hosted, or less capable.
That comparison is particularly interesting because memory may matter differently depending on the strength of the base model. A stronger model may already know how to solve many tasks with little help from previous cases, while a smaller model could potentially benefit much more from structured experience. Conversely, weaker models may also be worse at correctly interpreting or reusing retrieved cases. I would want to measure both possibilities.
I also would not want to assume that Case-Based Reasoning is the best memory architecture simply because it worked here.
CBR stores fairly complete task-solving experiences:
task
plan
trajectory
answer
reward
That is only one way to represent experience.
I would like to compare it with other approaches, such as storing only successful plans, storing tool-use patterns separately from task content, summarising trajectories into reusable lessons, extracting factual or procedural memories, maintaining semantic and episodic memory separately, or allowing the agent to continually rewrite and consolidate its own memory.
That leads to a broader experiment:
No memory
vs
Case-Based Reasoning
vs
episodic memory
vs
semantic memory
vs
procedural memory
vs
compressed lessons
The interesting question would no longer just be whether memory helps, but what form of memory helps which kinds of tasks.
There is also a more fundamental question behind the whole case representation itself. A complete previous case may be too large or too specific. Perhaps the most valuable thing to retrieve from an old task is not the full task and trajectory, but a small reusable abstraction such as:
When a question asks for a relationship between two entities, resolve the first entity explicitly before researching the second relationship.
That kind of memory is closer to a learned strategy than a remembered example.
So another direction would be to compare storing complete cases against storing distilled rules, strategies, or skills derived from those cases.
At that point, the experiment starts moving beyond Memento specifically and towards a larger question:
What should an agent actually remember from experience?
A different way to think about agent learning
What I like most about this experiment is that it changes where I look for improvement.
When an agent fails repeatedly, the immediate instinct does not always have to be:
I need a better model.
Maybe the model already has enough capability. Maybe the system is failing to accumulate experience.
A model can solve a task today and encounter an almost identical task tomorrow as if it had never seen anything before. That is strange when compared with how we think about intelligent systems.
An agent operating repeatedly in an environment should probably have some mechanism for becoming different because of what has happened to it. Fine-tuning is one way to achieve that. External memory is another.
External memory also has a useful property: the underlying model does not have to change.
Conclusion
In this pilot, I kept the planner and executor frozen and changed only whether the agent could retrieve previous task-solving experiences.
The baseline achieved 32.5% Exact Match and 47.1% F1.
With non-parametric Case-Based Reasoning, performance increased to 40.0% Exact Match and 52.6% F1.
That does not prove Memento's full claims, and this was not a full reproduction of the paper. It does, however, demonstrate the central idea at a small scale: an agent can improve without updating the weights of the model inside it.
Sometimes learning does not have to mean changing what the model knows.
It can mean changing what the agent remembers.
And the next question is probably not only whether an agent should have memory, but what kind of memory it should have at all.
Links
Paper: Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Replication notebook: Google Colab

