The problem
An agent picks its next tool by writing a function call. The model spells out the tool name one token at a time, the framework parses the text, and the call runs. This is how most agent frameworks work today.
This has three costs. Decoding a tool name takes about 10 forward passes on top of the prompt pass (6 to 14 for most ToolBench tools with this tokenizer), so each decision is slow. The model can spell a name that does not match any tool, and the call fails. And you get one name, not a ranking over the offered tools. Getting alternatives means more decoding, and the alternatives are mostly respellings of the first guess.
None of this is necessary. The agent's toolbox is a fixed list. The model does not need to write a name. It needs to pick one.
The idea
Netflix's recommendation team solved the same problem for movie titles. In GenRec, they argue that when your choices come from a catalog, you should not generate at all. Run the LLM once over the context, take a summary vector from the hidden states, and score every catalog item against it. The model can only pick items that exist, so it cannot hallucinate a title. There is no decoding step, so serving is one forward pass.
I wanted to know if this works when the catalog is an agent's toolbox. I had not seen it tested there, and tools are a harder catalog than movies. The set changes from task to task. A quarter of the tools the agent has to pick at test time were never the correct answer in training. And the right choice depends on a long history of calls and results, not a viewing record. The question was whether a scorer could match a generating agent's accuracy under those conditions, and what it takes to train one.
Almost. With GenRec's full training objective, three training seeds per system, and 1,352 held-out steps that no training run ever looked at, scoring the list lands about 2.6 points behind generating the name on first-guess accuracy, close enough that no seed pairing is decisive either way. On Mean Reciprocal Rank, the metric GenRec itself reports, the two are level. Scoring wins everything else.
| system | top-1 | tool picks | MRR | top-5 | halluc. | latency |
|---|---|---|---|---|---|---|
| generation | 66.6% | 62.7% | 0.764 | 90.7% | 0.9% | 632 ms |
| ActionRank, ranking loss only | 62.5% | 56.6% | 0.768 | 97.2% | 0.0% | 251 ms |
| ActionRank, GenRec's full objective | 64.0% | 58.7% | 0.778 | 97.5% | 0.0% | 247 ms |
| difference, last row vs generation | -2.6 | -4.0 | +0.014 | +6.8 | -0.9 | 2.6x |
Best value per column in bold; the difference row is the full-objective scorer relative to generation, in points. Mean of three training seeds per system, evaluated on 1,352 held-out steps outside every system's training-time monitoring. "Tool picks" is top-1 excluding steps where the right answer is to stop. MRR is Mean Reciprocal Rank over each system's five best guesses, the metric GenRec reports; it was computed afterwards from the saved predictions. All three use the same LoRA configuration and three backbone epochs on the same data. Latency is from a laptop GPU at batch size 1. Every system selects a tool name only; none generates arguments.
A note on scale. Everything here uses a 1.5B-parameter backbone, chosen because of compute. The absolute accuracies belong to that scale, and to a dataset whose labels cap top-1 well below 100% for any system. A larger backbone would likely lift both systems. Whether it narrows the gap is untested. GenRec reports that larger backbones rank better, and the scorer's weakest slice, tools it never trained on, depends on how well the model reads a description. The structural results do not depend on size: a scorer cannot name a tool that was not offered, and it never pays for decoding.
The rest of this post is how I got there. The first version of this experiment said the scorer lost badly. A pooling fix made it a tie. A clean evaluation and three seeds turned the tie into a 4-point loss. Then re-reading GenRec's recipe turned up a loss term I had skipped, and adding it recovered a third of that. Here is how the two approaches differ, then what I built, then what happened.
Generation (today's default)
ActionRank (scoring)
What I built
ActionRank takes the same prompt a generating agent would see: the task, the tool calls so far with their results, and the candidate tools with one-line descriptions. It runs Qwen2.5-1.5B-Instruct over that prompt once and scores each candidate from the hidden states.
GenRec scores items with a learned embedding table, one row per catalog item. I built that first. It works for tools that appear as training labels and collapses on tools that do not, because an embedding row that never received a positive example carries little information. That matters here: for 24% of the evaluation steps, the correct tool was never a training label. It may have appeared as a candidate, but the head never got a positive example for it.
So I replaced the table with what I call a span head. For each candidate, it pools the hidden states over that tool's description line inside the prompt and scores the pooled vector against the prompt summary. A tool the head never trained on still gets a real representation, because its description is right there in the context. It beat the table head everywhere I tried it, so the final comparison below uses it.
I trained it in two stages. First, freeze the backbone and train only the small scoring head, which takes about a minute on cached vectors. Second, add a LoRA adapter to the backbone and train adapter and head together, starting from the head the first stage produced.
One scope note. Both systems choose a tool name. Neither one generates the arguments for the call. A real agent still has to do that after the tool is chosen, so the latency numbers below are for the selection step, not the whole call.
How I tested it
- Data. ToolBench G1, the single-tool tasks. Each task lists between 2 and 11 candidate tools, about 6 on average, and comes with a reference trajectory from a stronger agent. I sliced every trajectory into decision steps and held out 15% of tasks for evaluation. Deciding to stop is a tool too, called
Finish. - Baseline. The same backbone, prompted to write the tool name. For the fair comparison, I fine-tuned it with the same LoRA configuration, the same data, and the same number of backbone epochs as the scorer. It starts from the base model, where the scorer's head starts from the frozen-backbone stage.
- Metrics. Top-1 is how often the single best guess is right. Top-5 is how often the right tool is in the best five, which measures the ranking. For the generator, that list is the greedy answer plus four beam-search alternatives. MRR, Mean Reciprocal Rank, averages 1 divided by the position of the right tool in that list, so first place scores 1, second 0.5, third 0.33, and a miss 0. It is the metric GenRec itself reports. Hallucination rate is how often the system names a tool that is not on the list. Latency is wall-clock time per decision.
- Evaluation. Two sets. A 500-step development set, the first 500 steps of the held-out split, which every intermediate experiment below uses. During training I watched accuracy on a prefix of it (200 steps for the scorer, 100 for the generator), so it is not a clean test set. And the final comparison: 1,352 held-out steps from 371 trajectories that no training run ever logged, evaluated once per seed with the analysis decided in advance.
What happened
The first comparison misled me
With the backbone frozen, the scorer beat the prompted generator by 26 points. That looked like a clean win. It was not. The prompted generator never outputs Finish, so it fails every step where stopping was correct. The scorer's head had trained on ToolBench and the generator had trained on nothing.
So I fine-tuned the generator. One epoch of LoRA took 12 minutes on an A100 and lifted it from 31.8% to 66.8%. Now it led the frozen scorer by 9 points overall and by 16 on tools that were never a training label. The honest summary at this point was: scoring buys safety and speed, and you pay for it in accuracy.
The scorer would not train
The obvious fix was to fine-tune the scorer's backbone too. I added the same adapter to the table-head scorer and trained for two epochs, checking accuracy on the first 200 evaluation steps after each one. It went 48.0%, 47.5%, 48.0%. The loss fell a little. Accuracy did not.
I checked the run. The adapter was changing the pooled vector by about 10%, and it was receiving gradients. It was learning. The metric just did not respond.
My best explanation is one detail. GenRec takes the hidden state at a single pooling position. I had been averaging over all 400 prompt tokens. To move that average, a small adapter has to shift hundreds of token states in the same direction at once. The generator has no such problem, because it reads one sharp distribution at the last token. I have not proven this is the mechanism, but it predicted the fix: I switched the scorer to last-token pooling and ran the identical recipe again.
With last-token pooling, the scorer trains. For a frozen backbone the pooling choice barely matters, which is why I missed it. It only matters once you fine-tune through it. From here on the scorer uses last-token pooling and the span head.
The matched comparison, on the development set
Now both systems get the same LoRA configuration, the same data, and the same three backbone epochs. The scorer's head starts from the frozen-backbone stage; the generator starts from the base model.
- Top-1 ties, here. 66.6% against 66.0%. At 500 examples that gap is noise. This is the number I first published. Keep reading.
- Top-5 does not. When the scorer misses, the right tool sits in its top five 98% of the time. The generator's beam-search alternatives are mostly respellings of its first guess, so its top-5 stops at 93%.
- The generator still hallucinates. Three epochs cut its rate from 1.4% to 0.6%, but it still names a tool that does not exist three times in 500 decisions. The scorer cannot.
- The scorer runs in less than half the time. One prefill, no decoding. 250 ms against 575 ms.
More epochs do not change this. At six epochs the scorer overfits and the generator plateaus.
The tie did not survive a clean evaluation
A code review pointed out that both training scripts print accuracy on a prefix of the held-out split after every epoch, and those steps sit inside my 500. I never picked a checkpoint on them, but I could see them while making decisions. That makes the 500 a development set. So I wrote down, in advance, a plan to evaluate the two frozen final models on the 1,352 held-out steps past that prefix, with the analysis fixed before running: a bootstrap over trajectories, not steps, and a 3-point margin for calling anything a tie.
The generator got 68.3%. The scorer got 63.5%. The tie was gone. On the 500, the scorer had stopped correctly 88% of the time against the generator's 78%. On the 1,352 they were level on stopping, and on the steps where an actual tool was the answer the generator led by 6 points.
One run each cannot separate a 5-point effect from training luck, so I trained two more of each system with different random seeds, same recipe, and evaluated them on the same 1,352 steps.
The scorer's three seeds land within 1.7 points of each other. The generator's third seed is 5 points below its siblings, and the breakdown says why: it learned to stop less reliably. Its Finish recall is 62% where the other two are 78% and 91%, while its accuracy on actual tool picks matches them. The generator's variance is almost entirely about when to stop. Across all nine scorer-versus-generator pairings, six favor the generator with confidence intervals clear of zero, and the three involving that third generator seed are a tie or inconclusive. Averaged over seeds, the generator leads by 4 points, and by 6 on tool picks.
What held on every seed, without exception: the scorer never picked a tool that was not offered, its top-5 stayed at 97% against the generator's 90 to 92%, and the latency ratio did not move.
Giving the scorer the rest of GenRec's recipe
Before writing that up, I went back through the GenRec post line by line to make sure the scorer I had built was the one they described. It was not, quite. GenRec trains its scorer with two losses at once: the ranking loss, and a plain language-modeling loss over the verbalized text and the answer. My scorer had only the ranking loss. My generator baseline, on the other hand, was trained with exactly that language-modeling loss. So the matched comparison above was a scorer trained with half of GenRec's recipe against a generator that got the half the scorer was missing.
The fix is one loss term. Same scorer, same data, same three epochs, same inference. During training, the model also predicts the prompt and the answer token by token, and that loss is added to the ranking loss. The ranking score is still read from the last prompt token, before the answer, so the answer cannot leak into it. Three seeds, evaluated once each on the same 1,352 steps.
The scorer gains about a point and a half overall, and the gain sits exactly where the theory says it should. On tools that were never a training label, two of three seeds jumped 7 points, from 44% to 51%, which is the reading-the-description problem the language-modeling loss targets. On tools the model trained on, the gain is under a point. The gap to the generator narrows from 4.1 points to 2.6, and under the pre-registered rule none of the nine seed pairings is decisive any more: the scorer no longer loses any pairing outright, and does not win any either.
One more number, because it is the one GenRec uses. GenRec never reports first-guess accuracy. Its offline metric is Mean Reciprocal Rank, which scores the whole ranking. On MRR the full-objective scorer gets 0.778 and the generator 0.764. Compared seed by seed with the same trajectory-level intervals, no pairing favors the generator: six of nine are a dead heat, within 0.006, and the three against the generator's weakest seed favor the scorer. The reason is visible in the predictions. The generator's first guess is right more often, but when it is wrong its alternatives are mostly respellings of that guess. The scorer's second and third choices are real candidates, and MRR credits that.
Everything structural is unchanged: zero hallucinated tools on every seed, top-5 between 97 and 98%, and 247 ms per decision on the laptop against the generator's 632.
Where the accuracy gap lives
With the ranking loss alone, the generator led by about 4 points on tools it trained on and by about 10 on tools that were never a training label. GenRec's language-modeling loss closed most of the second gap, to about 5 points, and left the first alone. What remains is a deficit on familiar tools that the loss does not touch. GenRec's Phase 1 domain adaptation, more training data, or a larger backbone are the levers left, and all three are things GenRec itself reports gains from.
What this means
Scoring is a trade, not a free win. GenRec's structural promises transferred exactly: the scorer cannot pick a tool that was not offered, it ranks every candidate from one forward pass, and it decides in less than half the time. What did not fully transfer is accuracy parity. With GenRec's full objective and three seeds, generating the name is still about 2.6 points more accurate, and about 4 points more accurate on the steps where a tool is actually chosen, though no seed pairing is decisive, and on GenRec's own ranking metric the two are level. Whether that trade is worth it depends on the agent. One that retries from the ranking, or that cannot afford an invalid call, gets a lot for 3 to 4 points. One that executes its first choice and can validate names cheaply gets less.
The scorer is also the more predictable system. Three seeds within 2 points, against a 5-point spread for the generator driven by how well each run learns to stop. That matters if you are going to train one and ship it.
Three things I would not have guessed going in. First, the untrained comparison is worthless. Fine-tune both sides before you compare, or the result mostly measures whether the generator knows when to stop. Second, pooling position is not a minor choice. With mean pooling the scorer did not improve through a LoRA adapter. Last-token pooling fixed it, with nothing else changed. Third, a result on a development set is a hypothesis. The tie was real on those 500 steps and gone on the next 1,352. Evaluate on data nobody looked at, then run more seeds, before you believe a number. And re-read the paper you are replicating after you have results, not just before. That is when you know which details matter.
Where this fits in practice
The scorer needs training data. Specifically, it needs traces: for each decision an agent faced, the context it had and the tool it should have called. I got mine from ToolBench. Most teams already have their own.
If you run agents in production with tracing turned on, you are already logging exactly this. Every run records the task, the tool calls, their results, and whether the run succeeded. Filter to the runs that ended well, slice them into decision steps, and you have a training set in the same shape I used here. Nothing about the recipe is specific to ToolBench.
That makes this a way to turn observability data into a better agent. The frozen-backbone version trains in about a minute on cached vectors, so you can retrain as traces accumulate. The scorer learns which tool your agent needs in each situation, it cannot call a tool you did not give it, and it makes each decision faster than generating the call. It gives up a few points of top-1 against a fine-tuned generator, so it fits best where an invalid call is expensive or where you use the ranking. It should keep improving as traces accumulate, and the traces are the same ones you were already keeping.
One caution. The scorer learns to imitate the traces you give it, so the filter matters. Train on runs that succeeded, or weight by outcome the way GenRec weights by engagement, and it learns good tool use. Train on everything, and it learns your agent's current habits, mistakes included.
Limitations and future expansions
- Each task lists about 6 candidates, so this is shortlist ranking, not retrieval. Random guessing already gets about 83% top-5. The real test for scoring is ranking all 6,372 tools in the catalog at once. That is the next experiment.
- The headline has three seeds per system on 1,352 steps. Everything else in the post is one run on the 500-step development set. One dataset and one backbone size throughout.
- This replicates GenRec's Phase 2. It does not do Phase 1, domain adaptation of the backbone before ranking, which GenRec credits with a 10 to 20% relative gain, nor reward weighting, which ToolBench has no signal for. Phase 1, more training data, and a larger backbone are the natural next steps, and GenRec reports gains from all three.
- The 1,352 steps were never monitored during training, but earlier frozen-head experiments did report aggregate numbers over the full held-out split, and those informed the choice of head and pooling. A truly independent test needs a predeclared train, dev, and test split with retraining. That is the next experiment.
- This measures tool-name selection only. Neither system generates arguments, so the latency numbers cover the selection step, not a complete tool call.
- The generator's top-5 comes from beam search, which is not the strongest ranking a generator can produce. A constrained decoder that can only emit candidate names would be a fairer zero-hallucination baseline, and it would isolate ranking quality and latency.
- A larger backbone might narrow the gap on tools that were never a training label, if that gap comes from how well the model reads descriptions. That is a hypothesis.
- Latencies come from Apple Silicon at batch size 1. I expect the ratio to hold elsewhere, but I have not measured it. The absolute numbers will not transfer.
The code, data pipeline, training scripts, configs, pre-registration documents, and per-run results tables are in the repo below. Checkpoints and per-step prediction files are not tracked; the evaluation script regenerates them.