Technical Report

ActionRank

Does Netflix's "score, don't generate" idea work for agent tool selection?

Yash Thapliyal · September 2026 · 14 min read

The problem

An agent picks its next tool by writing a function call. The model spells out the tool name one token at a time, the framework parses the text, and the call runs. This is how most agent frameworks work today.

This has three costs. Decoding a tool name takes about 10 forward passes on top of the prompt pass (6 to 14 for most ToolBench tools with this tokenizer), so each decision is slow. The model can spell a name that does not match any tool, and the call fails. And you get one name, not a ranking over the offered tools. Getting alternatives means more decoding, and the alternatives are mostly respellings of the first guess.

None of this is necessary. The agent's toolbox is a fixed list. The model does not need to write a name. It needs to pick one.

The idea

Netflix's recommendation team solved the same problem for movie titles. In GenRec, they argue that when your choices come from a catalog, you should not generate at all. Run the LLM once over the context, take a summary vector from the hidden states, and score every catalog item against it. The model can only pick items that exist, so it cannot hallucinate a title. There is no decoding step, so serving is one forward pass.

I wanted to know if this works when the catalog is an agent's toolbox. I had not seen it tested there, and tools are a harder catalog than movies. The set changes from task to task. A quarter of the tools the agent has to pick at test time were never the correct answer in training. And the right choice depends on a long history of calls and results, not a viewing record. The question was whether a scorer could match a generating agent's accuracy under those conditions, and what it takes to train one.

Almost. With GenRec's full training objective, three training seeds per system, and 1,352 held-out steps that no training run ever looked at, scoring the list lands about 2.6 points behind generating the name on first-guess accuracy, close enough that no seed pairing is decisive either way. On Mean Reciprocal Rank, the metric GenRec itself reports, the two are level. Scoring wins everything else.

systemtop-1tool picksMRRtop-5halluc.latency
generation66.6%62.7%0.76490.7%0.9%632 ms
ActionRank, ranking loss only62.5%56.6%0.76897.2%0.0%251 ms
ActionRank, GenRec's full objective64.0%58.7%0.77897.5%0.0%247 ms
difference, last row vs generation-2.6-4.0+0.014+6.8-0.92.6x

Best value per column in bold; the difference row is the full-objective scorer relative to generation, in points. Mean of three training seeds per system, evaluated on 1,352 held-out steps outside every system's training-time monitoring. "Tool picks" is top-1 excluding steps where the right answer is to stop. MRR is Mean Reciprocal Rank over each system's five best guesses, the metric GenRec reports; it was computed afterwards from the saved predictions. All three use the same LoRA configuration and three backbone epochs on the same data. Latency is from a laptop GPU at batch size 1. Every system selects a tool name only; none generates arguments.

A note on scale. Everything here uses a 1.5B-parameter backbone, chosen because of compute. The absolute accuracies belong to that scale, and to a dataset whose labels cap top-1 well below 100% for any system. A larger backbone would likely lift both systems. Whether it narrows the gap is untested. GenRec reports that larger backbones rank better, and the scorer's weakest slice, tools it never trained on, depends on how well the model reads a description. The structural results do not depend on size: a scorer cannot name a tool that was not offered, and it never pays for decoding.

The rest of this post is how I got there. The first version of this experiment said the scorer lost badly. A pooling fix made it a tie. A clean evaluation and three seeds turned the tie into a 4-point loss. Then re-reading GenRec's recipe turned up a loss term I had skipped, and adding it recovered a third of that. Here is how the two approaches differ, then what I built, then what happened.

Generation (today's default)

1. Run the model over the prompt.
2. Decode the tool name, one token per forward pass.
3. Parse the text and hope it names a real tool.
output: one name, no scores

ActionRank (scoring)

1. Run the model over the prompt, once.
2. Pool a vector for each candidate tool from its own description line in the prompt.
3. Score every candidate against the prompt summary. Softmax. Done.
output: a score for every candidate

What I built

ActionRank takes the same prompt a generating agent would see: the task, the tool calls so far with their results, and the candidate tools with one-line descriptions. It runs Qwen2.5-1.5B-Instruct over that prompt once and scores each candidate from the hidden states.

GenRec scores items with a learned embedding table, one row per catalog item. I built that first. It works for tools that appear as training labels and collapses on tools that do not, because an embedding row that never received a positive example carries little information. That matters here: for 24% of the evaluation steps, the correct tool was never a training label. It may have appeared as a candidate, but the head never got a positive example for it.

So I replaced the table with what I call a span head. For each candidate, it pools the hidden states over that tool's description line inside the prompt and scores the pooled vector against the prompt summary. A tool the head never trained on still gets a real representation, because its description is right there in the context. It beat the table head everywhere I tried it, so the final comparison below uses it.

I trained it in two stages. First, freeze the backbone and train only the small scoring head, which takes about a minute on cached vectors. Second, add a LoRA adapter to the backbone and train adapter and head together, starting from the head the first stage produced.

One scope note. Both systems choose a tool name. Neither one generates the arguments for the call. A real agent still has to do that after the tool is chosen, so the latency numbers below are for the selection step, not the whole call.

How I tested it

What happened

The first comparison misled me

Top-1 accuracy before matched fine-tuning
500 evaluation steps. The scorer's head is trained; the prompted generator is zero-shot.
generationActionRank
random pick among candidates: 20.8%

With the backbone frozen, the scorer beat the prompted generator by 26 points. That looked like a clean win. It was not. The prompted generator never outputs Finish, so it fails every step where stopping was correct. The scorer's head had trained on ToolBench and the generator had trained on nothing.

So I fine-tuned the generator. One epoch of LoRA took 12 minutes on an A100 and lifted it from 31.8% to 66.8%. Now it led the frozen scorer by 9 points overall and by 16 on tools that were never a training label. The honest summary at this point was: scoring buys safety and speed, and you pay for it in accuracy.

The scorer would not train

The obvious fix was to fine-tune the scorer's backbone too. I added the same adapter to the table-head scorer and trained for two epochs, checking accuracy on the first 200 evaluation steps after each one. It went 48.0%, 47.5%, 48.0%. The loss fell a little. Accuracy did not.

I checked the run. The adapter was changing the pooled vector by about 10%, and it was receiving gradients. It was learning. The metric just did not respond.

My best explanation is one detail. GenRec takes the hidden state at a single pooling position. I had been averaging over all 400 prompt tokens. To move that average, a small adapter has to shift hundreds of token states in the same direction at once. The generator has no such problem, because it reads one sharp distribution at the last token. I have not proven this is the mechanism, but it predicted the fix: I switched the scorer to last-token pooling and ran the identical recipe again.

Same adapter, same data, same head. Only the pooling position changed.
LoRA plus table head, top-1 on the first 200 evaluation steps by epoch. Epoch 0 is before adapter training.
mean poolinglast-token pooling

With last-token pooling, the scorer trains. For a frozen backbone the pooling choice barely matters, which is why I missed it. It only matters once you fine-tune through it. From here on the scorer uses last-token pooling and the span head.

The matched comparison, on the development set

Now both systems get the same LoRA configuration, the same data, and the same three backbone epochs. The scorer's head starts from the frozen-backbone stage; the generator starts from the base model.

Both systems fine-tuned for 3 epochs
500-step development set, batch size 1, same laptop GPU.
generationActionRank

More epochs do not change this. At six epochs the scorer overfits and the generator plateaus.

The tie did not survive a clean evaluation

A code review pointed out that both training scripts print accuracy on a prefix of the held-out split after every epoch, and those steps sit inside my 500. I never picked a checkpoint on them, but I could see them while making decisions. That makes the 500 a development set. So I wrote down, in advance, a plan to evaluate the two frozen final models on the 1,352 held-out steps past that prefix, with the analysis fixed before running: a bootstrap over trajectories, not steps, and a 3-point margin for calling anything a tie.

The generator got 68.3%. The scorer got 63.5%. The tie was gone. On the 500, the scorer had stopped correctly 88% of the time against the generator's 78%. On the 1,352 they were level on stopping, and on the steps where an actual tool was the answer the generator led by 6 points.

One run each cannot separate a 5-point effect from training luck, so I trained two more of each system with different random seeds, same recipe, and evaluated them on the same 1,352 steps.

Top-1 by training seed, 1,352 unmonitored held-out steps
Same recipe per system, different shuffle order and adapter initialization. Seed 1 is the original run.
generationActionRank

The scorer's three seeds land within 1.7 points of each other. The generator's third seed is 5 points below its siblings, and the breakdown says why: it learned to stop less reliably. Its Finish recall is 62% where the other two are 78% and 91%, while its accuracy on actual tool picks matches them. The generator's variance is almost entirely about when to stop. Across all nine scorer-versus-generator pairings, six favor the generator with confidence intervals clear of zero, and the three involving that third generator seed are a tie or inconclusive. Averaged over seeds, the generator leads by 4 points, and by 6 on tool picks.

What held on every seed, without exception: the scorer never picked a tool that was not offered, its top-5 stayed at 97% against the generator's 90 to 92%, and the latency ratio did not move.

Giving the scorer the rest of GenRec's recipe

Before writing that up, I went back through the GenRec post line by line to make sure the scorer I had built was the one they described. It was not, quite. GenRec trains its scorer with two losses at once: the ranking loss, and a plain language-modeling loss over the verbalized text and the answer. My scorer had only the ranking loss. My generator baseline, on the other hand, was trained with exactly that language-modeling loss. So the matched comparison above was a scorer trained with half of GenRec's recipe against a generator that got the half the scorer was missing.

The fix is one loss term. Same scorer, same data, same three epochs, same inference. During training, the model also predicts the prompt and the answer token by token, and that loss is added to the ranking loss. The ranking score is still read from the last prompt token, before the answer, so the answer cannot leak into it. Three seeds, evaluated once each on the same 1,352 steps.

Top-1 by training seed and objective, 1,352 unmonitored held-out steps
Adding GenRec's language-modeling loss to the scorer. Hover for each seed's accuracy on tools that were never a training label.
generationActionRank

The scorer gains about a point and a half overall, and the gain sits exactly where the theory says it should. On tools that were never a training label, two of three seeds jumped 7 points, from 44% to 51%, which is the reading-the-description problem the language-modeling loss targets. On tools the model trained on, the gain is under a point. The gap to the generator narrows from 4.1 points to 2.6, and under the pre-registered rule none of the nine seed pairings is decisive any more: the scorer no longer loses any pairing outright, and does not win any either.

One more number, because it is the one GenRec uses. GenRec never reports first-guess accuracy. Its offline metric is Mean Reciprocal Rank, which scores the whole ranking. On MRR the full-objective scorer gets 0.778 and the generator 0.764. Compared seed by seed with the same trajectory-level intervals, no pairing favors the generator: six of nine are a dead heat, within 0.006, and the three against the generator's weakest seed favor the scorer. The reason is visible in the predictions. The generator's first guess is right more often, but when it is wrong its alternatives are mostly respellings of that guess. The scorer's second and third choices are real candidates, and MRR credits that.

Everything structural is unchanged: zero hallucinated tools on every seed, top-5 between 97 and 98%, and 247 ms per decision on the laptop against the generator's 632.

Where the accuracy gap lives

Top-1 by whether the correct tool appeared in training
1,352 unmonitored steps, tool picks only, mean of three seeds, ranking-loss-only scorer. 651 steps need a tool that was a training label at least once; 329 need one that never was.
generationActionRank

With the ranking loss alone, the generator led by about 4 points on tools it trained on and by about 10 on tools that were never a training label. GenRec's language-modeling loss closed most of the second gap, to about 5 points, and left the first alone. What remains is a deficit on familiar tools that the loss does not touch. GenRec's Phase 1 domain adaptation, more training data, or a larger backbone are the levers left, and all three are things GenRec itself reports gains from.

What this means

Scoring is a trade, not a free win. GenRec's structural promises transferred exactly: the scorer cannot pick a tool that was not offered, it ranks every candidate from one forward pass, and it decides in less than half the time. What did not fully transfer is accuracy parity. With GenRec's full objective and three seeds, generating the name is still about 2.6 points more accurate, and about 4 points more accurate on the steps where a tool is actually chosen, though no seed pairing is decisive, and on GenRec's own ranking metric the two are level. Whether that trade is worth it depends on the agent. One that retries from the ranking, or that cannot afford an invalid call, gets a lot for 3 to 4 points. One that executes its first choice and can validate names cheaply gets less.

The scorer is also the more predictable system. Three seeds within 2 points, against a 5-point spread for the generator driven by how well each run learns to stop. That matters if you are going to train one and ship it.

Three things I would not have guessed going in. First, the untrained comparison is worthless. Fine-tune both sides before you compare, or the result mostly measures whether the generator knows when to stop. Second, pooling position is not a minor choice. With mean pooling the scorer did not improve through a LoRA adapter. Last-token pooling fixed it, with nothing else changed. Third, a result on a development set is a hypothesis. The tie was real on those 500 steps and gone on the next 1,352. Evaluate on data nobody looked at, then run more seeds, before you believe a number. And re-read the paper you are replicating after you have results, not just before. That is when you know which details matter.

Where this fits in practice

The scorer needs training data. Specifically, it needs traces: for each decision an agent faced, the context it had and the tool it should have called. I got mine from ToolBench. Most teams already have their own.

If you run agents in production with tracing turned on, you are already logging exactly this. Every run records the task, the tool calls, their results, and whether the run succeeded. Filter to the runs that ended well, slice them into decision steps, and you have a training set in the same shape I used here. Nothing about the recipe is specific to ToolBench.

That makes this a way to turn observability data into a better agent. The frozen-backbone version trains in about a minute on cached vectors, so you can retrain as traces accumulate. The scorer learns which tool your agent needs in each situation, it cannot call a tool you did not give it, and it makes each decision faster than generating the call. It gives up a few points of top-1 against a fine-tuned generator, so it fits best where an invalid call is expensive or where you use the ranking. It should keep improving as traces accumulate, and the traces are the same ones you were already keeping.

One caution. The scorer learns to imitate the traces you give it, so the filter matters. Train on runs that succeeded, or weight by outcome the way GenRec weights by engagement, and it learns good tool use. Train on everything, and it learns your agent's current habits, mistakes included.

Limitations and future expansions

The code, data pipeline, training scripts, configs, pre-registration documents, and per-run results tables are in the repo below. Checkpoints and per-step prediction files are not tracked; the evaluation script regenerates them.

GitHub repo → GenRec post → ← Back to all posts