XLM-R classifiers and embeddings for intent routing
How to route requests in several languages, handle more than one action, and choose between trained XLM-R and embedding retrieval. Includes Python and E5 results.

"I have moved. Send my order to the new address, but do not cancel it."
The requested action is change_address. Cancellation is mentioned because the customer wants to avoid it. This is a useful test for a router: the vocabulary overlaps, while the requested actions differ.
For the setup described here, I would keep the existing XLM-R classifier trained on labelled user requests as the default for single-action requests with stable intents. It already has examples of how your users express those requests. Use embedding retrieval when the list of functions changes frequently, or when you are starting without labelled data. Pass uncertain cases to an LLM with a short candidate list and the relevant conversation state.
The choice I would make
| Your situation | Start here | Why |
|---|---|---|
| A trained classifier and a stable set of supported intents | Keep XLM-R as the main router | The labelled requests already describe the decisions you need |
| New functions arrive regularly | Retrieve function descriptions | Add a description and its vector without expanding a fixed classification head |
| A new product with few labelled requests | Embedding retrieval | Establish a usable baseline while collecting examples |
| Stable intents and many requests asking for several actions | Multi-label XLM-R trained on complete intent sets | Predict several supported labels together |
| Compound requests with few labelled combinations | Identify tasks, then retrieve candidates for each | Give every action its own shortlist and preserve the original context |
| A question needs current policy or product documentation | Document retrieval and generation | The answer needs source material beyond an intent label |
A useful classifier can return the leading intent directly for clear requests and five alternatives for ambiguous ones. The LLM can then inspect context and ask for missing details. For the opening sentence, even a correct change_address prediction leaves the new address to collect.
If your product has a dozen short function definitions, sending all of them to the LLM is also reasonable. Shortlisting becomes more useful as the catalogue grows. For 100 definitions at 150 tokens each, moving to five cuts that part of the prompt from 15,000 tokens to 750. The rest of the prompt and the answer still contribute to the bill.
The cost of teaching the labels
The classifier already exists in this comparison because that is the setup we are discussing. Building one from scratch means collecting requests, agreeing on intent definitions, resolving inconsistent labels and keeping examples current as the product changes. Fine-tuning is one part of that work.
Consider these four examples from a fictional shop:
| Request | Training label |
|---|---|
| "Cancel order 1234." | cancel_order |
| "Do not cancel order 1234. Change its address." | change_address |
| "The parcel arrived broken. I want my money back." | refund_order |
| "Where is order 1234?" | track_order |
The difficult examples are usually close to another supported intent. Include questions about an action, negation, reports that an action already happened, and local shorthand. A set made mostly from "cancel my order" with small wording changes teaches a much narrower task than the one customers will bring.
XLM-R supplies a bidirectional encoder pretrained across 100 languages. Intent training adds the mapping from requests to your labels. The appeal here is multilingual traffic using a shared encoder; the examples still need to cover the languages and wording your users actually use.
A fixed classification head returns scores for its trained label set in one encoder pass. Adding a genuinely new intent requires changing that head and training it for the new class. Removing an unavailable function from the candidate list is an ordinary application filter. An existing class can also need new examples if its meaning changes.
What the opening sentence retrieves
I ran intfloat/multilingual-e5-small against nine function descriptions for the fictional shop. For the exact opening request, it returned:
| Rank | Function | Cosine score |
|---|---|---|
| 1 | change_address |
0.867806 |
| 2 | cancel_order |
0.851355 |
| 3 | update_email |
0.838387 |
| 4 | download_invoice |
0.835597 |
| 5 | track_order |
0.829509 |
It got this one right. A nearby request exposed the problem:
"Do not cancel anything. Tell me whether order 5901 has shipped."
Here, cancel_order came first at 0.855029 and track_order came second at 0.834838. Passing the shortlist onward preserves the tracking option. An LLM still has to read the request correctly; choosing the leading vector match would select the wrong operation.
The evaluation file contains 18 supported requests, nine in English and nine in Spanish, plus two unsupported requests. On the supported set, E5 ranked the correct intent first in 16 cases. It appeared in the first three in 17 cases and the first five in all 18. Five places cover more than half of this small registry, so the two first-choice errors tell us more than the perfect recall at five.
The misses are useful to inspect. A request to return an already delivered item ranked cancellation above refund. In Spanish, "El pedido todavia no ha salido; no lo envies" placed cancellation fourth. Both correct intents survived shortlisting, which is the reason to keep several candidates for review.
Unsupported requests also found close matches. "I want to upgrade an existing order to express shipping" scored 0.871750 against change_address, higher than the opening request's correct match. That is a concrete reason to evaluate rejection separately from ranking.
These are measured outputs on a small, hand-written fixture. They illustrate the errors; they cannot establish production accuracy or a winner over your trained classifier, whose checkpoint was unavailable for this run. The E5 run used model revision 614241f622f53c4eeff9890bdc4f31cfecc418b3, with CPU float32 inference on an Apple M3 Pro.
Classifier and retrieval code
Multilingual E5 base starts from XLM-R base, then receives retrieval training. The smaller E5 model used below starts from multilingual MiniLM, as documented in the E5 technical report. The choice here concerns a head trained to distinguish your intents versus a model trained to compare separately encoded texts.
Both implementations can return the same structure: scores for the full registry, plus up to five eligible candidates. Keep the full scores for rejection and unavailable-intent checks. Use one ranking helper so ties behave consistently:
Here, intents.json is your own function catalogue: a JSON array of objects with id and description fields.
import json
from pathlib import Path
registry = json.loads(
Path("intents.json").read_text(encoding="utf-8")
)
def top5(scores, allowed_ids):
candidates = [
{"id": intent_id, "score": float(score)}
for intent_id, score in scores.items()
if intent_id in allowed_ids
]
return sorted(
candidates, key=lambda item: (-item["score"], item["id"])
)[:5]
The trained XLM-R checkpoint
Load the checkpoint that has already learned your supported labels. Hugging Face's classification guide documents the saved model and id2label mapping. The snippet assumes those labels are the registry IDs.
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_dir = "./trained-intent-model"
tokenizer = AutoTokenizer.from_pretrained(model_dir)
classifier = AutoModelForSequenceClassification.from_pretrained(
model_dir
).eval()
def classifier_candidates(text, allowed_ids):
inputs = tokenizer(
text, return_tensors="pt", truncation=True, max_length=128
)
with torch.inference_mode():
values = classifier(**inputs).logits[0].softmax(dim=-1)
scores = {
classifier.config.id2label[i]: float(value)
for i, value in enumerate(values)
}
return {"all_scores": scores, "candidates": top5(scores, allowed_ids)}
The 128-token limit is the short-request policy used by this example. Match it to your trained model's input policy and provide the relevant conversation context. If the head includes out_of_scope, keep that score alongside the supported labels.
Retrieval over function descriptions
Encode the descriptions once, then compare each incoming request with the stored vectors. This is the bi-encoder pattern in the Sentence Transformers semantic search guide.
from sentence_transformers import SentenceTransformer
encoder = SentenceTransformer("intfloat/multilingual-e5-small")
vectors = encoder.encode(
["passage: " + item["description"] for item in registry],
normalize_embeddings=True,
convert_to_numpy=True,
)
def embedding_candidates(text, allowed_ids):
query = encoder.encode(
"query: " + text,
normalize_embeddings=True,
convert_to_numpy=True,
)
similarities = vectors @ query
scores = {
item["id"]: float(similarities[i])
for i, item in enumerate(registry)
}
return {"all_scores": scores, "candidates": top5(scores, allowed_ids)}
The E5 model card specifies query: for requests and passage: for descriptions, including non-English inputs. Normalised vectors make the dot product equal cosine similarity. E5-small produces 384-dimensional vectors and truncates inputs at 512 tokens.
For 1,000 functions, a float32 description matrix takes 1,000 * 384 * 4, about 1.5 MB. Exact matrix search is a sensible starting point for a catalogue this size. You can add a search service later for scale, filtering or operational needs.
Give the retriever a description that explains the operation. "Change the delivery address of an existing order before shipment" is more useful than a function signature alone. Several examples per function can help, but aggregate matches into distinct IDs so five refund paraphrases do not occupy all five places.
You can also index the same labelled requests used to train the classifier and retrieve similar examples. E5 recommends query: on both sides for request-to-request paraphrase matching. Keep evaluation requests out of that index. This gives the retrieval baseline access to the same examples, which makes the comparison more useful than giving it one vague description per function.
Choose one scorer for the request
The diagram shows alternative implementations. Select one scorer in the application configuration. Run both on the same evaluation requests when comparing them; each production request can use the chosen implementation.
Send candidate IDs, descriptions and required arguments to the LLM, together with the context that explains references such as "my order" or "send it there". For the opening sentence, a useful next response is:
{
"status": "clarify",
"intent_id": null,
"clarifying_question": "What is the new delivery address for this order?"
}
The review stage can return choose, clarify, unsupported or blocked. Application code validates the selected ID, arguments, permissions and current order state before execution.
Filter eligible functions before taking five. Also retain stronger unavailable matches as separate context. If cancellation leads the ranking but the order has shipped, the review stage needs that information to explain the restriction. Promoting a weaker tracking match would answer a different request.
If you do need a hybrid, give its stages explicit jobs. A classifier could choose a broad category, then embeddings retrieve functions within that category. This helps when a large catalogue changes inside otherwise stable categories. Check category errors carefully, because an early wrong category can exclude the correct function. Combining both scorers adds work; I would introduce that extra stage for a demonstrated retrieval problem.
When one message asks for several things
"Cancel order 5901 and send me its invoice" asks for cancel_order and download_invoice. "Do not cancel order 5901; send me its invoice" asks for just the invoice. Your labels need to preserve that difference.
If these combinations are common and the supported intents are stable, I would train multi-label XLM-R. Label each request with every action it asks for, using a multi-hot target such as [1, 0, 0, 0, 1, ...]. Hugging Face's XLM-R documentation supports multi_label_classification: binary cross-entropy during training, then a sigmoid score for each label. The existing checkpoint can supply the starting weights. It needs further training on those complete label sets before the new scores are useful.
After training, use thresholds selected on held-out multi-intent requests. This function takes that trained model and its threshold map:
def multiple_intents(model, tokenizer, text, allowed_ids, thresholds):
if model.config.problem_type != "multi_label_classification":
raise ValueError("Load a checkpoint trained on multi-hot intent labels.")
inputs = tokenizer(
text, return_tensors="pt", truncation=True, max_length=128
)
model.eval()
with torch.inference_mode():
values = model(**inputs).logits[0].sigmoid()
scores = {
model.config.id2label[i]: float(value)
for i, value in enumerate(values)
}
return sorted(
intent_id for intent_id, score in scores.items()
if intent_id in allowed_ids and score >= thresholds[intent_id]
)
It can return zero, one or several IDs. Keep out_of_scope outside the execution allowlist. Changing a saved model's configuration flag and swapping in sigmoid is not a substitute for training.
For a changing catalogue or a shortage of labelled combinations, I would identify the requested tasks first, then retrieve function candidates for each task. Your existing single-label classifier can also score those individual tasks. Keep the original message alongside them so negation, references and conditions survive. Splitting strings on "and" would make a fairly terrible planner.
I tried three versions of the same compound request with the pinned E5 model and the existing registry:
| Request | Cancellation rank | Invoice rank |
|---|---|---|
| "Cancel order 5901 and send me its invoice." | 1 | 2 |
| "Cancela el pedido 5901 y envíame su factura." | 1 | 3 |
| "Cancela order 5901 and send me the invoice." | 2 | 1 |
Both actions reached the top five in every case. With two manually written clauses per request, each intended function ranked first for its clause. This checks retrieval on six supplied clauses; automatic task extraction still needs its own evaluation.
The final plan needs arguments, conditions and action order. "Cancel it if it has not shipped" requires an order-state check. "Cancel orders 5901 and 5902" needs two task records even though a multi-label head returns cancel_order once. If a message also asks for unsupported gift wrapping, preserve that task so the response explains what the shop can do.
Several languages, one set of function IDs
An English request and a Spanish request can both map to change_address. Keep that ID fixed across languages. The same applies to requests that mix languages, such as the third row above.
XLM-R's shared encoder makes multilingual intent training practical. I would keep the trained classifier when its labels are stable and its language coverage matches the traffic. Adding Portuguese reuses the existing head if the intents stay the same; add local examples and continue training where the errors warrant it. English-only training can transfer to other languages, but XLM-R's 100-language pretraining does not establish equal routing quality in all of them.
For a new language with few labelled requests, multilingual embeddings offer a useful starting point. English function descriptions can be matched against Spanish requests in the same vector space, as the multilingual E5 report describes. The E5-small model card also warns about weaker results in low-resource languages. Keep the query: and passage: prefixes for non-English text.
Add reviewed local descriptions or example requests when product vocabulary needs them, and aggregate their matches by function ID. Otherwise, five translations of cancellation could fill the whole shortlist. Evaluate mixed-language messages as their own group, including negation and informal spelling. Translating everything into English adds another model decision that can alter those details.
For multilingual policy questions, retrieve the relevant documents and let the answer stage use that evidence. A classifier can choose the documentation route, while RAG supplies the current policy. For supported actions, the choice remains trained labels versus retrieved function candidates, with separate shortlists when the message contains several tasks.
Handle unsupported and uncertain requests
A closed-set classifier or description retriever will rank supported functions even for "Translate this poem". Give the application a rejection path for requests beyond its catalogue. The CLINC out-of-scope study shows why known-intent classification and unsupported-request detection need separate evaluation.
For single-label XLM-R, use realistic unsupported examples if you train an out_of_scope label. Include near misses such as "Change my airline booking", not just unrelated sentences. Choose acceptance rules from the strongest score and its gap from competing classes using separate held-out supported and unsupported data. For multi-label models, check each label's threshold and whether the predicted set covers the request; two strong scores can represent two wanted actions.
Single-label classifier probabilities can need calibration. Guo and colleagues describe temperature scaling as a practical method. Embedding cosine scores need their own acceptance rules; E5 often produces similarities around 0.7 to 1.0. Their reliability comes from checking what happens at those scores on your requests.
Track accepted accuracy alongside coverage, plus unsupported requests sent down the supported path. A stricter threshold may prevent mistakes while turning away useful traffic. Use clarification for missing details and a blocked response for an understood action prevented by order state.
Where document RAG fits
"What is the refund policy for damaged items?" needs policy text. "Refund order 1234" needs an operation and its arguments. The RAG paper describes generation conditioned on retrieved external material, which fits the policy question.
An intent classifier can route that question to a documentation tool, which retrieves the relevant policy. Embedding function descriptions helps select an operation; embedding policy passages helps supply evidence for an answer. A workflow can use both at different stages.
You can improve retrieval with task-specific pairs and hard negatives, or add a cross-encoder reranker that reads the request and description together. A frozen sentence encoder with logistic regression, or SetFit, is another supervised baseline if you want cheaper training than updating the full XLM-R encoder.
Check the errors that affect deployment
Compare top-one accuracy and recall at five on the same supported requests and registry version. Recall at five measures how often the correct intent reaches the LLM. A selector restricted to that list can succeed only when the right candidate is included.
For multiple intents, also report label precision and recall, plus exact-set accuracy: did the prediction include every requested label and no extra ones? For shortlist routing, count how often every requested action survives task extraction and gets its correct function in its candidate list.
Keep related paraphrases and conversations together when splitting data. Reserve separate data for threshold selection and final evaluation. Break results down by language, mixed-language requests and intent count, then inspect unsupported and ambiguous cases separately. The small English/Spanish runs above cover only their listed requests.
Measure the whole request path: p50 and p95 latency, prompt tokens and the share of requests sent to the LLM. XLM-R and E5-small have different model sizes and serving costs, so the deployed combination matters. A useful classifier may avoid some routing calls; a compact retriever may be easier to serve on the hardware you have.
For this shop, I would keep the existing classifier for stable actions in the languages it handles well. Frequent compound requests would justify multi-label training. For new functions, new languages with little labelled data, and documentation, I would start with retrieval and preserve a separate shortlist for each requested action.