machine learning

Build Your Own Decision Model from a Language Model

Build Your Own Decision Model from a Language Model

A language model usually looks like a writer. You give it a prompt, and it keeps producing tokens until it reaches a stopping point. That works well for essays, code, and conversations, but it is wasteful when the answer must be one of a few known choices.

A decision model has a narrower job: inspect the input, score a fixed set of options, and select one. This resembles a system-one-style decision process, where the model makes a fast choice instead of composing a long explanation. The useful question is how to turn a language model into that kind of fixed-choice system without retraining an enormous network.

The next token is already a classifier

A causal language model predicts what comes next. A token is a small piece of text, such as a word, part of a word, punctuation, or whitespace attached to a word. After reading a prompt that ends with Answer:, the model produces a vector of logits: raw scores for every token in its vocabulary.

A function called softmax converts those scores into numbers that add up to one. During normal generation, the model samples or selects one token, adds it to the prompt, and predicts the next token again. For a multiple-choice task, we can stop after the first answer token and inspect only the labels we allow, such as A, B, C, D, and E.

Masking is the usual term for removing forbidden choices. We set their logits to negative infinity, which makes softmax assign them zero probability. In code, selecting only the allowed logits and applying softmax to that smaller vector produces the same result.

Code that scores instead of generating

Here is a compact decision model using Qwen/Qwen3-1.7B. The important detail is that the code never calls generate. It performs one forward pass over the prompt, reads the scores at the final position, and normalizes the five allowed labels.

import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = 'Qwen/Qwen3-1.7B'
labels = ['A', 'B', 'C', 'D', 'E']

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
 model_id,
 torch_dtype='auto',
 device_map='auto',
)
model.eval

def make_prompt(item):
 lines = [item['question']]
 lines.extend(f'{label}. {item[label]}' for label in labels)
 messages = [{
 'role': 'user',
 'content': '\n'.join(lines) + '\nAnswer:'
 }]
 return tokenizer.apply_chat_template(
 messages,
 tokenize=False,
 add_generation_prompt=True,
 enable_thinking=False,
 )

def one_token_id(label):
 # The leading space must match the expected continuation.
 ids = tokenizer.encode(f' {label}', add_special_tokens=False)
 if len(ids)!= 1:
 raise ValueError(f'{label} is not represented by one token')
 return ids[0]

item = json.load(open('question.json'))
text = make_prompt(item)
inputs = tokenizer(text, return_tensors='pt').to(model.device)
option_ids = [one_token_id(label) for label in labels]

with torch.inference_mode:
 logits = model(**inputs).logits[0, -1].float
 option_logits = logits[option_ids]
 probabilities = torch.softmax(option_logits, dim=-1)

winner = labels[probabilities.argmax.item]
print(f'prediction: {winner}')
for label, probability in zip(labels, probabilities.tolist):
 print(label, round(probability, 4))

Tokenization deserves attention here. Many tokenizers represent A and A as different tokens, so the candidate must match the exact continuation expected after the prompt. The assertion also protects against labels that split into multiple tokens. If your choices are long phrases rather than short labels, score the complete token sequences instead of looking at the first token alone.

The enable_thinking=False setting keeps the model focused on producing the answer label immediately. If the model writes an internal reasoning block first, the final answer is no longer the next token and this one-step method no longer measures the intended decision.

A probability is not a promise

The five numbers printed by this program are conditional probabilities. They answer a question like: “Given that the next token must be one of these five labels, which label does the model prefer?” They do not automatically mean that the winning answer has that chance of being factually correct.

Imagine a question such as “Where would you expect to see a crane?” with choices involving a construction site, a wetland, a hardware store, and a shipping dock. The model may assign 99 percent to one option because the wording strongly resembles examples it has seen, even though the question is ambiguous. Constrained decoding makes the output tidy, but it does not create new knowledge or guarantee correctness.

This is the difference between accuracy and calibration. Accuracy asks whether the selected label is right. Calibration asks whether confidence matches reality. If a model makes 100 predictions near 70 percent confidence, a well-calibrated model should be correct roughly 70 times.

Temperature changes confidence, not the winner

A common post-training calibration method is temperature scaling. Post-training means the original model weights stay fixed while a small correction is learned afterward. Given option logits z, calibrated probabilities are calculated as:

def calibrated_probs(option_logits, temperature):
 return torch.softmax(option_logits / temperature, dim=-1)

A temperature of 1.0 leaves the distribution unchanged. A value greater than 1.0 flattens the distribution, reducing overconfidence. A value below 1.0 sharpens it. Because every option is divided by the same positive number, the ordering normally stays the same, so calibration changes the confidence scale without changing the selected answer.

The temperature should be fitted on a validation set by minimizing negative log-likelihood, a loss that penalizes confident mistakes more heavily than cautious ones. A convenient implementation learns log_temperature, then uses log_temperature.exp so the resulting temperature cannot become negative. Do not fit this parameter on the final test set; that would make the evaluation optimistic.

One useful summary is expected calibration error. Divide predictions into confidence bins, compare average confidence with actual accuracy in each bin, and weight the gaps by the number of examples. A model that says 98 percent but is correct only 70 percent of the time will show a large calibration error even if its raw accuracy is respectable.

Evaluation is part of the model

A fixed-choice interface makes evaluation much clearer. Run the decision model over a holdout set with known answers, record the winning label, and measure accuracy. In the original small-model experiment on CommonsenseQA, the unmodified setup reached roughly 59 percent accuracy, while a quick fine-tune raised that result to about 62 percent. Those figures are useful reference points, not guarantees: prompt formatting, tokenizer behavior, model revisions, and dataset splits can all move the result.

Keep accuracy and calibration as separate measurements. Also inspect errors by label. A model that predicts A too often may look acceptable overall while failing badly whenever the correct answer is D or E. Randomizing the order of answer choices during evaluation can reveal this kind of position bias.

Several small implementation details have an outsized effect. Use the same chat template for calibration and testing. Verify that every label is the token you intended. Preserve a separate calibration split. If you need to enforce the answer during actual generation, apply the vocabulary mask with a constrained decoder or logits processor; the scoring code above ranks choices but does not prevent a later generator from emitting other text.

Building a decision model does not turn a language model into an oracle. It gives the model a smaller action space, exposes the scores behind its choice, and creates room to measure whether those scores deserve trust. One next-token distribution can be enough to choose an option. Calibration determines how honestly that choice is presented.

ahsan

ahsan

Hello! I am Mr Ahsan, the writer of the Website. I am from Netherland. I like to write about technology and the news around it.

Comments (0)

No comments yet. Be the first to respond!

Leave a Comment

Your comment will be visible after review.