· ai / train
Train Study
from datasets import load_dataset
from transformers import AutoTokenizer
from transformers import DataCollatorWithPadding
from transformers import TrainingArguments
from transformers import AutoModelForSequenceClassification
import evaluate, numpy as np
raw_datasets = load_dataset("glue", "mrpc")
checkpoint = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
def tokenize_function(example):
return tokenizer(example["sentence1"], example["sentence2"], truncation=True)
tokenized_datasets = raw_datasets.map(tokenize_function, batched=True)
data_collator = DataCollatorWithPadding(tokenizer=tokenizer)
training_args = TrainingArguments("test-trainer", eval_strategy="epoch")
model = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)
metric = evaluate.load("glue", "mrpc")
def compute_metrics(eval_preds):
logits, labels = eval_preds
predictions = np.argmax(logits, axis=-1)
return metric.compute(predictions=predictions, references=labels)
from transformers import Trainer
trainer = Trainer(
model,
training_args,
train_dataset=tokenized_datasets["train"],
eval_dataset=tokenized_datasets["validation"],
data_collator=data_collator,
processing_class=tokenizer,
compute_metrics=compute_metrics,
)
trainer.train()
Why Training Runs for Three Epochs
The training did not run for only two epochs because the dataset is small. In this code:
training_args = TrainingArguments(
"test-trainer",
eval_strategy="epoch"
)
num_train_epochs is not set explicitly. Its default value in TrainingArguments is 3.0, so the final training output correctly reports:
'epoch': 3.0
At the same time, eval_strategy="epoch" runs evaluation at the end of every epoch. That is why the output contains:
epoch: 1.0
epoch: 2.0
epoch: 3.0
The epoch: 1.09 and epoch: 2.18 entries are intermediate training log points, not additional epochs.
To train for exactly two epochs, set the value explicitly:
training_args = TrainingArguments(
output_dir="test-trainer",
eval_strategy="epoch",
num_train_epochs=2,
)
The MRPC dataset has about 3,668 training examples and 408 validation examples. Its small size mainly means that each epoch finishes quickly; it does not change the default number of epochs. The result from this run was:
- Epoch 1: accuracy
0.694 - Epoch 2: accuracy
0.809 - Epoch 3: accuracy
0.838
The validation metric was still improving at epoch 3. However, because MRPC is a small dataset, watch for overfitting when increasing the number of epochs.
Understanding the MRPC Task and Evaluation Logs
This example is a binary classification task from the GLUE benchmark. More specifically, it uses the Microsoft Research Paraphrase Corpus (MRPC). Each input is a pair of sentences:
sentence1
sentence2
The model predicts whether the two sentences express the same or a very similar meaning. This is a paraphrase judgment; it is not a test of whether the two strings are exactly identical.
The labels are normally interpreted as:
label = 0 # not equivalent: the meanings are different
label = 1 # equivalent: the meanings are the same or similar
The model is configured with two output classes:
model = AutoModelForSequenceClassification.from_pretrained(
checkpoint,
num_labels=2,
)
For a batch, the classification head therefore returns two logits per example:
logits.shape == (batch_size, 2)
Logits are scores before softmax. The prediction code selects the index of the larger score:
predictions = np.argmax(logits, axis=-1)
For example, if one row is [1.2, 2.8], the predicted class is 1. The two scores do not need to be probabilities, and they do not have to add up to one.
What the Metrics Mean
eval_loss is the loss on the validation set. Because this is a single-label classification model with num_labels=2, the model uses cross-entropy loss when labels are supplied. The loss penalizes the model more strongly when it assigns a low probability to the correct class. The value reported for evaluation is an aggregate over the validation batches.
For example:
eval_loss: 0.625
eval_loss: 0.430
eval_loss: 0.518
The second evaluation has the lowest loss, while the third evaluation can still have better accuracy or F1. Loss and classification metrics measure different aspects of the predictions, so model quality should not be judged from loss alone.
eval_accuracy is the fraction of validation examples classified correctly:
With about 408 validation examples, an accuracy of 0.838235 corresponds to approximately:
408 * 0.838235 ~= 342 correct predictions
eval_f1 is the F1 score. MRPC commonly reports both accuracy and F1 because the class distribution is not perfectly balanced. F1 is the harmonic mean of precision and recall:
In this example, evaluate.load("glue", "mrpc") computes the metrics defined for the MRPC task, including accuracy and F1.
eval_runtime is the time spent on one validation run, measured in seconds. For example:
eval_runtime: 13.813
This is about 13.8 seconds for that evaluation. It covers the evaluation loop, including iterating over the validation dataloader, running forward passes, collecting predictions, and computing the requested metrics. It is not the total training time. The total time for training is reported separately as:
train_runtime: 93.639
Why Evaluation Throughput Changes
eval_samples_per_second is approximately calculated as:
Therefore, when the number of validation samples is fixed, a longer runtime produces a lower throughput. The observed values were:
13.813 seconds -> 29.537 samples/second
2.643 seconds -> 154.366 samples/second
3.168 seconds -> 128.791 samples/second
The first evaluation can be slower because it may include one-time initialization costs such as CUDA setup, memory allocation, kernel loading, data loading, or cache initialization. CPU/GPU contention and other background work can also affect the measurement. These costs matter more when the validation set is small. Consequently, the first 29.537 samples/second value may not represent steady-state throughput; the later values are more representative of this run. Some variation between later evaluations is still normal.
Why the Logged Epoch Can Be Fractional
Evaluation is configured by epoch, but ordinary training logs are configured by steps. The default logging strategy is step-based, with logging_steps=500. The train split has about 3,668 examples, and the default training batch size is 8, so one epoch contains approximately:
ceil(3668 / 8) = 459 training steps
After 500 update steps, the progress expressed in epochs is approximately:
\[\text{epoch} = \frac{500}{459} \approx 1.09\]Thus, a log entry such as:
{'loss': ..., 'epoch': 1.09}
means that training has progressed about 1.09 epochs at that step. It does not mean that a separate partial epoch was configured.
The log sequence can be read as follows:
eval ... epoch: 1.0 -> the first epoch ended, then evaluation ran
loss ... epoch: 1.09 -> a step-based training log after entering epoch 2
eval ... epoch: 2.0 -> the second epoch ended, then evaluation ran
loss ... epoch: 2.18 -> another step-based training log during epoch 3
eval ... epoch: 3.0 -> the third epoch ended and training finished
The exact fractional values depend on the dataset size, batch size, gradient accumulation, and logging interval. The important distinction is that 1.0, 2.0, and 3.0 are epoch-boundary evaluations, while 1.09 and 2.18 are intermediate step-based logs. The run still trained for the default three complete epochs.
The metric-loading pattern in this example is already the recommended form:
metric = evaluate.load("glue", "mrpc")
def compute_metrics(eval_preds):
logits, labels = eval_preds
predictions = np.argmax(logits, axis=-1)
return metric.compute(
predictions=predictions,
references=labels,
)
Loading the metric once avoids repeating its initialization every time compute_metrics is called. It can reduce a small amount of evaluation overhead while leaving the computed results unchanged.
Detecting Overfitting in a Longer Run
The three-epoch run was still improving, but a longer run shows a clearer overfitting pattern. The following results summarize an extended run:
| Epoch | Training loss | Validation loss | Accuracy | F1 |
|---|---|---|---|---|
| 1 | not directly reported | 0.448 | 0.787 | 0.861 |
| 2 | not directly reported | 0.398 | 0.833 | 0.875 |
| 3 | about 0.20 | 0.573 | 0.868 | 0.906 |
| 4 | about 0.09 | 0.694 | 0.865 | 0.906 |
| 5 | about 0.06 | 0.946 | 0.858 | 0.901 |
| 10 | about 0.01 | 1.015 | 0.858 | 0.901 |
The training loss falls from about 0.55 to almost zero, which means that the model is fitting the training examples increasingly closely. In contrast, validation loss reaches its minimum at epoch 2 and then increases. Accuracy and F1 reach their best values around epoch 3 and then stop improving or decline slightly. This separation is the characteristic pattern to watch for:
training loss continues to decrease
validation loss starts to increase
validation metrics peak and then stagnate or decrease
The choice of the best epoch depends on the selected objective. Epoch 3 is best by F1 and accuracy in this run, while epoch 2 is best by validation loss. A lower loss is not guaranteed to produce a higher thresholded accuracy or F1 score, because loss also reflects the model’s confidence in each prediction. The validation set contains about 408 examples, so changing one prediction changes accuracy by approximately:
\[\frac{1}{408} \approx 0.245\%\]Therefore, small accuracy changes can be sampling noise from a relatively small validation set. The stronger signal here is the sustained increase in validation loss while training loss approaches zero.
Keeping the Best Checkpoint
If the goal is to run for at most 10 epochs but return the checkpoint with the best validation F1, configure evaluation and saving at the same interval and enable best-model loading:
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="test-trainer",
eval_strategy="epoch",
save_strategy="epoch",
num_train_epochs=10,
load_best_model_at_end=True,
metric_for_best_model="f1",
greater_is_better=True,
)
This configuration still allows training to continue for up to 10 epochs, but restores the checkpoint with the best evaluation F1 when training ends. In the example above, epoch 3 is the apparent best checkpoint by F1; the displayed values are rounded, so an exact tie or a different ordering in the unrounded values should be resolved from the Trainer’s recorded metric values. The selected metric can instead be loss when the validation loss is the objective; in that case, use metric_for_best_model="loss" and greater_is_better=False.
Stopping When Improvement Ends
To stop training after the selected metric fails to improve for two evaluation calls, add EarlyStoppingCallback:
from transformers import EarlyStoppingCallback
trainer = Trainer(
model,
training_args,
train_dataset=tokenized_datasets["train"],
eval_dataset=tokenized_datasets["validation"],
data_collator=data_collator,
processing_class=tokenizer,
compute_metrics=compute_metrics,
callbacks=[
EarlyStoppingCallback(early_stopping_patience=2),
],
)
Here, early_stopping_patience=2 means that training can stop after two evaluation calls without a new best value for the selected metric. Early stopping depends on the best-model tracking configuration, so load_best_model_at_end=True and a valid metric_for_best_model should remain enabled. With epoch-based evaluation and saving, the callback observes the metric at the same checkpoints and avoids the step/save interval mismatch described in the Transformers documentation.
grad_norm is not a direct overfitting measure. It describes the size of the gradients, not how well the model generalizes. For this run, the more useful evidence is the divergence between training loss and validation loss together with the plateau or decline in validation metrics.
References
- TrainingArguments documentation: documents the
num_train_epochs=3.0default and thateval_strategy="epoch"evaluates at the end of each epoch. - GLUE MRPC dataset: lists the MRPC train and validation split sizes.
- BERT sequence classification documentation: documents the classification head, logits shape, and cross-entropy behavior for multiple labels.
- Evaluate documentation: documents loading an evaluation module and computing metrics from predictions and references.
- Transformers callbacks documentation: documents
EarlyStoppingCallback, its patience behavior, and its dependency on best-model tracking.