Learning Record learning from practice

· ai / train

Train Study

from datasets import load_dataset
from transformers import AutoTokenizer
from transformers import DataCollatorWithPadding
from transformers import TrainingArguments
from transformers import AutoModelForSequenceClassification
import evaluate, numpy as np

raw_datasets = load_dataset("glue", "mrpc")

checkpoint = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)

def tokenize_function(example):
    return tokenizer(example["sentence1"], example["sentence2"], truncation=True)

tokenized_datasets = raw_datasets.map(tokenize_function, batched=True)

data_collator = DataCollatorWithPadding(tokenizer=tokenizer)

training_args = TrainingArguments("test-trainer", eval_strategy="epoch")

model = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)

metric = evaluate.load("glue", "mrpc")

def compute_metrics(eval_preds):
    logits, labels = eval_preds
    predictions = np.argmax(logits, axis=-1)
    return metric.compute(predictions=predictions, references=labels)

from transformers import Trainer

trainer = Trainer(
    model,
    training_args,
    train_dataset=tokenized_datasets["train"],
    eval_dataset=tokenized_datasets["validation"],
    data_collator=data_collator,
    processing_class=tokenizer,
    compute_metrics=compute_metrics,
)

trainer.train()

Why Training Runs for Three Epochs

The training did not run for only two epochs because the dataset is small. In this code:

training_args = TrainingArguments(
    "test-trainer",
    eval_strategy="epoch"
)

num_train_epochs is not set explicitly. Its default value in TrainingArguments is 3.0, so the final training output correctly reports:

'epoch': 3.0

At the same time, eval_strategy="epoch" runs evaluation at the end of every epoch. That is why the output contains:

epoch: 1.0
epoch: 2.0
epoch: 3.0

The epoch: 1.09 and epoch: 2.18 entries are intermediate training log points, not additional epochs.

To train for exactly two epochs, set the value explicitly:

training_args = TrainingArguments(
    output_dir="test-trainer",
    eval_strategy="epoch",
    num_train_epochs=2,
)

The MRPC dataset has about 3,668 training examples and 408 validation examples. Its small size mainly means that each epoch finishes quickly; it does not change the default number of epochs. The result from this run was:

  • Epoch 1: accuracy 0.694
  • Epoch 2: accuracy 0.809
  • Epoch 3: accuracy 0.838

The validation metric was still improving at epoch 3. However, because MRPC is a small dataset, watch for overfitting when increasing the number of epochs.

Understanding the MRPC Task and Evaluation Logs

This example is a binary classification task from the GLUE benchmark. More specifically, it uses the Microsoft Research Paraphrase Corpus (MRPC). Each input is a pair of sentences:

sentence1
sentence2

The model predicts whether the two sentences express the same or a very similar meaning. This is a paraphrase judgment; it is not a test of whether the two strings are exactly identical.

The labels are normally interpreted as:

label = 0  # not equivalent: the meanings are different
label = 1  # equivalent: the meanings are the same or similar

The model is configured with two output classes:

model = AutoModelForSequenceClassification.from_pretrained(
    checkpoint,
    num_labels=2,
)

For a batch, the classification head therefore returns two logits per example:

logits.shape == (batch_size, 2)

Logits are scores before softmax. The prediction code selects the index of the larger score:

predictions = np.argmax(logits, axis=-1)

For example, if one row is [1.2, 2.8], the predicted class is 1. The two scores do not need to be probabilities, and they do not have to add up to one.

What the Metrics Mean

eval_loss is the loss on the validation set. Because this is a single-label classification model with num_labels=2, the model uses cross-entropy loss when labels are supplied. The loss penalizes the model more strongly when it assigns a low probability to the correct class. The value reported for evaluation is an aggregate over the validation batches.

For example:

eval_loss: 0.625
eval_loss: 0.430
eval_loss: 0.518

The second evaluation has the lowest loss, while the third evaluation can still have better accuracy or F1. Loss and classification metrics measure different aspects of the predictions, so model quality should not be judged from loss alone.

eval_accuracy is the fraction of validation examples classified correctly:

\[\text{accuracy} = \frac{\text{number of correct predictions}} {\text{number of validation examples}}\]

With about 408 validation examples, an accuracy of 0.838235 corresponds to approximately:

408 * 0.838235 ~= 342 correct predictions

eval_f1 is the F1 score. MRPC commonly reports both accuracy and F1 because the class distribution is not perfectly balanced. F1 is the harmonic mean of precision and recall:

\[F1 = \frac{2 \times \text{precision} \times \text{recall}} {\text{precision} + \text{recall}}\]

In this example, evaluate.load("glue", "mrpc") computes the metrics defined for the MRPC task, including accuracy and F1.

eval_runtime is the time spent on one validation run, measured in seconds. For example:

eval_runtime: 13.813

This is about 13.8 seconds for that evaluation. It covers the evaluation loop, including iterating over the validation dataloader, running forward passes, collecting predictions, and computing the requested metrics. It is not the total training time. The total time for training is reported separately as:

train_runtime: 93.639

Why Evaluation Throughput Changes

eval_samples_per_second is approximately calculated as:

\[\text{eval\_samples\_per\_second} = \frac{\text{number of validation samples}} {\text{eval\_runtime}}\]

Therefore, when the number of validation samples is fixed, a longer runtime produces a lower throughput. The observed values were:

13.813 seconds -> 29.537 samples/second
2.643 seconds  -> 154.366 samples/second
3.168 seconds  -> 128.791 samples/second

The first evaluation can be slower because it may include one-time initialization costs such as CUDA setup, memory allocation, kernel loading, data loading, or cache initialization. CPU/GPU contention and other background work can also affect the measurement. These costs matter more when the validation set is small. Consequently, the first 29.537 samples/second value may not represent steady-state throughput; the later values are more representative of this run. Some variation between later evaluations is still normal.

Why the Logged Epoch Can Be Fractional

Evaluation is configured by epoch, but ordinary training logs are configured by steps. The default logging strategy is step-based, with logging_steps=500. The train split has about 3,668 examples, and the default training batch size is 8, so one epoch contains approximately:

ceil(3668 / 8) = 459 training steps

After 500 update steps, the progress expressed in epochs is approximately:

\[\text{epoch} = \frac{500}{459} \approx 1.09\]

Thus, a log entry such as:

{'loss': ..., 'epoch': 1.09}

means that training has progressed about 1.09 epochs at that step. It does not mean that a separate partial epoch was configured.

The log sequence can be read as follows:

eval ... epoch: 1.0  -> the first epoch ended, then evaluation ran
loss ... epoch: 1.09 -> a step-based training log after entering epoch 2
eval ... epoch: 2.0  -> the second epoch ended, then evaluation ran
loss ... epoch: 2.18 -> another step-based training log during epoch 3
eval ... epoch: 3.0  -> the third epoch ended and training finished

The exact fractional values depend on the dataset size, batch size, gradient accumulation, and logging interval. The important distinction is that 1.0, 2.0, and 3.0 are epoch-boundary evaluations, while 1.09 and 2.18 are intermediate step-based logs. The run still trained for the default three complete epochs.

The metric-loading pattern in this example is already the recommended form:

metric = evaluate.load("glue", "mrpc")

def compute_metrics(eval_preds):
    logits, labels = eval_preds
    predictions = np.argmax(logits, axis=-1)
    return metric.compute(
        predictions=predictions,
        references=labels,
    )

Loading the metric once avoids repeating its initialization every time compute_metrics is called. It can reduce a small amount of evaluation overhead while leaving the computed results unchanged.

Detecting Overfitting in a Longer Run

The three-epoch run was still improving, but a longer run shows a clearer overfitting pattern. The following results summarize an extended run:

Epoch Training loss Validation loss Accuracy F1
1 not directly reported 0.448 0.787 0.861
2 not directly reported 0.398 0.833 0.875
3 about 0.20 0.573 0.868 0.906
4 about 0.09 0.694 0.865 0.906
5 about 0.06 0.946 0.858 0.901
10 about 0.01 1.015 0.858 0.901

The training loss falls from about 0.55 to almost zero, which means that the model is fitting the training examples increasingly closely. In contrast, validation loss reaches its minimum at epoch 2 and then increases. Accuracy and F1 reach their best values around epoch 3 and then stop improving or decline slightly. This separation is the characteristic pattern to watch for:

training loss continues to decrease
validation loss starts to increase
validation metrics peak and then stagnate or decrease

The choice of the best epoch depends on the selected objective. Epoch 3 is best by F1 and accuracy in this run, while epoch 2 is best by validation loss. A lower loss is not guaranteed to produce a higher thresholded accuracy or F1 score, because loss also reflects the model’s confidence in each prediction. The validation set contains about 408 examples, so changing one prediction changes accuracy by approximately:

\[\frac{1}{408} \approx 0.245\%\]

Therefore, small accuracy changes can be sampling noise from a relatively small validation set. The stronger signal here is the sustained increase in validation loss while training loss approaches zero.

Keeping the Best Checkpoint

If the goal is to run for at most 10 epochs but return the checkpoint with the best validation F1, configure evaluation and saving at the same interval and enable best-model loading:

from transformers import TrainingArguments

training_args = TrainingArguments(
    output_dir="test-trainer",
    eval_strategy="epoch",
    save_strategy="epoch",
    num_train_epochs=10,
    load_best_model_at_end=True,
    metric_for_best_model="f1",
    greater_is_better=True,
)

This configuration still allows training to continue for up to 10 epochs, but restores the checkpoint with the best evaluation F1 when training ends. In the example above, epoch 3 is the apparent best checkpoint by F1; the displayed values are rounded, so an exact tie or a different ordering in the unrounded values should be resolved from the Trainer’s recorded metric values. The selected metric can instead be loss when the validation loss is the objective; in that case, use metric_for_best_model="loss" and greater_is_better=False.

Stopping When Improvement Ends

To stop training after the selected metric fails to improve for two evaluation calls, add EarlyStoppingCallback:

from transformers import EarlyStoppingCallback

trainer = Trainer(
    model,
    training_args,
    train_dataset=tokenized_datasets["train"],
    eval_dataset=tokenized_datasets["validation"],
    data_collator=data_collator,
    processing_class=tokenizer,
    compute_metrics=compute_metrics,
    callbacks=[
        EarlyStoppingCallback(early_stopping_patience=2),
    ],
)

Here, early_stopping_patience=2 means that training can stop after two evaluation calls without a new best value for the selected metric. Early stopping depends on the best-model tracking configuration, so load_best_model_at_end=True and a valid metric_for_best_model should remain enabled. With epoch-based evaluation and saving, the callback observes the metric at the same checkpoints and avoids the step/save interval mismatch described in the Transformers documentation.

grad_norm is not a direct overfitting measure. It describes the size of the gradients, not how well the model generalizes. For this run, the more useful evidence is the divergence between training loss and validation loss together with the plateau or decline in validation metrics.

References