Imbalanced Classification with XGBoost: scale_pos_weight, sample_weight, and Threshold Tuning

Class imbalance changes the distribution of training signals, not merely the choice of evaluation metric. In binary gradient boosting, each boosting round fits trees to gradients of the current objective. When one class has many more training rows, that class can contribute more row-level gradient signals unless the objective is reweighted. The optimizer may therefore reduce overall loss by predicting the majority more reliably while continuing to miss many minority-class instances. The objective is doing what it was asked to do; the problem is that aggregate performance can conceal the kind of error that matters operationally.

This distinction matters because the remedy depends on what the model must deliver. Reweighting can improve a ranking-oriented model, while calibrated probability estimation requires preserving the class prevalence represented by the data. Threshold tuning serves a third purpose: choosing where to convert fixed model scores into class labels. Training weights, probability interpretation, and decision thresholds should therefore be treated as separate design decisions.

What imbalance changes in the training objective

The documented default eval_metric for classification is logloss. A tree can improve this probabilistic objective while offering little improvement for the rare positive class, because most rows may agree with a majority-oriented prediction. Hard-label error can hide the same problem: it aggregates false positives and false negatives without showing which class contributes each mistake.

Gradient boosting does not reserve a fixed set of trees for each class. Its updates follow the gradients of the weighted training objective. Without additional weighting, a skewed data set can therefore produce an optimization process dominated by the majority class. Early trees may repeatedly find patterns that improve predictions for common cases, and later trees may receive weak or inconsistent positive-class gradients from the comparatively small positive sample.

This does not mean that every imbalanced problem fails, or that one metric must always be used. It means that class prevalence, the per-row contribution to the loss, and the operational definition of success must be considered together. Changing an evaluation metric does not change those training dynamics; changing the objective’s weights does.

scale_pos_weight: class balancing during training

The XGBoost parameter reference documents scale_pos_weight as the parameter that controls the balance between positive and negative weights. Its default is 1, meaning that no additional class-balancing adjustment is requested.

The same reference gives this starting value:

sum(negative instances) / sum(positive instances)

This is the documented value to consider for an imbalanced binary problem. It uses the aggregate representation of the two classes rather than requiring a practitioner to invent a weight from intuition. It is a starting point, not a universal optimum, and the reference does not establish that the resulting model is best for any particular data set.

How it changes the model

scale_pos_weight affects the weights used while learning the trees. Its influence is therefore baked into the tree structure and leaf outputs. A differently weighted objective can lead the booster to devote more of its capacity to patterns that distinguish positive instances, even when positive rows remain the minority.

The parameter does not supply an inference-time decision threshold. A model trained with scale_pos_weight can assign different scores—and can consequently produce different hard predictions at the default threshold—because the learned function changed. But the explicit classification rule remains the one implemented by the evaluation metric: XGBoost’s error metric treats prediction values larger than 0.5 as positive. Supplying error@t changes that threshold explicitly.

This distinction prevents a common category error:

  • scale_pos_weight changes the objective used to learn the model.
  • error@t changes how predictions are classified during evaluation.
  • A precision-recall threshold sweep changes the operating point without changing the fitted model.

Because class weighting changes the distribution seen by the learner, it also has implications for predicted probabilities. That consequence is desirable in some ranking applications and undesirable when probabilities must retain their original prevalence interpretation.

sample_weight: control at the instance level

The XGBoost Python API defines sample_weight as instance weights. Its signature places it directly in XGBModel.fit:

fit(
    X,
    y,
    *,
    sample_weight=None,
    base_margin=None,
    eval_set=None,
    verbose=True,
    xgb_model=None,
    sample_weight_eval_set=None,
    base_margin_eval_set=None,
    feature_weights=None,
)

The documented default is None. When weights are supplied, the training call takes the form:

model.fit(X_train, y_train, sample_weight=sample_weight)

Each training observation receives its corresponding weight in the objective. The weights must remain aligned with the rows of X and y; an accidental reordering would attach an intended weight to the wrong instance. The signature also provides sample_weight_eval_set, allowing weights to be supplied for evaluation sets rather than silently assuming that training weights describe validation observations.

How it differs from scale_pos_weight

Both parameters influence training rather than directly setting a prediction threshold, but they operate at different levels of control:

  • scale_pos_weight is a single class-balancing parameter for the positive-versus-negative weighting problem.
  • sample_weight is an instance-weight vector and can distinguish observations that would otherwise receive the same class-level treatment.
  • scale_pos_weight provides a documented class-ratio starting point; sample_weight requires the practitioner to construct and maintain the desired weight for each row.
  • scale_pos_weight changes aggregate class influence. sample_weight controls row-level influence directly and may encode class importance, observation-specific costs, or other intended distinctions.

The extra control offered by sample_weight has a real operational cost. A weight must be created, stored, ordered, and carried through preprocessing and model fitting. It must also be audited at validation and prediction pipeline boundaries. The vector becomes another data artifact whose correspondence to the training rows must be preserved. scale_pos_weight avoids that per-row bookkeeping when the only intended intervention is class balancing.

Neither parameter makes a hard-label threshold irrelevant. If the goal is to control false positives and false negatives after fitting, threshold tuning addresses that decision directly. Reweighting can alter which observations receive high scores, but it does not encode the application’s final classification rule.

Preserve class representation before tuning

Imbalance affects validation design before it affects hyperparameter selection. If positive observations are concentrated in one split, validation results can reflect an accidental change in class representation rather than the model’s general behavior.

The XGBoost scikit-learn interface documentation shows a stratified holdout split:

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, stratify=y, random_state=94
)

Passing stratify=y asks the split to preserve class representation instead of allocating rows purely at random. The documented random_state makes that split reproducible across calls.

For repeated validation, StratifiedKFold provides class-wise stratified folds while preserving the percentage of samples belonging to each class:

from sklearn.model_selection import StratifiedKFold

stratified_cv = StratifiedKFold(
    n_splits=5,
    shuffle=True,
    random_state=94,
)

The documented default for n_splits is 5, and the documented default for shuffle is False. Setting shuffle=True with an integer random_state provides reproducible shuffled folds.

Stratification solves what the scikit-learn documentation calls an engineering problem rather than a statistical one: it keeps each class represented across the splits. It does not rebalance the training objective, create additional positive rows, or make rare-class modeling easier by itself. Its role is to make the evaluation procedure less dependent on where minority observations happen to fall.

This separation is important. Stratification controls where rows are evaluated. scale_pos_weight and sample_weight control how rows influence training. A threshold controls which predicted scores become positive labels.

error, error@t, and the classification threshold

The XGBoost parameter reference defines error as:

[ \text{error} = \frac{#(\text{wrong cases})}{#(\text{all cases})}. ]

For binary classification, a prediction value larger than 0.5 is regarded as positive; the others are regarded as negative. This metric is therefore tied to a particular operating point.

The variant error@t replaces that default with a supplied numerical threshold t. For example, asking for error at a different threshold changes which predictions count as positive and can change the resulting error. It does not refit the trees, alter their learned scores, or change logloss.

Threshold-dependent error is useful when the hard-label operating point is the final output. Under imbalance, however, the aggregate error can remain dominated by the majority class. A lower overall error does not by itself establish that enough positive instances are being detected. The application must decide which balance of false positives and false negatives is acceptable.

The documented classification default remains logloss, not error. This distinction reflects the difference between evaluating hard classifications and evaluating the quality of probabilistic predictions. Choosing a threshold sweep later does not turn that classification decision back into a training change.

What auc and aucpr measure

The parameter reference defines auc as the area under the receiver operating characteristic curve. For binary classification, XGBoost expects an objective such as binary:logistic, or a similar objective that works with probabilities. AUC summarizes how well observations are ordered as the positive decision threshold varies; it is not a statement that a predicted probability equals the observed event frequency.

That distinction explains the official two-branch guidance. If the only concern is overall ranking performance measured by AUC, probability calibration is not part of the objective, and the documentation permits balancing positive and negative weights with scale_pos_weight.

aucpr is the area under the precision-recall curve. Unlike mean average precision, XGBoost documents aucpr as an interpolated area computed with continuous interpolation. For binary classification, its classification requirements and restrictions became similar to those of auc after XGBoost 1.6.

The precision-recall view exposes a different aspect of classifier behavior. Its axes are directly connected to positive detections:

[ \text{precision} = \frac{tp}{tp + fp} ]

[ \text{recall} = \frac{tp}{tp + fn} ]

Here, tp is the number of true positives, fp the number of false positives, and fn the number of false negatives.

The precision_recall_curve documentation provides two useful boundary facts. The first precision and recall values correspond to a classifier that always predicts the positive class: precision equals the class balance and recall is 1.0. The final precision and recall values are 1. and 0., respectively, and have no corresponding threshold.

For distributed XGBoost workloads, the parameter reference also warns that AUC calculation is not identical to the single-machine calculation: the distributed result is a weighted average over worker-level AUC values and is sensitive to how data are distributed. When precision and reproducibility are important, the documentation recommends another metric rather than treating the distributed AUC as exact.

Sweep a precision-recall threshold without retraining

A fitted model’s positive scores can be retained while the classification threshold changes. The scikit-learn function computes precision-recall pairs at different probability thresholds:

from sklearn.metrics import precision_recall_curve

precision, recall, thresholds = precision_recall_curve(y, probability)

At each threshold, predictions above the threshold contribute to the confusion matrix from which precision and recall are computed. Moving the threshold changes the operating point: it accepts a different balance of false positives and false negatives. The booster, its trees, its leaf outputs, and the underlying probability scores remain unchanged.

This is the same operational distinction represented by error@t, but the precision-recall sweep shows the trade-off rather than reporting only one aggregate error. The returned endpoints require care: the final precision-recall pair has no corresponding threshold, so the last precision and recall values must not be treated as if they came from the final entry in thresholds.

The threshold should be selected according to the application’s requirements. Neither the default 0.5 nor the PR curve establishes a universally correct operating point. What matters is whether the resulting precision and recall are appropriate for the decisions that will be made from the predictions.

Threshold tuning has strict boundaries:

  • It does not retrain or rebalance XGBoost.
  • It does not change the ordering used by auc or aucpr.
  • It does not repair probability calibration.
  • It changes only the map from fixed model scores to positive and negative decisions.

For that reason, threshold selection belongs after model fitting, not inside the definition of scale_pos_weight.

When calibrated probabilities are the requirement

A calibrated probability is not merely a score whose positive class has been selected successfully. It is a probability whose numerical value agrees with the event frequency represented by the underlying data.

Class rebalancing changes the class distribution seen by the learner. When scale_pos_weight, or a class-balancing design implemented through sample_weight, increases the influence of positive rows, the model is trained under a different effective class prior. Its raw probability outputs can therefore cease to represent prevalence in the original population. Moving the threshold cannot correct this: threshold tuning discards information by producing hard labels, while calibration concerns the probability values themselves.

The official XGBoost parameter-tuning tutorial makes the distinction explicit. It presents two branches:

  • If only overall performance measured by AUC matters, balance positive and negative weights with scale_pos_weight and use AUC for evaluation.
  • If the goal is to predict the right probability, do not rebalance the data set; set max_delta_step to a finite value, with 1 given as an example, to help convergence.

For the probability branch, max_delta_step has a documented default of 0. A value of 0 imposes no constraint. A positive value limits each leaf output update and makes the update more conservative. The parameter reference says that 1–10 may help control updates in logistic regression when the class is extremely imbalanced.

Unlike scale_pos_weight, max_delta_step does not rebalance positive and negative weights. It changes how far a tree leaf may move during an update. The tutorial presents it as a convergence aid for the probability-oriented branch, not as a guarantee that every resulting probability is calibrated.

A disciplined workflow

A reliable imbalanced-classification workflow keeps the decisions separate:

  • Use stratified splitting and stratified cross-validation so class representation is preserved across evaluation sets.
  • Decide whether the model is intended to rank cases, estimate probabilities, or produce hard labels.
  • For ranking-focused training, evaluate scale_pos_weight or per-instance sample_weight rather than assuming that an unweighted fit reflects the desired balance.
  • For probability-focused training, avoid class rebalancing when raw probabilities must retain the original prevalence; consider the documented max_delta_step alternative.
  • Evaluate classification quality with metrics appropriate to the target, including the distinction between error, error@t, auc, and aucpr.
  • Sweep a precision-recall threshold only after obtaining fixed model scores, and remember that the result changes labels rather than the model.

The central rule is simple: class weights determine what the boosting objective learns, probability interpretation determines what its numerical outputs mean, and threshold tuning determines where those outputs become decisions. Treating those as interchangeable is the most common source of an imbalanced model that appears strong on one metric but fails in practice.