{
  "id": 173852,
  "title": "My overall auc for one CV score lower than the average of each fold",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173852",
  "author_name": "",
  "post_date": "2020-08-11T02:27:57.435412300Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone, the average of val_auc score of my 5 folds should be about 0.93 round according to my own calculation while the overall auc only has 0.904. I 'm not show whether it is normal, and how can I deal with the problem.  Here I mainly use the baseline model of Chris <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a></p>",
  "messages": [
    {
      "id": "965955",
      "postDate": "08/11/2020 02:27:57",
      "content": "<p>Hi everyone, the average of val_auc score of my 5 folds should be about 0.93 round according to my own calculation while the overall auc only has 0.904. I 'm not show whether it is normal, and how can I deal with the problem.  Here I mainly use the baseline model of Chris <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a></p>",
      "rawMarkdown": "Hi everyone, the average of val_auc score of my 5 folds should be about 0.93 round according to my own calculation while the overall auc only has 0.904. I 'm not show whether it is normal, and how can I deal with the problem.  Here I mainly use the baseline model of Chris [here](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords)",
      "votes": null
    },
    {
      "id": "966119",
      "postDate": "08/11/2020 06:42:36",
      "content": "<p>This can happen if you train each fold a different number of epochs which happens when we save the best epoch's fold model for each fold. One solution is min-max normalization. Take the predictions for each fold and scale them between 0 and 1.</p>\n<p>If predictions for fold 1 are in variable <code>preds</code> then do this </p>\n<pre><code>new_preds = preds - np.min(preds)\nnew_preds = new_preds / np.max(new_preds)\n</code></pre>\n<p>So in my popular notebook, find this code</p>\n<pre><code>pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \noof_pred.append( np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1) )\n</code></pre>\n<p>And change to this code</p>\n<pre><code>pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \npred = np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1)\nnew_preds = preds - np.min(preds)\nnew_preds = new_preds / np.max(new_preds)\noof_pred.append( new_preds )\n</code></pre>\n<p>Then when the 5 folds get combined, they will match up better. AUC is a ranking metric, so if the fold predictions are not aligned, you can observe what you observed.</p>",
      "rawMarkdown": "This can happen if you train each fold a different number of epochs which happens when we save the best epoch's fold model for each fold. One solution is min-max normalization. Take the predictions for each fold and scale them between 0 and 1.\n\nIf predictions for fold 1 are in variable `preds` then do this \n\n    new_preds = preds - np.min(preds)\n    new_preds = new_preds / np.max(new_preds)\n  \nSo in my popular notebook, find this code\n\n    pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \n    oof_pred.append( np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1) )\n\nAnd change to this code\n\n    pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \n    pred = np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1)\n    new_preds = preds - np.min(preds)\n    new_preds = new_preds / np.max(new_preds)\n    oof_pred.append( new_preds )\n\nThen when the 5 folds get combined, they will match up better. AUC is a ranking metric, so if the fold predictions are not aligned, you can observe what you observed.",
      "votes": null
    },
    {
      "id": "966120",
      "postDate": "08/11/2020 06:42:55",
      "content": "<p>Maybe your score on different folds has different scale? For example if on the first fold score has mean -1 and on the second score has mean +50, then before concatenation you should rescale you scores on different folds to lead them to the same scale with mean 0.</p>",
      "rawMarkdown": "Maybe your score on different folds has different scale? For example if on the first fold score has mean -1 and on the second score has mean +50, then before concatenation you should rescale you scores on different folds to lead them to the same scale with mean 0.",
      "votes": null
    },
    {
      "id": "966122",
      "postDate": "08/11/2020 06:45:14",
      "content": "<p>Yes, exactly Sxwat. I posted code below showing how to rescale each fold.</p>",
      "rawMarkdown": "Yes, exactly Sxwat. I posted code below showing how to rescale each fold.",
      "votes": null
    },
    {
      "id": "966125",
      "postDate": "08/11/2020 06:47:45",
      "content": "<p>But note that it isn't so important to fix your problem. You can use the average of the 5 folds or the scaled merging of the 5 folds. Try both but the most important thing is to have a CV score (whether it be the 5 average or 5 merge) that goes up and down when LB goes up and down.</p>",
      "rawMarkdown": "But note that it isn't so important to fix your problem. You can use the average of the 5 folds or the scaled merging of the 5 folds. Try both but the most important thing is to have a CV score (whether it be the 5 average or 5 merge) that goes up and down when LB goes up and down.",
      "votes": null
    },
    {
      "id": "966195",
      "postDate": "08/11/2020 08:22:47",
      "content": "<p>Appreciate much !</p>",
      "rawMarkdown": "Appreciate much !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 966119,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/11/2020 06:42:36",
      "content": "<p>This can happen if you train each fold a different number of epochs which happens when we save the best epoch's fold model for each fold. One solution is min-max normalization. Take the predictions for each fold and scale them between 0 and 1.</p>\n<p>If predictions for fold 1 are in variable <code>preds</code> then do this </p>\n<pre><code>new_preds = preds - np.min(preds)\nnew_preds = new_preds / np.max(new_preds)\n</code></pre>\n<p>So in my popular notebook, find this code</p>\n<pre><code>pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \noof_pred.append( np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1) )\n</code></pre>\n<p>And change to this code</p>\n<pre><code>pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \npred = np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1)\nnew_preds = preds - np.min(preds)\nnew_preds = new_preds / np.max(new_preds)\noof_pred.append( new_preds )\n</code></pre>\n<p>Then when the 5 folds get combined, they will match up better. AUC is a ranking metric, so if the fold predictions are not aligned, you can observe what you observed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 966125,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/11/2020 06:47:45",
          "content": "<p>But note that it isn't so important to fix your problem. You can use the average of the 5 folds or the scaled merging of the 5 folds. Try both but the most important thing is to have a CV score (whether it be the 5 average or 5 merge) that goes up and down when LB goes up and down.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 966195,
          "author_name": "shenjiaxjtlu",
          "author_url": "",
          "post_date": "08/11/2020 08:22:47",
          "content": "<p>Appreciate much !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 966120,
      "author_name": "aybatov",
      "author_url": "",
      "post_date": "08/11/2020 06:42:55",
      "content": "<p>Maybe your score on different folds has different scale? For example if on the first fold score has mean -1 and on the second score has mean +50, then before concatenation you should rescale you scores on different folds to lead them to the same scale with mean 0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 966122,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/11/2020 06:45:14",
          "content": "<p>Yes, exactly Sxwat. I posted code below showing how to rescale each fold.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "965955": "Hi everyone, the average of val_auc score of my 5 folds should be about 0.93 round according to my own calculation while the overall auc only has 0.904. I 'm not show whether it is normal, and how can I deal with the problem.  Here I mainly use the baseline model of Chris [here](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords)",
    "966119": "This can happen if you train each fold a different number of epochs which happens when we save the best epoch's fold model for each fold. One solution is min-max normalization. Take the predictions for each fold and scale them between 0 and 1.\n\nIf predictions for fold 1 are in variable `preds` then do this \n\n    new_preds = preds - np.min(preds)\n    new_preds = new_preds / np.max(new_preds)\n  \nSo in my popular notebook, find this code\n\n    pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \n    oof_pred.append( np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1) )\n\nAnd change to this code\n\n    pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:TTA*ct_valid,] \n    pred = np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1)\n    new_preds = preds - np.min(preds)\n    new_preds = new_preds / np.max(new_preds)\n    oof_pred.append( new_preds )\n\nThen when the 5 folds get combined, they will match up better. AUC is a ranking metric, so if the fold predictions are not aligned, you can observe what you observed.",
    "966120": "Maybe your score on different folds has different scale? For example if on the first fold score has mean -1 and on the second score has mean +50, then before concatenation you should rescale you scores on different folds to lead them to the same scale with mean 0.",
    "966122": "Yes, exactly Sxwat. I posted code below showing how to rescale each fold.",
    "966125": "But note that it isn't so important to fix your problem. You can use the average of the 5 folds or the scaled merging of the 5 folds. Try both but the most important thing is to have a CV score (whether it be the 5 average or 5 merge) that goes up and down when LB goes up and down.",
    "966195": "Appreciate much !"
  },
  "source": "meta"
}