{
  "id": 346563,
  "title": "Help! Stacking OOF gone wrong?",
  "url": "/competitions/amex-default-prediction/discussion/346563",
  "author_name": "",
  "post_date": "2022-08-20T08:15:32.836769800Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>So I have a bunch of OOF preds from models that I've built, all using the same exact fold indices. And I tried to build a stacked model on top of them. The results are completely confounding me </p>\n<p>Fold 1 Kaggle Metric = 0.8004509963707986 </p>\n<p>Fold 2 Kaggle Metric = 0.7978457968168704 </p>\n<p>Fold 3 Kaggle Metric = 0.8022105945948493 </p>\n<p>Fold 4 Kaggle Metric = 0.7963073864926411 </p>\n<p>Fold 5 Kaggle Metric = 0.7976026103183025 </p>\n<p><strong>OVERALL CV Kaggle Metric = 0.7616026728560368</strong></p>\n<p>Each fold validation metric seems high but the overall OOF metric seems obscenely low!</p>\n<p>Despite this if I run inference on the test set, I get an <strong>LB score of 0.798</strong></p>\n<p>No idea what is going on. Appreciate any feedback on this!</p>",
  "messages": [
    {
      "id": "1906824",
      "postDate": "08/20/2022 08:15:32",
      "content": "<p>So I have a bunch of OOF preds from models that I've built, all using the same exact fold indices. And I tried to build a stacked model on top of them. The results are completely confounding me </p>\n<p>Fold 1 Kaggle Metric = 0.8004509963707986 </p>\n<p>Fold 2 Kaggle Metric = 0.7978457968168704 </p>\n<p>Fold 3 Kaggle Metric = 0.8022105945948493 </p>\n<p>Fold 4 Kaggle Metric = 0.7963073864926411 </p>\n<p>Fold 5 Kaggle Metric = 0.7976026103183025 </p>\n<p><strong>OVERALL CV Kaggle Metric = 0.7616026728560368</strong></p>\n<p>Each fold validation metric seems high but the overall OOF metric seems obscenely low!</p>\n<p>Despite this if I run inference on the test set, I get an <strong>LB score of 0.798</strong></p>\n<p>No idea what is going on. Appreciate any feedback on this!</p>",
      "rawMarkdown": "So I have a bunch of OOF preds from models that I've built, all using the same exact fold indices. And I tried to build a stacked model on top of them. The results are completely confounding me \n\nFold 1 Kaggle Metric = 0.8004509963707986 \n\nFold 2 Kaggle Metric = 0.7978457968168704 \n\nFold 3 Kaggle Metric = 0.8022105945948493 \n\nFold 4 Kaggle Metric = 0.7963073864926411 \n\nFold 5 Kaggle Metric = 0.7976026103183025 \n\n**OVERALL CV Kaggle Metric = 0.7616026728560368**\n\nEach fold validation metric seems high but the overall OOF metric seems obscenely low!\n\nDespite this if I run inference on the test set, I get an **LB score of 0.798**\n\nNo idea what is going on. Appreciate any feedback on this!",
      "votes": null
    },
    {
      "id": "1906846",
      "postDate": "08/20/2022 08:47:35",
      "content": "<p>you should plot probability distribution of each fold. You might notice something interesting</p>",
      "rawMarkdown": "you should plot probability distribution of each fold. You might notice something interesting",
      "votes": null
    },
    {
      "id": "1906905",
      "postDate": "08/20/2022 10:20:13",
      "content": "<p>I will let the screenshots and results tell the story. Thank you for the insight <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> !!</p>\n<p>Fold Predictions (Before)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fb12a72c4dba7c0f2f299460652f35482%2FFoldpreds_before.PNG?generation=1660990518784433&amp;alt=media\" alt=\"\"></p>\n<p>OOF (Before)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fd9ecde46ffac56a04097da75519d1816%2FOOFpreds_before.PNG?generation=1660990587869710&amp;alt=media\" alt=\"\"></p>\n<p>Applied the following fix, to normalize each of the fold predictions</p>\n<p><code>oof_preds = (oof_preds - np.min(oof_preds)) / (np.max(oof_preds) - np.min(oof_preds))</code></p>\n<p>Fold Predictions (After)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F27202d33798a71c97b5d6921f3e11308%2FFoldpreds_after.PNG?generation=1660990697986344&amp;alt=media\" alt=\"\"></p>\n<p>OOF (After)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F8923a59dbfda186a7f967226e8b355fa%2FOOFpreds_after.PNG?generation=1660990712625072&amp;alt=media\" alt=\"\"></p>\n<h1>New results</h1>\n<p>Fold 1 Kaggle Metric = 0.8004509963707986 </p>\n<p>Fold 2 Kaggle Metric = 0.7978457968168704 </p>\n<p>Fold 3 Kaggle Metric = 0.8022105945948493 </p>\n<p>Fold 4 Kaggle Metric = 0.7963073864926411 </p>\n<p>Fold 5 Kaggle Metric = 0.7976026103183025 </p>\n<p>OVERALL CV Kaggle Metric = 0.7988348356023588</p>",
      "rawMarkdown": "I will let the screenshots and results tell the story. Thank you for the insight @raddar !!\n\nFold Predictions (Before)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fb12a72c4dba7c0f2f299460652f35482%2FFoldpreds_before.PNG?generation=1660990518784433&alt=media)\n\n\nOOF (Before)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fd9ecde46ffac56a04097da75519d1816%2FOOFpreds_before.PNG?generation=1660990587869710&alt=media)\n\n\nApplied the following fix, to normalize each of the fold predictions\n\n`oof_preds = (oof_preds - np.min(oof_preds)) / (np.max(oof_preds) - np.min(oof_preds))`\n\n\nFold Predictions (After)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F27202d33798a71c97b5d6921f3e11308%2FFoldpreds_after.PNG?generation=1660990697986344&alt=media)\n\n\nOOF (After)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F8923a59dbfda186a7f967226e8b355fa%2FOOFpreds_after.PNG?generation=1660990712625072&alt=media)\n\n\n# New results\n\nFold 1 Kaggle Metric = 0.8004509963707986 \n\nFold 2 Kaggle Metric = 0.7978457968168704 \n\nFold 3 Kaggle Metric = 0.8022105945948493 \n\nFold 4 Kaggle Metric = 0.7963073864926411 \n\nFold 5 Kaggle Metric = 0.7976026103183025 \n\nOVERALL CV Kaggle Metric = 0.7988348356023588",
      "votes": null
    },
    {
      "id": "1907531",
      "postDate": "08/20/2022 21:17:54",
      "content": "<p>good strategy is to use scipy's rankdata on each fold. min/max still leaves room for errors - rankdata is perfect for competition metric</p>",
      "rawMarkdown": "good strategy is to use scipy's rankdata on each fold. min/max still leaves room for errors - rankdata is perfect for competition metric",
      "votes": null
    },
    {
      "id": "1907876",
      "postDate": "08/21/2022 06:58:12",
      "content": "<p>Thanks for the tip! I will use that instead.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Thanks for the tip! I will use that instead.\n\nGood luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1906846,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "08/20/2022 08:47:35",
      "content": "<p>you should plot probability distribution of each fold. You might notice something interesting</p>",
      "votes": null,
      "replies": [
        {
          "id": 1906905,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "08/20/2022 10:20:13",
          "content": "<p>I will let the screenshots and results tell the story. Thank you for the insight <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> !!</p>\n<p>Fold Predictions (Before)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fb12a72c4dba7c0f2f299460652f35482%2FFoldpreds_before.PNG?generation=1660990518784433&amp;alt=media\" alt=\"\"></p>\n<p>OOF (Before)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fd9ecde46ffac56a04097da75519d1816%2FOOFpreds_before.PNG?generation=1660990587869710&amp;alt=media\" alt=\"\"></p>\n<p>Applied the following fix, to normalize each of the fold predictions</p>\n<p><code>oof_preds = (oof_preds - np.min(oof_preds)) / (np.max(oof_preds) - np.min(oof_preds))</code></p>\n<p>Fold Predictions (After)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F27202d33798a71c97b5d6921f3e11308%2FFoldpreds_after.PNG?generation=1660990697986344&amp;alt=media\" alt=\"\"></p>\n<p>OOF (After)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F8923a59dbfda186a7f967226e8b355fa%2FOOFpreds_after.PNG?generation=1660990712625072&amp;alt=media\" alt=\"\"></p>\n<h1>New results</h1>\n<p>Fold 1 Kaggle Metric = 0.8004509963707986 </p>\n<p>Fold 2 Kaggle Metric = 0.7978457968168704 </p>\n<p>Fold 3 Kaggle Metric = 0.8022105945948493 </p>\n<p>Fold 4 Kaggle Metric = 0.7963073864926411 </p>\n<p>Fold 5 Kaggle Metric = 0.7976026103183025 </p>\n<p>OVERALL CV Kaggle Metric = 0.7988348356023588</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1907531,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "08/20/2022 21:17:54",
          "content": "<p>good strategy is to use scipy's rankdata on each fold. min/max still leaves room for errors - rankdata is perfect for competition metric</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1907876,
          "author_name": "illidan7",
          "author_url": "",
          "post_date": "08/21/2022 06:58:12",
          "content": "<p>Thanks for the tip! I will use that instead.</p>\n<p>Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1906824": "So I have a bunch of OOF preds from models that I've built, all using the same exact fold indices. And I tried to build a stacked model on top of them. The results are completely confounding me \n\nFold 1 Kaggle Metric = 0.8004509963707986 \n\nFold 2 Kaggle Metric = 0.7978457968168704 \n\nFold 3 Kaggle Metric = 0.8022105945948493 \n\nFold 4 Kaggle Metric = 0.7963073864926411 \n\nFold 5 Kaggle Metric = 0.7976026103183025 \n\n**OVERALL CV Kaggle Metric = 0.7616026728560368**\n\nEach fold validation metric seems high but the overall OOF metric seems obscenely low!\n\nDespite this if I run inference on the test set, I get an **LB score of 0.798**\n\nNo idea what is going on. Appreciate any feedback on this!",
    "1906846": "you should plot probability distribution of each fold. You might notice something interesting",
    "1906905": "I will let the screenshots and results tell the story. Thank you for the insight @raddar !!\n\nFold Predictions (Before)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fb12a72c4dba7c0f2f299460652f35482%2FFoldpreds_before.PNG?generation=1660990518784433&alt=media)\n\n\nOOF (Before)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2Fd9ecde46ffac56a04097da75519d1816%2FOOFpreds_before.PNG?generation=1660990587869710&alt=media)\n\n\nApplied the following fix, to normalize each of the fold predictions\n\n`oof_preds = (oof_preds - np.min(oof_preds)) / (np.max(oof_preds) - np.min(oof_preds))`\n\n\nFold Predictions (After)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F27202d33798a71c97b5d6921f3e11308%2FFoldpreds_after.PNG?generation=1660990697986344&alt=media)\n\n\nOOF (After)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F471225%2F8923a59dbfda186a7f967226e8b355fa%2FOOFpreds_after.PNG?generation=1660990712625072&alt=media)\n\n\n# New results\n\nFold 1 Kaggle Metric = 0.8004509963707986 \n\nFold 2 Kaggle Metric = 0.7978457968168704 \n\nFold 3 Kaggle Metric = 0.8022105945948493 \n\nFold 4 Kaggle Metric = 0.7963073864926411 \n\nFold 5 Kaggle Metric = 0.7976026103183025 \n\nOVERALL CV Kaggle Metric = 0.7988348356023588",
    "1907531": "good strategy is to use scipy's rankdata on each fold. min/max still leaves room for errors - rankdata is perfect for competition metric",
    "1907876": "Thanks for the tip! I will use that instead.\n\nGood luck!"
  },
  "source": "meta"
}