{
  "id": 175781,
  "title": "Stacking on meta-data including Ensembling and Pseudo Labeling",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175781",
  "author_name": "",
  "post_date": "2020-08-19T11:56:59.919034100Z",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2F1dc91cf15788b749cb3e79ac3c55b09b%2FModel.png?generation=1597838067015231&amp;alt=media\" alt=\"\"></p>\n<h3>Strategy</h3>\n<ol>\n<li>Split the train set in k folds [we have used <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's TFrecords Id for splitting data]</li>\n<li>Fit a first-stage model on k-1 folds and predict the kth fold</li>\n<li>Repeat 2) to predict each fold</li>\n<li>We now have the (out-of-folds) prediction of the k folds</li>\n<li>Split these out-of folds predictions in p folds</li>\n<li>Fit a second stage (stacker) model on p-1 folds and predict the pth fold</li>\n<li>Repeat 6) to predict each fold</li>\n<li>The CV error of the second stage is calculated on each predicted fold<br>\n<a href=\"https://www.kaggle.com/general/18793\" target=\"_blank\">reference</a></li>\n</ol>\n<p>Check Notebook : <a href=\"https://www.kaggle.com/vatsalparsaniya/oof-pseudo-labeling-stacking-with-metadata\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "977293",
      "postDate": "08/19/2020 11:56:59",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2F1dc91cf15788b749cb3e79ac3c55b09b%2FModel.png?generation=1597838067015231&amp;alt=media\" alt=\"\"></p>\n<h3>Strategy</h3>\n<ol>\n<li>Split the train set in k folds [we have used <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's TFrecords Id for splitting data]</li>\n<li>Fit a first-stage model on k-1 folds and predict the kth fold</li>\n<li>Repeat 2) to predict each fold</li>\n<li>We now have the (out-of-folds) prediction of the k folds</li>\n<li>Split these out-of folds predictions in p folds</li>\n<li>Fit a second stage (stacker) model on p-1 folds and predict the pth fold</li>\n<li>Repeat 6) to predict each fold</li>\n<li>The CV error of the second stage is calculated on each predicted fold<br>\n<a href=\"https://www.kaggle.com/general/18793\" target=\"_blank\">reference</a></li>\n</ol>\n<p>Check Notebook : <a href=\"https://www.kaggle.com/vatsalparsaniya/oof-pseudo-labeling-stacking-with-metadata\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2F1dc91cf15788b749cb3e79ac3c55b09b%2FModel.png?generation=1597838067015231&alt=media)\n\n\n### Strategy\n1. Split the train set in k folds [we have used @cdeotte 's TFrecords Id for splitting data]\n2. Fit a first-stage model on k-1 folds and predict the kth fold\n3. Repeat 2) to predict each fold\n4. We now have the (out-of-folds) prediction of the k folds\n5. Split these out-of folds predictions in p folds\n6. Fit a second stage (stacker) model on p-1 folds and predict the pth fold\n7. Repeat 6) to predict each fold\n8. The CV error of the second stage is calculated on each predicted fold\n[reference](https://www.kaggle.com/general/18793)\n\nCheck Notebook : [here](https://www.kaggle.com/vatsalparsaniya/oof-pseudo-labeling-stacking-with-metadata)",
      "votes": null
    },
    {
      "id": "979618",
      "postDate": "08/21/2020 02:02:11",
      "content": "<p>Good job,thanks for sharing  </p>",
      "rawMarkdown": "Good job,thanks for sharing",
      "votes": null
    },
    {
      "id": "980702",
      "postDate": "08/21/2020 19:32:50",
      "content": "<p>Did stacking work for you?</p>",
      "rawMarkdown": "Did stacking work for you?",
      "votes": null
    },
    {
      "id": "980808",
      "postDate": "08/21/2020 22:11:43",
      "content": "<pre><code>clf1 = lgb.LGBMClassifier\nclf5 = RandomForestClassifier\nclf2 = LogisticRegression\nSCF = StackingClassifier(classifiers=[clf1, clf5], \n                         meta_classifier=clf2,\n                         use_probas=True,\n                         average_probas=True)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2Fdbbe02c7b9fe65fa773378d1f60928cf%2FCapture.png?generation=1598047671230499&amp;alt=media\" alt=\"\"></p>\n<p>Yes, StackingClassifier has given better ROC-Score than other Classifier for us.</p>",
      "rawMarkdown": "```\nclf1 = lgb.LGBMClassifier\nclf5 = RandomForestClassifier\nclf2 = LogisticRegression\nSCF = StackingClassifier(classifiers=[clf1, clf5], \n                         meta_classifier=clf2,\n                         use_probas=True,\n                         average_probas=True)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2Fdbbe02c7b9fe65fa773378d1f60928cf%2FCapture.png?generation=1598047671230499&alt=media)\n\nYes, StackingClassifier has given better ROC-Score than other Classifier for us.",
      "votes": null
    },
    {
      "id": "980822",
      "postDate": "08/21/2020 22:31:08",
      "content": "<p>I don't understand, your LB score is 0.9409. How does it relate to the above picture?</p>",
      "rawMarkdown": "I don't understand, your LB score is 0.9409. How does it relate to the above picture?",
      "votes": null
    },
    {
      "id": "980823",
      "postDate": "08/21/2020 22:31:45",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "980839",
      "postDate": "08/21/2020 23:08:31",
      "content": "<p>Image you are seeing above is showing CV score with Classifier. <br>\nwith StackingClassifier public LB: 0.9606 and  Private LB: 0.9406<br>\n<code>your LB score is 0.9409. How does it relate to the above picture?</code><br>\nPrivate Score 0.9409 which is weighted ensembling of StackingClassifier,  LGBM ,  RFC.</p>",
      "rawMarkdown": "Image you are seeing above is showing CV score with Classifier. \nwith StackingClassifier public LB: 0.9606 and  Private LB: 0.9406\n`your LB score is 0.9409. How does it relate to the above picture?`\nPrivate Score 0.9409 which is weighted ensembling of StackingClassifier,  LGBM ,  RFC.",
      "votes": null
    },
    {
      "id": "981247",
      "postDate": "08/22/2020 10:12:51",
      "content": "<p>Let me try again.  Do you get a higher private LB score with stacking?</p>\n<p>Answering with CV score or public LB score is not relevant.</p>\n<p>In my case I got better CV and public LB, but slightly worse private LB.  This is called overfitting and it means stacking did not help me.</p>",
      "rawMarkdown": "Let me try again.  Do you get a higher private LB score with stacking?\n\nAnswering with CV score or public LB score is not relevant.\n\nIn my case I got better CV and public LB, but slightly worse private LB.  This is called overfitting and it means stacking did not help me.",
      "votes": null
    },
    {
      "id": "983179",
      "postDate": "08/24/2020 05:16:52",
      "content": "<p>No, in that context stacking doesn't work.</p>\n<p>How stacking work for us,<br>\nwe wad finalized 3 high CV ensembling files before stacking, after applying stacking on combination of prediction and metadata it was giving us nice correlation between CV and Public LB (not higher Privet LB). so, we thought it works for us.</p>",
      "rawMarkdown": "No, in that context stacking doesn't work.\n\nHow stacking work for us,\nwe wad finalized 3 high CV ensembling files before stacking, after applying stacking on combination of prediction and metadata it was giving us nice correlation between CV and Public LB (not higher Privet LB). so, we thought it works for us.",
      "votes": null
    },
    {
      "id": "983524",
      "postDate": "08/24/2020 11:29:05",
      "content": "<p>Thanks for the answer, this is consistent with my experience here.</p>\n<blockquote>\n  <p>No, in that context stacking doesn't work.</p>\n</blockquote>\n<p>Do you know of any other definition of working?  Goal is to get best private LB, right?</p>",
      "rawMarkdown": "Thanks for the answer, this is consistent with my experience here.\n\n> No, in that context stacking doesn't work.\n\nDo you know of any other definition of working?  Goal is to get best private LB, right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 979618,
      "author_name": "",
      "author_url": "",
      "post_date": "08/21/2020 02:02:11",
      "content": "<p>Good job,thanks for sharing  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 980702,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/21/2020 19:32:50",
      "content": "<p>Did stacking work for you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 980808,
          "author_name": "vatsalparsaniya",
          "author_url": "",
          "post_date": "08/21/2020 22:11:43",
          "content": "<pre><code>clf1 = lgb.LGBMClassifier\nclf5 = RandomForestClassifier\nclf2 = LogisticRegression\nSCF = StackingClassifier(classifiers=[clf1, clf5], \n                         meta_classifier=clf2,\n                         use_probas=True,\n                         average_probas=True)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2Fdbbe02c7b9fe65fa773378d1f60928cf%2FCapture.png?generation=1598047671230499&amp;alt=media\" alt=\"\"></p>\n<p>Yes, StackingClassifier has given better ROC-Score than other Classifier for us.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980822,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/21/2020 22:31:08",
          "content": "<p>I don't understand, your LB score is 0.9409. How does it relate to the above picture?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980839,
          "author_name": "vatsalparsaniya",
          "author_url": "",
          "post_date": "08/21/2020 23:08:31",
          "content": "<p>Image you are seeing above is showing CV score with Classifier. <br>\nwith StackingClassifier public LB: 0.9606 and  Private LB: 0.9406<br>\n<code>your LB score is 0.9409. How does it relate to the above picture?</code><br>\nPrivate Score 0.9409 which is weighted ensembling of StackingClassifier,  LGBM ,  RFC.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 981247,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/22/2020 10:12:51",
          "content": "<p>Let me try again.  Do you get a higher private LB score with stacking?</p>\n<p>Answering with CV score or public LB score is not relevant.</p>\n<p>In my case I got better CV and public LB, but slightly worse private LB.  This is called overfitting and it means stacking did not help me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 983179,
          "author_name": "vatsalparsaniya",
          "author_url": "",
          "post_date": "08/24/2020 05:16:52",
          "content": "<p>No, in that context stacking doesn't work.</p>\n<p>How stacking work for us,<br>\nwe wad finalized 3 high CV ensembling files before stacking, after applying stacking on combination of prediction and metadata it was giving us nice correlation between CV and Public LB (not higher Privet LB). so, we thought it works for us.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 983524,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/24/2020 11:29:05",
          "content": "<p>Thanks for the answer, this is consistent with my experience here.</p>\n<blockquote>\n  <p>No, in that context stacking doesn't work.</p>\n</blockquote>\n<p>Do you know of any other definition of working?  Goal is to get best private LB, right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 980823,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "08/21/2020 22:31:45",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "977293": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2F1dc91cf15788b749cb3e79ac3c55b09b%2FModel.png?generation=1597838067015231&alt=media)\n\n\n### Strategy\n1. Split the train set in k folds [we have used @cdeotte 's TFrecords Id for splitting data]\n2. Fit a first-stage model on k-1 folds and predict the kth fold\n3. Repeat 2) to predict each fold\n4. We now have the (out-of-folds) prediction of the k folds\n5. Split these out-of folds predictions in p folds\n6. Fit a second stage (stacker) model on p-1 folds and predict the pth fold\n7. Repeat 6) to predict each fold\n8. The CV error of the second stage is calculated on each predicted fold\n[reference](https://www.kaggle.com/general/18793)\n\nCheck Notebook : [here](https://www.kaggle.com/vatsalparsaniya/oof-pseudo-labeling-stacking-with-metadata)",
    "979618": "Good job,thanks for sharing",
    "980702": "Did stacking work for you?",
    "980808": "```\nclf1 = lgb.LGBMClassifier\nclf5 = RandomForestClassifier\nclf2 = LogisticRegression\nSCF = StackingClassifier(classifiers=[clf1, clf5], \n                         meta_classifier=clf2,\n                         use_probas=True,\n                         average_probas=True)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2907842%2Fdbbe02c7b9fe65fa773378d1f60928cf%2FCapture.png?generation=1598047671230499&alt=media)\n\nYes, StackingClassifier has given better ROC-Score than other Classifier for us.",
    "980822": "I don't understand, your LB score is 0.9409. How does it relate to the above picture?",
    "980823": "Thanks for sharing",
    "980839": "Image you are seeing above is showing CV score with Classifier. \nwith StackingClassifier public LB: 0.9606 and  Private LB: 0.9406\n`your LB score is 0.9409. How does it relate to the above picture?`\nPrivate Score 0.9409 which is weighted ensembling of StackingClassifier,  LGBM ,  RFC.",
    "981247": "Let me try again.  Do you get a higher private LB score with stacking?\n\nAnswering with CV score or public LB score is not relevant.\n\nIn my case I got better CV and public LB, but slightly worse private LB.  This is called overfitting and it means stacking did not help me.",
    "983179": "No, in that context stacking doesn't work.\n\nHow stacking work for us,\nwe wad finalized 3 high CV ensembling files before stacking, after applying stacking on combination of prediction and metadata it was giving us nice correlation between CV and Public LB (not higher Privet LB). so, we thought it works for us.",
    "983524": "Thanks for the answer, this is consistent with my experience here.\n\n> No, in that context stacking doesn't work.\n\nDo you know of any other definition of working?  Goal is to get best private LB, right?"
  },
  "source": "meta"
}