{
  "id": 170494,
  "title": "LGB/MLP/RIDGE STACKING Effnet OOF with META",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170494",
  "author_name": "",
  "post_date": "2020-07-27T23:48:24.728425400Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>i created an extension notebook based on my previous work:\n<a href=\"https://www.kaggle.com/chen2222/ridge-lgb-nn-on-meta-data-optuna-focal-loss\">https://www.kaggle.com/chen2222/ridge-lgb-nn-on-meta-data-optuna-focal-loss</a></p>\n\n<p>The new notebook is here: \n<a href=\"https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta\">https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta</a></p>\n\n<p>This notebook will show you how to stack your neural network out of sample (oof) outputs with meta features. Three models will be applied to the stacking data and you can combine them at the end: Ridge, Multilayer Perceptron and LightGBM. You can compare this approach with directly using meta in your neural network, or even blend both approaches to get a more robust model. Be aware of data leakage.</p>\n\n<p>Input files: out of sample and test predictions are obtained using Chris Deotte's kernel <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords</a> ** The files and parameters in this notebook are only for demonstration purposes. You may want to use your own.**</p>\n\n<p>You can certainly apply more trails to find better parameters or experiment with target mean encoding, binning your numeric outputs etc. Please be aware of overfitting. This notebook gives the opportunities to explore these options. It also allows you to explore different imputing strategies: mean, constant etc, as well as different binning strategies, kmeans, uniform, quantile etc. This allows you to build different models based on different datasets to improve robustness.</p>\n\n<p>If you want to further tune LightGBM, you can use this kernel: \n<a href=\"https://www.kaggle.com/chen2222/lightgbm-tuning-step-by-step-optuna-0-122-lb\">https://www.kaggle.com/chen2222/lightgbm-tuning-step-by-step-optuna-0-122-lb</a></p>\n\n<p>If you want to check correlation between your oof files and eliminate some features, you can use this kernel: \n<a href=\"https://www.kaggle.com/chen2222/feature-selection-mdi-perm-rfe-in-depth-review\">https://www.kaggle.com/chen2222/feature-selection-mdi-perm-rfe-in-depth-review</a></p>\n\n<p>If you find any bugs, please let me know. Please upvote if you find these notebooks helpful and I really appreciate your support.</p>",
  "messages": [
    {
      "id": "948427",
      "postDate": "07/27/2020 23:48:24",
      "content": "<p>i created an extension notebook based on my previous work:\n<a href=\"https://www.kaggle.com/chen2222/ridge-lgb-nn-on-meta-data-optuna-focal-loss\">https://www.kaggle.com/chen2222/ridge-lgb-nn-on-meta-data-optuna-focal-loss</a></p>\n\n<p>The new notebook is here: \n<a href=\"https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta\">https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta</a></p>\n\n<p>This notebook will show you how to stack your neural network out of sample (oof) outputs with meta features. Three models will be applied to the stacking data and you can combine them at the end: Ridge, Multilayer Perceptron and LightGBM. You can compare this approach with directly using meta in your neural network, or even blend both approaches to get a more robust model. Be aware of data leakage.</p>\n\n<p>Input files: out of sample and test predictions are obtained using Chris Deotte's kernel <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords</a> ** The files and parameters in this notebook are only for demonstration purposes. You may want to use your own.**</p>\n\n<p>You can certainly apply more trails to find better parameters or experiment with target mean encoding, binning your numeric outputs etc. Please be aware of overfitting. This notebook gives the opportunities to explore these options. It also allows you to explore different imputing strategies: mean, constant etc, as well as different binning strategies, kmeans, uniform, quantile etc. This allows you to build different models based on different datasets to improve robustness.</p>\n\n<p>If you want to further tune LightGBM, you can use this kernel: \n<a href=\"https://www.kaggle.com/chen2222/lightgbm-tuning-step-by-step-optuna-0-122-lb\">https://www.kaggle.com/chen2222/lightgbm-tuning-step-by-step-optuna-0-122-lb</a></p>\n\n<p>If you want to check correlation between your oof files and eliminate some features, you can use this kernel: \n<a href=\"https://www.kaggle.com/chen2222/feature-selection-mdi-perm-rfe-in-depth-review\">https://www.kaggle.com/chen2222/feature-selection-mdi-perm-rfe-in-depth-review</a></p>\n\n<p>If you find any bugs, please let me know. Please upvote if you find these notebooks helpful and I really appreciate your support.</p>",
      "rawMarkdown": "i created an extension notebook based on my previous work:\nhttps://www.kaggle.com/chen2222/ridge-lgb-nn-on-meta-data-optuna-focal-loss\n\nThe new notebook is here: \nhttps://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta\n\nThis notebook will show you how to stack your neural network out of sample (oof) outputs with meta features. Three models will be applied to the stacking data and you can combine them at the end: Ridge, Multilayer Perceptron and LightGBM. You can compare this approach with directly using meta in your neural network, or even blend both approaches to get a more robust model. Be aware of data leakage.\n\nInput files: out of sample and test predictions are obtained using Chris Deotte's kernel https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords ** The files and parameters in this notebook are only for demonstration purposes. You may want to use your own.**\n\nYou can certainly apply more trails to find better parameters or experiment with target mean encoding, binning your numeric outputs etc. Please be aware of overfitting. This notebook gives the opportunities to explore these options. It also allows you to explore different imputing strategies: mean, constant etc, as well as different binning strategies, kmeans, uniform, quantile etc. This allows you to build different models based on different datasets to improve robustness.\n\nIf you want to further tune LightGBM, you can use this kernel: \nhttps://www.kaggle.com/chen2222/lightgbm-tuning-step-by-step-optuna-0-122-lb\n\nIf you want to check correlation between your oof files and eliminate some features, you can use this kernel: \nhttps://www.kaggle.com/chen2222/feature-selection-mdi-perm-rfe-in-depth-review\n\nIf you find any bugs, please let me know. Please upvote if you find these notebooks helpful and I really appreciate your support.",
      "votes": null
    },
    {
      "id": "948486",
      "postDate": "07/28/2020 02:17:33",
      "content": "<p>Thanks for your work!</p>",
      "rawMarkdown": "Thanks for your work!",
      "votes": null
    },
    {
      "id": "948488",
      "postDate": "07/28/2020 02:22:00",
      "content": "<p>Please upvote notebook if you found them helpful.</p>",
      "rawMarkdown": "Please upvote notebook if you found them helpful.",
      "votes": null
    },
    {
      "id": "948532",
      "postDate": "07/28/2020 03:34:48",
      "content": "<p>What is the best score(CV, LB) you could generate through stacking? I have a similar approach but I use <code>sklearn.ensemble.RandomForestClassifier</code> with no hyper-parameter optimisation and my LB is my PB.</p>",
      "rawMarkdown": "What is the best score(CV, LB) you could generate through stacking? I have a similar approach but I use `sklearn.ensemble.RandomForestClassifier` with no hyper-parameter optimisation and my LB is my PB.",
      "votes": null
    },
    {
      "id": "948547",
      "postDate": "07/28/2020 04:00:40",
      "content": "<p>LB is higher than my cv in most cases. I couldn't find a reliable strategy so far. But you could certainly try to incorporate in more models to see whether it can improve your performance, and yes, optimizing parameters is important you can check my kernel see the change.  </p>",
      "rawMarkdown": "LB is higher than my cv in most cases. I couldn't find a reliable strategy so far. But you could certainly try to incorporate in more models to see whether it can improve your performance, and yes, optimizing parameters is important you can check my kernel see the change.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 948486,
      "author_name": "vincentrenzw",
      "author_url": "",
      "post_date": "07/28/2020 02:17:33",
      "content": "<p>Thanks for your work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 948488,
          "author_name": "chen2222",
          "author_url": "",
          "post_date": "07/28/2020 02:22:00",
          "content": "<p>Please upvote notebook if you found them helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 948532,
      "author_name": "spoon69",
      "author_url": "",
      "post_date": "07/28/2020 03:34:48",
      "content": "<p>What is the best score(CV, LB) you could generate through stacking? I have a similar approach but I use <code>sklearn.ensemble.RandomForestClassifier</code> with no hyper-parameter optimisation and my LB is my PB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 948547,
          "author_name": "chen2222",
          "author_url": "",
          "post_date": "07/28/2020 04:00:40",
          "content": "<p>LB is higher than my cv in most cases. I couldn't find a reliable strategy so far. But you could certainly try to incorporate in more models to see whether it can improve your performance, and yes, optimizing parameters is important you can check my kernel see the change.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "948427": "i created an extension notebook based on my previous work:\nhttps://www.kaggle.com/chen2222/ridge-lgb-nn-on-meta-data-optuna-focal-loss\n\nThe new notebook is here: \nhttps://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta\n\nThis notebook will show you how to stack your neural network out of sample (oof) outputs with meta features. Three models will be applied to the stacking data and you can combine them at the end: Ridge, Multilayer Perceptron and LightGBM. You can compare this approach with directly using meta in your neural network, or even blend both approaches to get a more robust model. Be aware of data leakage.\n\nInput files: out of sample and test predictions are obtained using Chris Deotte's kernel https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords ** The files and parameters in this notebook are only for demonstration purposes. You may want to use your own.**\n\nYou can certainly apply more trails to find better parameters or experiment with target mean encoding, binning your numeric outputs etc. Please be aware of overfitting. This notebook gives the opportunities to explore these options. It also allows you to explore different imputing strategies: mean, constant etc, as well as different binning strategies, kmeans, uniform, quantile etc. This allows you to build different models based on different datasets to improve robustness.\n\nIf you want to further tune LightGBM, you can use this kernel: \nhttps://www.kaggle.com/chen2222/lightgbm-tuning-step-by-step-optuna-0-122-lb\n\nIf you want to check correlation between your oof files and eliminate some features, you can use this kernel: \nhttps://www.kaggle.com/chen2222/feature-selection-mdi-perm-rfe-in-depth-review\n\nIf you find any bugs, please let me know. Please upvote if you find these notebooks helpful and I really appreciate your support.",
    "948486": "Thanks for your work!",
    "948488": "Please upvote notebook if you found them helpful.",
    "948532": "What is the best score(CV, LB) you could generate through stacking? I have a similar approach but I use `sklearn.ensemble.RandomForestClassifier` with no hyper-parameter optimisation and my LB is my PB.",
    "948547": "LB is higher than my cv in most cases. I couldn't find a reliable strategy so far. But you could certainly try to incorporate in more models to see whether it can improve your performance, and yes, optimizing parameters is important you can check my kernel see the change."
  },
  "source": "meta"
}