{
  "id": 174058,
  "title": "Weights for different models - Ensemble",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174058",
  "author_name": "",
  "post_date": "2020-08-12T05:56:13.870483Z",
  "votes": 2,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Is there a good strategy to decide how much to weight different models all with good performance, only slight differences (val<em>loss/auc</em>score) when creating an ensemble. I read the different threads on this and suggestions such as rank and power ensembles. But those methods seem to weight all models the same (I may have missed something). </p>",
  "messages": [
    {
      "id": "967262",
      "postDate": "08/12/2020 05:56:13",
      "content": "<p>Is there a good strategy to decide how much to weight different models all with good performance, only slight differences (val<em>loss/auc</em>score) when creating an ensemble. I read the different threads on this and suggestions such as rank and power ensembles. But those methods seem to weight all models the same (I may have missed something). </p>",
      "rawMarkdown": "Is there a good strategy to decide how much to weight different models all with good performance, only slight differences (val_loss/auc_score) when creating an ensemble. I read the different threads on this and suggestions such as rank and power ensembles. But those methods seem to weight all models the same (I may have missed something).",
      "votes": null
    },
    {
      "id": "967346",
      "postDate": "08/12/2020 07:17:19",
      "content": "<p>you can select weights that maximize AUC on OOF predictions</p>",
      "rawMarkdown": "you can select weights that maximize AUC on OOF predictions",
      "votes": null
    },
    {
      "id": "967353",
      "postDate": "08/12/2020 07:25:09",
      "content": "<p>Is there someplace which shows how to do this. If you have 5 models and you record their AUC on OOF predictions as: [0.89,0.91,0.92,0.93,0.94]</p>\n<p>in this scenario how would you weight each model to heavily favor the last model, rather than a simple 1/5</p>",
      "rawMarkdown": "Is there someplace which shows how to do this. If you have 5 models and you record their AUC on OOF predictions as: [0.89,0.91,0.92,0.93,0.94]\n\nin this scenario how would you weight each model to heavily favor the last model, rather than a simple 1/5",
      "votes": null
    },
    {
      "id": "967368",
      "postDate": "08/12/2020 07:50:34",
      "content": "<p>I do not think the 0.89 one is \"with good perf\" except it brings something special.<br>\nBy intuition, I would suggest [0, 0.1, 0.2, 0.3, 0.4] or [0.05, 0.05, 0.2, 0.3, 0.4] if you wish to keep them all.</p>",
      "rawMarkdown": "I do not think the 0.89 one is \"with good perf\" except it brings something special.\nBy intuition, I would suggest [0, 0.1, 0.2, 0.3, 0.4] or [0.05, 0.05, 0.2, 0.3, 0.4] if you wish to keep them all.",
      "votes": null
    },
    {
      "id": "967371",
      "postDate": "08/12/2020 07:54:17",
      "content": "<p>Thanks. Looks like this is based on intuition, not a formula or algorithm. </p>",
      "rawMarkdown": "Thanks. Looks like this is based on intuition, not a formula or algorithm.",
      "votes": null
    },
    {
      "id": "967378",
      "postDate": "08/12/2020 08:03:04",
      "content": "<p>You can use scipy.optimize. It searches the parameters that maximize a custom function.</p>",
      "rawMarkdown": "You can use scipy.optimize. It searches the parameters that maximize a custom function.",
      "votes": null
    },
    {
      "id": "967400",
      "postDate": "08/12/2020 08:18:35",
      "content": "<p>Please, Can you expand OOF ?</p>",
      "rawMarkdown": "Please, Can you expand OOF ?",
      "votes": null
    },
    {
      "id": "967425",
      "postDate": "08/12/2020 08:43:12",
      "content": "<p>Out of Fold (OOF). If we use K-fold cross validation, then when training on every fold there is a holdout set (out of fold), and we can make predictions on that data-set. When we complete our K-fold training then every example in our data will have one OOF prediction. </p>",
      "rawMarkdown": "Out of Fold (OOF). If we use K-fold cross validation, then when training on every fold there is a holdout set (out of fold), and we can make predictions on that data-set. When we complete our K-fold training then every example in our data will have one OOF prediction.",
      "votes": null
    },
    {
      "id": "967433",
      "postDate": "08/12/2020 08:47:46",
      "content": "<p>Thank you very much.</p>",
      "rawMarkdown": "Thank you very much.",
      "votes": null
    },
    {
      "id": "967505",
      "postDate": "08/12/2020 09:50:04",
      "content": "<p>I made a notebook <a href=\"https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification\" target=\"_blank\">Simple OOF Ensembling methods for Classification</a> to show some simple techniques</p>",
      "rawMarkdown": "I made a notebook [Simple OOF Ensembling methods for Classification](https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification) to show some simple techniques",
      "votes": null
    },
    {
      "id": "968286",
      "postDate": "08/12/2020 20:39:46",
      "content": "<p><a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> thank you for sharing the notebook. Question: would it be more reasonable to build bayesian not on \"raw\" oof predictions but on oofs, calibrated by any of the methods you applied (rank, power or avg)? </p>",
      "rawMarkdown": "steubk thank you for sharing the notebook. Question: would it be more reasonable to build bayesian not on \"raw\" oof predictions but on oofs, calibrated by any of the methods you applied (rank, power or avg)?",
      "votes": null
    },
    {
      "id": "968362",
      "postDate": "08/12/2020 23:43:02",
      "content": "<p>I've made a notebook on this, <a href=\"https://www.kaggle.com/ipythonx/efficientnet-b6-oof-weights-finder-seed-42/notebook\" target=\"_blank\">OOF Weights Finder</a>, modeled on E6, the same validation scheme on four experiments. The <code>scipy.optimize</code> is used to find the optimized weights that will maximize the <code>roc_auc</code>. And finally, the coefficient will use in the prediction phase.  </p>\n<p><a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> has published some <a href=\"https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification\" target=\"_blank\">nice strategy</a> on this. However, one silly thing bugging me is that, in both approaches, the sum of the coefficient wasn't summed to 1, I expected though after giving long iteration. I think I'm missing something here. Anyone can please give some pointer on this?</p>",
      "rawMarkdown": "I've made a notebook on this, [OOF Weights Finder](https://www.kaggle.com/ipythonx/efficientnet-b6-oof-weights-finder-seed-42/notebook), modeled on E6, the same validation scheme on four experiments. The `scipy.optimize` is used to find the optimized weights that will maximize the `roc_auc`. And finally, the coefficient will use in the prediction phase.  \n\n@steubk has published some [nice strategy](https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification) on this. However, one silly thing bugging me is that, in both approaches, the sum of the coefficient wasn't summed to 1, I expected though after giving long iteration. I think I'm missing something here. Anyone can please give some pointer on this?",
      "votes": null
    },
    {
      "id": "968420",
      "postDate": "08/13/2020 02:08:25",
      "content": "<p>It's pretty simple. Gather all your OOF predictions in a dataset. Define a weight vector. Multiply the matrix of prediction with the vector, tada weighted predictions. Now put all that in a function of the weight vector and use an optimizer with objective maximize auc (or minimize -auc).</p>\n<p>Alternatively, use a meta model and train it.</p>",
      "rawMarkdown": "It's pretty simple. Gather all your OOF predictions in a dataset. Define a weight vector. Multiply the matrix of prediction with the vector, tada weighted predictions. Now put all that in a function of the weight vector and use an optimizer with objective maximize auc (or minimize -auc).\n\nAlternatively, use a meta model and train it.",
      "votes": null
    },
    {
      "id": "968623",
      "postDate": "08/13/2020 06:29:29",
      "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> As usual, there is not a definitive rule. You should experiment as much as possible.<br>\nAnyway for the 3 models in notebook I also tried weighted rank and weighted power and  bayesian optimization for weighted rank  give the best OOF result:<br>\n<code>x = c0*df[  models[0] ].rank()/df[  models[0] ].rank().max() + c1*df[ models[1]].rank()/df[  models[2] ].rank().max() + c2*df[ models[2]].rank()/df[  models[2] ].rank().max()</code><br>\nauc weighted rank: 0.9322<br>\nauc weighted avg: 0.9316<br>\nauc rank: 0.9312<br>\nauc weighted power: 0.9307<br>\nauc avg: 0.9296<br>\nauc power: 0.9277</p>",
      "rawMarkdown": "@dunklerwald As usual, there is not a definitive rule. You should experiment as much as possible.\n\n\nAnyway for the 3 models in notebook I also tried weighted rank and weighted power and  bayesian optimization for weighted rank  give the best OOF result:\n\n\n`  x = c0*df[  models[0] ].rank()/df[  models[0] ].rank().max() + c1*df[ models[1]].rank()/df[  models[2] ].rank().max() + c2*df[ models[2]].rank()/df[  models[2] ].rank().max()`\n\n\nauc weighted rank: 0.9322\nauc weighted avg: 0.9316\nauc rank: 0.9312\nauc weighted power: 0.9307\nauc avg: 0.9296\nauc power: 0.9277",
      "votes": null
    },
    {
      "id": "968717",
      "postDate": "08/13/2020 07:46:17",
      "content": "<p><a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> thanks a lot!</p>",
      "rawMarkdown": "steubk thanks a lot!",
      "votes": null
    },
    {
      "id": "969191",
      "postDate": "08/13/2020 14:32:12",
      "content": "<p>In one of my experiment, we've found the following CV boost and corresponding LB scores</p>\n<pre><code>- BayesianOptimization: CV: 93215, LB: 94.90\n- ScipyOptimize: CV: 0.93201, LB: 94.91\n</code></pre>",
      "rawMarkdown": "In one of my experiment, we've found the following CV boost and corresponding LB scores\n\n```\n- BayesianOptimization: CV: 93215, LB: 94.90\n- ScipyOptimize: CV: 0.93201, LB: 94.91\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 967346,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "08/12/2020 07:17:19",
      "content": "<p>you can select weights that maximize AUC on OOF predictions</p>",
      "votes": null,
      "replies": [
        {
          "id": 967353,
          "author_name": "sebastianji",
          "author_url": "",
          "post_date": "08/12/2020 07:25:09",
          "content": "<p>Is there someplace which shows how to do this. If you have 5 models and you record their AUC on OOF predictions as: [0.89,0.91,0.92,0.93,0.94]</p>\n<p>in this scenario how would you weight each model to heavily favor the last model, rather than a simple 1/5</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 967368,
          "author_name": "vicioussong",
          "author_url": "",
          "post_date": "08/12/2020 07:50:34",
          "content": "<p>I do not think the 0.89 one is \"with good perf\" except it brings something special.<br>\nBy intuition, I would suggest [0, 0.1, 0.2, 0.3, 0.4] or [0.05, 0.05, 0.2, 0.3, 0.4] if you wish to keep them all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 967371,
          "author_name": "sebastianji",
          "author_url": "",
          "post_date": "08/12/2020 07:54:17",
          "content": "<p>Thanks. Looks like this is based on intuition, not a formula or algorithm. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 967378,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/12/2020 08:03:04",
          "content": "<p>You can use scipy.optimize. It searches the parameters that maximize a custom function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 967400,
          "author_name": "khanfashee",
          "author_url": "",
          "post_date": "08/12/2020 08:18:35",
          "content": "<p>Please, Can you expand OOF ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 967425,
          "author_name": "sebastianji",
          "author_url": "",
          "post_date": "08/12/2020 08:43:12",
          "content": "<p>Out of Fold (OOF). If we use K-fold cross validation, then when training on every fold there is a holdout set (out of fold), and we can make predictions on that data-set. When we complete our K-fold training then every example in our data will have one OOF prediction. </p>",
          "votes": null,
          "replies": [
            {
              "id": 967433,
              "author_name": "khanfashee",
              "author_url": "",
              "post_date": "08/12/2020 08:47:46",
              "content": "<p>Thank you very much.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 968420,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "08/13/2020 02:08:25",
          "content": "<p>It's pretty simple. Gather all your OOF predictions in a dataset. Define a weight vector. Multiply the matrix of prediction with the vector, tada weighted predictions. Now put all that in a function of the weight vector and use an optimizer with objective maximize auc (or minimize -auc).</p>\n<p>Alternatively, use a meta model and train it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 967505,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "08/12/2020 09:50:04",
      "content": "<p>I made a notebook <a href=\"https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification\" target=\"_blank\">Simple OOF Ensembling methods for Classification</a> to show some simple techniques</p>",
      "votes": null,
      "replies": [
        {
          "id": 968286,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/12/2020 20:39:46",
          "content": "<p><a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> thank you for sharing the notebook. Question: would it be more reasonable to build bayesian not on \"raw\" oof predictions but on oofs, calibrated by any of the methods you applied (rank, power or avg)? </p>",
          "votes": null,
          "replies": [
            {
              "id": 968623,
              "author_name": "steubk",
              "author_url": "",
              "post_date": "08/13/2020 06:29:29",
              "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> As usual, there is not a definitive rule. You should experiment as much as possible.<br>\nAnyway for the 3 models in notebook I also tried weighted rank and weighted power and  bayesian optimization for weighted rank  give the best OOF result:<br>\n<code>x = c0*df[  models[0] ].rank()/df[  models[0] ].rank().max() + c1*df[ models[1]].rank()/df[  models[2] ].rank().max() + c2*df[ models[2]].rank()/df[  models[2] ].rank().max()</code><br>\nauc weighted rank: 0.9322<br>\nauc weighted avg: 0.9316<br>\nauc rank: 0.9312<br>\nauc weighted power: 0.9307<br>\nauc avg: 0.9296<br>\nauc power: 0.9277</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 968717,
              "author_name": "dunklerwald",
              "author_url": "",
              "post_date": "08/13/2020 07:46:17",
              "content": "<p><a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> thanks a lot!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 968362,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "08/12/2020 23:43:02",
      "content": "<p>I've made a notebook on this, <a href=\"https://www.kaggle.com/ipythonx/efficientnet-b6-oof-weights-finder-seed-42/notebook\" target=\"_blank\">OOF Weights Finder</a>, modeled on E6, the same validation scheme on four experiments. The <code>scipy.optimize</code> is used to find the optimized weights that will maximize the <code>roc_auc</code>. And finally, the coefficient will use in the prediction phase.  </p>\n<p><a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> has published some <a href=\"https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification\" target=\"_blank\">nice strategy</a> on this. However, one silly thing bugging me is that, in both approaches, the sum of the coefficient wasn't summed to 1, I expected though after giving long iteration. I think I'm missing something here. Anyone can please give some pointer on this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 969191,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "08/13/2020 14:32:12",
      "content": "<p>In one of my experiment, we've found the following CV boost and corresponding LB scores</p>\n<pre><code>- BayesianOptimization: CV: 93215, LB: 94.90\n- ScipyOptimize: CV: 0.93201, LB: 94.91\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "967262": "Is there a good strategy to decide how much to weight different models all with good performance, only slight differences (val_loss/auc_score) when creating an ensemble. I read the different threads on this and suggestions such as rank and power ensembles. But those methods seem to weight all models the same (I may have missed something).",
    "967346": "you can select weights that maximize AUC on OOF predictions",
    "967353": "Is there someplace which shows how to do this. If you have 5 models and you record their AUC on OOF predictions as: [0.89,0.91,0.92,0.93,0.94]\n\nin this scenario how would you weight each model to heavily favor the last model, rather than a simple 1/5",
    "967368": "I do not think the 0.89 one is \"with good perf\" except it brings something special.\nBy intuition, I would suggest [0, 0.1, 0.2, 0.3, 0.4] or [0.05, 0.05, 0.2, 0.3, 0.4] if you wish to keep them all.",
    "967371": "Thanks. Looks like this is based on intuition, not a formula or algorithm.",
    "967378": "You can use scipy.optimize. It searches the parameters that maximize a custom function.",
    "967400": "Please, Can you expand OOF ?",
    "967425": "Out of Fold (OOF). If we use K-fold cross validation, then when training on every fold there is a holdout set (out of fold), and we can make predictions on that data-set. When we complete our K-fold training then every example in our data will have one OOF prediction.",
    "967433": "Thank you very much.",
    "967505": "I made a notebook [Simple OOF Ensembling methods for Classification](https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification) to show some simple techniques",
    "968286": "steubk thank you for sharing the notebook. Question: would it be more reasonable to build bayesian not on \"raw\" oof predictions but on oofs, calibrated by any of the methods you applied (rank, power or avg)?",
    "968362": "I've made a notebook on this, [OOF Weights Finder](https://www.kaggle.com/ipythonx/efficientnet-b6-oof-weights-finder-seed-42/notebook), modeled on E6, the same validation scheme on four experiments. The `scipy.optimize` is used to find the optimized weights that will maximize the `roc_auc`. And finally, the coefficient will use in the prediction phase.  \n\n@steubk has published some [nice strategy](https://www.kaggle.com/steubk/simple-oof-ensembling-methods-for-classification) on this. However, one silly thing bugging me is that, in both approaches, the sum of the coefficient wasn't summed to 1, I expected though after giving long iteration. I think I'm missing something here. Anyone can please give some pointer on this?",
    "968420": "It's pretty simple. Gather all your OOF predictions in a dataset. Define a weight vector. Multiply the matrix of prediction with the vector, tada weighted predictions. Now put all that in a function of the weight vector and use an optimizer with objective maximize auc (or minimize -auc).\n\nAlternatively, use a meta model and train it.",
    "968623": "@dunklerwald As usual, there is not a definitive rule. You should experiment as much as possible.\n\n\nAnyway for the 3 models in notebook I also tried weighted rank and weighted power and  bayesian optimization for weighted rank  give the best OOF result:\n\n\n`  x = c0*df[  models[0] ].rank()/df[  models[0] ].rank().max() + c1*df[ models[1]].rank()/df[  models[2] ].rank().max() + c2*df[ models[2]].rank()/df[  models[2] ].rank().max()`\n\n\nauc weighted rank: 0.9322\nauc weighted avg: 0.9316\nauc rank: 0.9312\nauc weighted power: 0.9307\nauc avg: 0.9296\nauc power: 0.9277",
    "968717": "steubk thanks a lot!",
    "969191": "In one of my experiment, we've found the following CV boost and corresponding LB scores\n\n```\n- BayesianOptimization: CV: 93215, LB: 94.90\n- ScipyOptimize: CV: 0.93201, LB: 94.91\n```"
  },
  "source": "meta"
}