{
  "id": 80979,
  "title": "Model ensembling for AUC metric",
  "url": "/competitions/histopathologic-cancer-detection/discussion/80979",
  "author_name": "Joni Juvonen",
  "post_date": "2019-02-18T11:51:20.086000",
  "votes": 3,
  "comment_count": 11,
  "views": 0,
  "content": "<p>What is your ensembling strategy when targeting the AUC metric?</p>\n\n<p>The area under the ROC curve (AUC) compares the rank of predictions to the rank of correct labels. When ensembling multiple models, more confident but not necessarily more accurate models have a larger impact.</p>\n\n<p>There are at least three different ensembling methods for AUC:</p>\n\n<ul>\n<li><strong>Simple average</strong> (weighted or uniform)</li>\n<li><strong>Rank average</strong> (average out the rankings of each model). Pros = confident models have equal impact.</li>\n<li><strong>Power average</strong> Power scaling aims to normalize the models: </li>\n</ul>\n\n<p><em>PowerAverage = (Pred1^p + Pred2^p + ... + Predn^p) / n</em></p>\n\n<hr>\n\n<p><strong>Relevant links:</strong></p>\n\n<p><a href=\"https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e\">AUC ensembling blog post</a></p>\n\n<p><a href=\"https://mlwave.com/kaggle-ensembling-guide/\">Kaggle ensembling guide</a></p>",
  "messages": [
    {
      "id": 473697,
      "postDate": "2019-02-18T11:51:20.087Z",
      "content": "<p>What is your ensembling strategy when targeting the AUC metric?</p>\n\n<p>The area under the ROC curve (AUC) compares the rank of predictions to the rank of correct labels. When ensembling multiple models, more confident but not necessarily more accurate models have a larger impact.</p>\n\n<p>There are at least three different ensembling methods for AUC:</p>\n\n<ul>\n<li><strong>Simple average</strong> (weighted or uniform)</li>\n<li><strong>Rank average</strong> (average out the rankings of each model). Pros = confident models have equal impact.</li>\n<li><strong>Power average</strong> Power scaling aims to normalize the models: </li>\n</ul>\n\n<p><em>PowerAverage = (Pred1^p + Pred2^p + ... + Predn^p) / n</em></p>\n\n<hr>\n\n<p><strong>Relevant links:</strong></p>\n\n<p><a href=\"https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e\">AUC ensembling blog post</a></p>\n\n<p><a href=\"https://mlwave.com/kaggle-ensembling-guide/\">Kaggle ensembling guide</a></p>",
      "rawMarkdown": "What is your ensembling strategy when targeting the AUC metric?\n\nThe area under the ROC curve (AUC) compares the rank of predictions to the rank of correct labels. When ensembling multiple models, more confident but not necessarily more accurate models have a larger impact.\n\nThere are at least three different ensembling methods for AUC:\n\n - **Simple average** (weighted or uniform)\n - **Rank average** (average out the rankings of each model). Pros = confident models have equal impact.\n - **Power average** Power scaling aims to normalize the models: \n \n*PowerAverage = (Pred1^p + Pred2^p + ... + Predn^p) / n*\n\n---------------------------------\n**Relevant links:**\n\n[AUC ensembling blog post](https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e)\n\n[Kaggle ensembling guide](https://mlwave.com/kaggle-ensembling-guide/)",
      "votes": 3
    },
    {
      "id": 474305,
      "postDate": "2019-02-19T07:29:51.037Z",
      "content": "<p>Okay, I made some tests.</p>\n\n<ul>\n<li>Ensembling 5-fold CV models and two different CNN architectures.</li>\n<li>Each single model scoring between LB 0.958 and LB 0.966.</li>\n<li>I used ranked weights in simple average and power average (best\nsingle model = 1 , second best = 0.5, third = 0.25, etc.) but models with LB ties (diff &lt; 0.1) had the same weights.</li>\n<li>All ensembled models correlated more or less (correlations between 0.875 and 0.975)</li>\n<li>Ranking correlations between 0.84 and 0.96</li>\n</ul>\n\n<h2>Results</h2>\n\n<ul>\n<li><strong>Simple average</strong> <code>LB 0.9709</code></li>\n<li><strong>Power of 2 average</strong> <code>LB 0.9702</code></li>\n<li><strong>Power of 4 average</strong> <code>LB 0.9695</code></li>\n<li><strong>Rank average</strong> <code>LB 0.9689</code></li>\n</ul>\n\n<h2>My conclusion</h2>\n\n<p>Using the most intuitive simple average seems to work the best here. At least when the models are correlated and similar weights are used. I wonder if anyone has different results or other ensembling methods to try.</p>",
      "rawMarkdown": "Okay, I made some tests.\n\n - Ensembling 5-fold CV models and two different CNN architectures.\n - Each single model scoring between LB 0.958 and LB 0.966.\n - I used ranked weights in simple average and power average (best\n   single model = 1 , second best = 0.5, third = 0.25, etc.) but models with LB ties (diff &lt; 0.1) had the same weights.\n - All ensembled models correlated more or less (correlations between 0.875 and 0.975)\n - Ranking correlations between 0.84 and 0.96\n\n## Results\n\n - **Simple average** `LB 0.9709`\n - **Power of 2 average** `LB 0.9702`\n - **Power of 4 average** `LB 0.9695`\n - **Rank average** `LB 0.9689`\n\n## My conclusion\n\nUsing the most intuitive simple average seems to work the best here. At least when the models are correlated and similar weights are used. I wonder if anyone has different results or other ensembling methods to try.",
      "votes": 4,
      "replies": [
        {
          "id": 479935,
          "postDate": "2019-02-27T15:38:19.313Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 480145,
          "postDate": "2019-02-27T20:42:04.120Z",
          "content": "<p>I used DenseNet169 similar to my <a href=\"https://www.kaggle.com/qitvision/a-complete-ml-pipeline-fast-ai\">public kernel</a> but trained the second cycle for longer (16 epochs) and used the full size (96x96) images. se_ResNext50 is also working well.</p>",
          "rawMarkdown": "I used DenseNet169 similar to my [public kernel](https://www.kaggle.com/qitvision/a-complete-ml-pipeline-fast-ai) but trained the second cycle for longer (16 epochs) and used the full size (96x96) images. se_ResNext50 is also working well."
        },
        {
          "id": 489815,
          "postDate": "2019-03-14T04:50:21.977Z",
          "content": "<p>I have tried Densenet121, Densenet169 and Nasnet-mobile with TTA and each model could LB over 0.9720. </p>",
          "rawMarkdown": "I have tried Densenet121, Densenet169 and Nasnet-mobile with TTA and each model could LB over 0.9720. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 484639,
      "postDate": "2019-03-06T09:19:39.837Z",
      "content": "<p>How are you managing the model ensembling? Are you training each model sequentially on the same kernel, or having separate kernels for each model, then average out their probabilities?</p>\n\n<p>The reason why I'm asking is that my models are taking over 6 hours to train. 😔</p>",
      "rawMarkdown": "How are you managing the model ensembling? Are you training each model sequentially on the same kernel, or having separate kernels for each model, then average out their probabilities?\n\nThe reason why I'm asking is that my models are taking over 6 hours to train. 😔",
      "votes": 1,
      "replies": [
        {
          "id": 484654,
          "postDate": "2019-03-06T09:40:22.223Z",
          "content": "<p>I have separate kernels for each model and one kernel for ensembling where I average out the predictions. You can import output files from your committed kernels using the <code>Add data</code> button on the right panel.</p>\n\n<p>My models are also training for 6-7 hours.</p>",
          "rawMarkdown": "I have separate kernels for each model and one kernel for ensembling where I average out the predictions. You can import output files from your committed kernels using the `Add data` button on the right panel.\n\nMy models are also training for 6-7 hours."
        }
      ]
    },
    {
      "id": 474378,
      "postDate": "2019-02-19T09:26:00.620Z",
      "content": "<p>I compared simple average with power average. Simple average gave better results.</p>",
      "rawMarkdown": "I compared simple average with power average. Simple average gave better results.",
      "votes": 1
    },
    {
      "id": 474423,
      "postDate": "2019-02-19T10:20:25.267Z",
      "content": "<p>I tried to fit different second level models (logistic regressions, RF, etc), I now have 7 different models trained, all 5 folds. So I used their OOF predictions (probabilities, not bottle neck features, that will be next step) as features and trained second level models on that.... it was worse, than just simple averaging of predictions :)</p>",
      "rawMarkdown": "I tried to fit different second level models (logistic regressions, RF, etc), I now have 7 different models trained, all 5 folds. So I used their OOF predictions (probabilities, not bottle neck features, that will be next step) as features and trained second level models on that.... it was worse, than just simple averaging of predictions :)",
      "votes": 2,
      "replies": [
        {
          "id": 474562,
          "postDate": "2019-02-19T14:56:20.567Z",
          "content": "<p>Thanks for the stacking idea! Too bad I haven't saved my OOF predictions. Perhaps adding some classical image features (GLCM, LBP,...) would bring something to second level models.</p>",
          "rawMarkdown": "Thanks for the stacking idea! Too bad I haven't saved my OOF predictions. Perhaps adding some classical image features (GLCM, LBP,...) would bring something to second level models."
        },
        {
          "id": 475170,
          "postDate": "2019-02-20T11:26:57.950Z",
          "content": "<p>nah, just average everything, should be good</p>",
          "rawMarkdown": "nah, just average everything, should be good",
          "votes": 1
        }
      ]
    },
    {
      "id": 500513,
      "postDate": "2019-03-26T06:33:09.437Z",
      "content": "<p>On my side, I assembled 32 from 2 very different models with 3 strategies : \n- pure averaging\n- \"voting system\", then averaging only the ones that mostly voted Yes/No\n-Best confidence : the results with biggest difference to mean (0.5) - the most confident - choose the answer, and then final answer is averaged for those who \"agreed\"</p>\n\n<p>Funny results, with sometimes the voting system getting better results than simple averaging. \"Confidence\" also sometimes wins and it can be further improved </p>",
      "rawMarkdown": "On my side, I assembled 32 from 2 very different models with 3 strategies : \n- pure averaging\n- \"voting system\", then averaging only the ones that mostly voted Yes/No\n-Best confidence : the results with biggest difference to mean (0.5) - the most confident - choose the answer, and then final answer is averaged for those who \"agreed\"\n\nFunny results, with sometimes the voting system getting better results than simple averaging. \"Confidence\" also sometimes wins and it can be further improved "
    }
  ],
  "comments": [
    {
      "id": 474305,
      "author_name": "Joni Juvonen",
      "author_url": "",
      "post_date": "2019-02-19T07:29:51.037000",
      "content": "<p>Okay, I made some tests.</p>\n\n<ul>\n<li>Ensembling 5-fold CV models and two different CNN architectures.</li>\n<li>Each single model scoring between LB 0.958 and LB 0.966.</li>\n<li>I used ranked weights in simple average and power average (best\nsingle model = 1 , second best = 0.5, third = 0.25, etc.) but models with LB ties (diff &lt; 0.1) had the same weights.</li>\n<li>All ensembled models correlated more or less (correlations between 0.875 and 0.975)</li>\n<li>Ranking correlations between 0.84 and 0.96</li>\n</ul>\n\n<h2>Results</h2>\n\n<ul>\n<li><strong>Simple average</strong> <code>LB 0.9709</code></li>\n<li><strong>Power of 2 average</strong> <code>LB 0.9702</code></li>\n<li><strong>Power of 4 average</strong> <code>LB 0.9695</code></li>\n<li><strong>Rank average</strong> <code>LB 0.9689</code></li>\n</ul>\n\n<h2>My conclusion</h2>\n\n<p>Using the most intuitive simple average seems to work the best here. At least when the models are correlated and similar weights are used. I wonder if anyone has different results or other ensembling methods to try.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 479935,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-27T15:38:19.313000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 480145,
          "author_name": "Joni Juvonen",
          "author_url": "",
          "post_date": "2019-02-27T20:42:04.120000",
          "content": "<p>I used DenseNet169 similar to my <a href=\"https://www.kaggle.com/qitvision/a-complete-ml-pipeline-fast-ai\">public kernel</a> but trained the second cycle for longer (16 epochs) and used the full size (96x96) images. se_ResNext50 is also working well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 489815,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-03-14T04:50:21.977000",
          "content": "<p>I have tried Densenet121, Densenet169 and Nasnet-mobile with TTA and each model could LB over 0.9720. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 484639,
      "author_name": "christeekayel",
      "author_url": "",
      "post_date": "2019-03-06T09:19:39.837000",
      "content": "<p>How are you managing the model ensembling? Are you training each model sequentially on the same kernel, or having separate kernels for each model, then average out their probabilities?</p>\n\n<p>The reason why I'm asking is that my models are taking over 6 hours to train. 😔</p>",
      "votes": 1,
      "replies": [
        {
          "id": 484654,
          "author_name": "Joni Juvonen",
          "author_url": "",
          "post_date": "2019-03-06T09:40:22.223000",
          "content": "<p>I have separate kernels for each model and one kernel for ensembling where I average out the predictions. You can import output files from your committed kernels using the <code>Add data</code> button on the right panel.</p>\n\n<p>My models are also training for 6-7 hours.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 474378,
      "author_name": "FelipeKitamura, MD, PhD",
      "author_url": "",
      "post_date": "2019-02-19T09:26:00.620000",
      "content": "<p>I compared simple average with power average. Simple average gave better results.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 474423,
      "author_name": "kvigly",
      "author_url": "",
      "post_date": "2019-02-19T10:20:25.267000",
      "content": "<p>I tried to fit different second level models (logistic regressions, RF, etc), I now have 7 different models trained, all 5 folds. So I used their OOF predictions (probabilities, not bottle neck features, that will be next step) as features and trained second level models on that.... it was worse, than just simple averaging of predictions :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 474562,
          "author_name": "Joni Juvonen",
          "author_url": "",
          "post_date": "2019-02-19T14:56:20.567000",
          "content": "<p>Thanks for the stacking idea! Too bad I haven't saved my OOF predictions. Perhaps adding some classical image features (GLCM, LBP,...) would bring something to second level models.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 475170,
          "author_name": "kvigly",
          "author_url": "",
          "post_date": "2019-02-20T11:26:57.950000",
          "content": "<p>nah, just average everything, should be good</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 500513,
      "author_name": "Simon Caby",
      "author_url": "",
      "post_date": "2019-03-26T06:33:09.437000",
      "content": "<p>On my side, I assembled 32 from 2 very different models with 3 strategies : \n- pure averaging\n- \"voting system\", then averaging only the ones that mostly voted Yes/No\n-Best confidence : the results with biggest difference to mean (0.5) - the most confident - choose the answer, and then final answer is averaged for those who \"agreed\"</p>\n\n<p>Funny results, with sometimes the voting system getting better results than simple averaging. \"Confidence\" also sometimes wins and it can be further improved </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "473697": "What is your ensembling strategy when targeting the AUC metric?\n\nThe area under the ROC curve (AUC) compares the rank of predictions to the rank of correct labels. When ensembling multiple models, more confident but not necessarily more accurate models have a larger impact.\n\nThere are at least three different ensembling methods for AUC:\n\n - **Simple average** (weighted or uniform)\n - **Rank average** (average out the rankings of each model). Pros = confident models have equal impact.\n - **Power average** Power scaling aims to normalize the models: \n \n*PowerAverage = (Pred1^p + Pred2^p + ... + Predn^p) / n*\n\n---------------------------------\n**Relevant links:**\n\n[AUC ensembling blog post](https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e)\n\n[Kaggle ensembling guide](https://mlwave.com/kaggle-ensembling-guide/)",
    "474305": "Okay, I made some tests.\n\n - Ensembling 5-fold CV models and two different CNN architectures.\n - Each single model scoring between LB 0.958 and LB 0.966.\n - I used ranked weights in simple average and power average (best\n   single model = 1 , second best = 0.5, third = 0.25, etc.) but models with LB ties (diff &lt; 0.1) had the same weights.\n - All ensembled models correlated more or less (correlations between 0.875 and 0.975)\n - Ranking correlations between 0.84 and 0.96\n\n## Results\n\n - **Simple average** `LB 0.9709`\n - **Power of 2 average** `LB 0.9702`\n - **Power of 4 average** `LB 0.9695`\n - **Rank average** `LB 0.9689`\n\n## My conclusion\n\nUsing the most intuitive simple average seems to work the best here. At least when the models are correlated and similar weights are used. I wonder if anyone has different results or other ensembling methods to try.",
    "484639": "How are you managing the model ensembling? Are you training each model sequentially on the same kernel, or having separate kernels for each model, then average out their probabilities?\n\nThe reason why I'm asking is that my models are taking over 6 hours to train. 😔",
    "474378": "I compared simple average with power average. Simple average gave better results.",
    "474423": "I tried to fit different second level models (logistic regressions, RF, etc), I now have 7 different models trained, all 5 folds. So I used their OOF predictions (probabilities, not bottle neck features, that will be next step) as features and trained second level models on that.... it was worse, than just simple averaging of predictions :)",
    "500513": "On my side, I assembled 32 from 2 very different models with 3 strategies : \n- pure averaging\n- \"voting system\", then averaging only the ones that mostly voted Yes/No\n-Best confidence : the results with biggest difference to mean (0.5) - the most confident - choose the answer, and then final answer is averaged for those who \"agreed\"\n\nFunny results, with sometimes the voting system getting better results than simple averaging. \"Confidence\" also sometimes wins and it can be further improved "
  }
}