{
  "id": 211221,
  "title": "Ensembling and stacking to maximize AUC",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/211221",
  "author_name": "",
  "post_date": "2021-01-14T09:19:32.807293700Z",
  "votes": 24,
  "comment_count": 4,
  "views": 0,
  "content": "<p>One of the question coming up as we get stuck on improving individual models further is how to best combine multiple models.</p>\n<p>Firstly, there's methods for forming averages, which we can apply to each potential diagnosis separately (after all, the competition metric is the average of the AUCs for each diagnosis - i.e. it is maximized if we maximize the AUC for each diagnosis). Of course, we may wish to use the information of what our model says on other diagnoses in order to take it into account (e.g. being pretty sure on some diagnosis that does not co-occur with another one, probably should affect what we predict for the second diagnosis - but that requires a \"proper\" model).</p>\n<ul>\n<li>Simple average of probabilities such as in <a href=\"https://www.kaggle.com/tpothjuan/efficientnetb7-resnet50-ensemble-tfrecords\" target=\"_blank\">this notebook</a></li>\n<li>Average probabilities with some transformation without back-transformation: e.g. averaging <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194\" target=\"_blank\">square roots of probabilities</a> without backtransformation or stretching from [min predicted, max predicted] to [0, 1], averaging power of 2 or higher (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">\"Power Averaging\"</a>).</li>\n<li>Average probabilities with some transformation with back-transformation (e.g. logit with inv-logit back-transformation)</li>\n<li>Average ranks (we might just as well dividing by the number of observations in the end to get back on the [0,1] scale, in case there's a check that probabilities are between 0 and 1): I first saw this in Abhishek Thakur's \"Approaching (almost) any Machine Learning Problem\" book, it's also discussed in the <a href=\"https://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">Kaggle ensembling guide</a> and it was also pointed out <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564\" target=\"_blank\">on this forum before</a>, upside or downside is that it does not really matter how sure models are about cases only how the predictions are ordered</li>\n</ul>\n<p>Secondly, there's models, which of course would need aligned cross-validation and out-of-fold predictions for all models to implement (while the simple averages can all be done without - you just don't have much an idea of what to expect beyond what the LB tells you):</p>\n<ul>\n<li>Weighted average (in one of the forms above) chosen so as to minimize the AUC (e.g. using <code>scipy.optimize</code>) applied to any of the averaging techniques above - sort of the simplest model and does still not account that probabilities for one class maybe should affect the probabilities for other classes</li>\n<li>Logistic regression with features consisting of e.g. ranks, probabilities or some transformation of probabilities (downside loss function is not directly aligned with competition metric) - possibly with something based on some simple average as an offset term</li>\n<li>Linear regression on ranks: Basically the idea is to do some modelling of the ranks to improve what we achieve via rank averaging.</li>\n<li>Neural network with <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204690\" target=\"_blank\">AUC loss function</a>: We probably want to keep such a neural network shallow in order to avoid overfitting, but it could be a nice attempt at combining a loss function that aligned to the competition metric with something that can take the probabilities for other categories into account</li>\n</ul>\n<p>Previous competitions to look at for ensembling ideas would include those that also used AUC, in terms of medical imaging there's mostly <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation\" target=\"_blank\">SIIM-ISIC Melanoma Classification</a>. As pointed out in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207775\" target=\"_blank\">another post</a>, the top solutions are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175504\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175324\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/177563\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175633\" target=\"_blank\">here</a>, and of course the <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/overview/evaluation\" target=\"_blank\">Histopathologic Cancer Detection playground</a>. Interesting takeaways on ensembling:</p>\n<ul>\n<li>First place used external data (previous competitions) also for validation and had two ensembles: one ensemble with all out-of-folds and one with the out-of-folds only from the current competition</li>\n<li>Both first and second place seem to just have used rank averages, while the third place seemed to <a href=\"https://github.com/Masdevallia/3rd-place-kaggle-siim-isic-melanoma-classification/blob/master/ensembling.py\" target=\"_blank\">average with some weights based on meta-data</a> (not quite clear to me)<br>\nHowever, when we go beyond medical imaging, there's also e.g. <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783\" target=\"_blank\">where we find this nice discussion</a>.</li>\n<li><a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/80979\" target=\"_blank\">Discussion</a> of techniques for ensembling</li>\n<li>Non-imaging <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/overview/evaluation\" target=\"_blank\">Santander Customer Transaction Prediction</a>: several teams such as <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88902\" target=\"_blank\">3rd place solution</a> and <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88897\" target=\"_blank\">5th place</a> used a neural network for stacking; also there's a a nice <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783\" target=\"_blank\">discussion thread</a></li>\n<li>I've probably missed a lot of useful other stuff in the long history of Kaggle.</li>\n</ul>\n<p>What does not seem to make any sense, at all: voting.</p>\n<p>I guess in the end, we might just as well also apply some monotone transformation to our predicted probabilities to improve their calibration. It's not like it will help (or harm) the competition metric, but I kind of would feel bad for doing well with a good ordering of predicted probabilities between 0.95 and 0.999 (when in fact the 0.95s should be close to zero, really)…</p>\n<p><strong>Update:</strong> It was pointed out in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/213007\" target=\"_blank\">another discussion post</a> that there's a <a href=\"https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e\" target=\"_blank\">really nice Medium post</a> about AUC ensembling.</p>",
  "messages": [
    {
      "id": "1152564",
      "postDate": "01/14/2021 09:19:32",
      "content": "<p>One of the question coming up as we get stuck on improving individual models further is how to best combine multiple models.</p>\n<p>Firstly, there's methods for forming averages, which we can apply to each potential diagnosis separately (after all, the competition metric is the average of the AUCs for each diagnosis - i.e. it is maximized if we maximize the AUC for each diagnosis). Of course, we may wish to use the information of what our model says on other diagnoses in order to take it into account (e.g. being pretty sure on some diagnosis that does not co-occur with another one, probably should affect what we predict for the second diagnosis - but that requires a \"proper\" model).</p>\n<ul>\n<li>Simple average of probabilities such as in <a href=\"https://www.kaggle.com/tpothjuan/efficientnetb7-resnet50-ensemble-tfrecords\" target=\"_blank\">this notebook</a></li>\n<li>Average probabilities with some transformation without back-transformation: e.g. averaging <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194\" target=\"_blank\">square roots of probabilities</a> without backtransformation or stretching from [min predicted, max predicted] to [0, 1], averaging power of 2 or higher (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">\"Power Averaging\"</a>).</li>\n<li>Average probabilities with some transformation with back-transformation (e.g. logit with inv-logit back-transformation)</li>\n<li>Average ranks (we might just as well dividing by the number of observations in the end to get back on the [0,1] scale, in case there's a check that probabilities are between 0 and 1): I first saw this in Abhishek Thakur's \"Approaching (almost) any Machine Learning Problem\" book, it's also discussed in the <a href=\"https://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">Kaggle ensembling guide</a> and it was also pointed out <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564\" target=\"_blank\">on this forum before</a>, upside or downside is that it does not really matter how sure models are about cases only how the predictions are ordered</li>\n</ul>\n<p>Secondly, there's models, which of course would need aligned cross-validation and out-of-fold predictions for all models to implement (while the simple averages can all be done without - you just don't have much an idea of what to expect beyond what the LB tells you):</p>\n<ul>\n<li>Weighted average (in one of the forms above) chosen so as to minimize the AUC (e.g. using <code>scipy.optimize</code>) applied to any of the averaging techniques above - sort of the simplest model and does still not account that probabilities for one class maybe should affect the probabilities for other classes</li>\n<li>Logistic regression with features consisting of e.g. ranks, probabilities or some transformation of probabilities (downside loss function is not directly aligned with competition metric) - possibly with something based on some simple average as an offset term</li>\n<li>Linear regression on ranks: Basically the idea is to do some modelling of the ranks to improve what we achieve via rank averaging.</li>\n<li>Neural network with <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204690\" target=\"_blank\">AUC loss function</a>: We probably want to keep such a neural network shallow in order to avoid overfitting, but it could be a nice attempt at combining a loss function that aligned to the competition metric with something that can take the probabilities for other categories into account</li>\n</ul>\n<p>Previous competitions to look at for ensembling ideas would include those that also used AUC, in terms of medical imaging there's mostly <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation\" target=\"_blank\">SIIM-ISIC Melanoma Classification</a>. As pointed out in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207775\" target=\"_blank\">another post</a>, the top solutions are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175504\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175324\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/177563\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175633\" target=\"_blank\">here</a>, and of course the <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/overview/evaluation\" target=\"_blank\">Histopathologic Cancer Detection playground</a>. Interesting takeaways on ensembling:</p>\n<ul>\n<li>First place used external data (previous competitions) also for validation and had two ensembles: one ensemble with all out-of-folds and one with the out-of-folds only from the current competition</li>\n<li>Both first and second place seem to just have used rank averages, while the third place seemed to <a href=\"https://github.com/Masdevallia/3rd-place-kaggle-siim-isic-melanoma-classification/blob/master/ensembling.py\" target=\"_blank\">average with some weights based on meta-data</a> (not quite clear to me)<br>\nHowever, when we go beyond medical imaging, there's also e.g. <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783\" target=\"_blank\">where we find this nice discussion</a>.</li>\n<li><a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/80979\" target=\"_blank\">Discussion</a> of techniques for ensembling</li>\n<li>Non-imaging <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/overview/evaluation\" target=\"_blank\">Santander Customer Transaction Prediction</a>: several teams such as <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88902\" target=\"_blank\">3rd place solution</a> and <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88897\" target=\"_blank\">5th place</a> used a neural network for stacking; also there's a a nice <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783\" target=\"_blank\">discussion thread</a></li>\n<li>I've probably missed a lot of useful other stuff in the long history of Kaggle.</li>\n</ul>\n<p>What does not seem to make any sense, at all: voting.</p>\n<p>I guess in the end, we might just as well also apply some monotone transformation to our predicted probabilities to improve their calibration. It's not like it will help (or harm) the competition metric, but I kind of would feel bad for doing well with a good ordering of predicted probabilities between 0.95 and 0.999 (when in fact the 0.95s should be close to zero, really)…</p>\n<p><strong>Update:</strong> It was pointed out in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/213007\" target=\"_blank\">another discussion post</a> that there's a <a href=\"https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e\" target=\"_blank\">really nice Medium post</a> about AUC ensembling.</p>",
      "rawMarkdown": "One of the question coming up as we get stuck on improving individual models further is how to best combine multiple models.\n\nFirstly, there's methods for forming averages, which we can apply to each potential diagnosis separately (after all, the competition metric is the average of the AUCs for each diagnosis - i.e. it is maximized if we maximize the AUC for each diagnosis). Of course, we may wish to use the information of what our model says on other diagnoses in order to take it into account (e.g. being pretty sure on some diagnosis that does not co-occur with another one, probably should affect what we predict for the second diagnosis - but that requires a \"proper\" model).\n\n* Simple average of probabilities such as in [this notebook](https://www.kaggle.com/tpothjuan/efficientnetb7-resnet50-ensemble-tfrecords)\n* Average probabilities with some transformation without back-transformation: e.g. averaging [square roots of probabilities](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194) without backtransformation or stretching from [min predicted, max predicted] to [0, 1], averaging power of 2 or higher ([\"Power Averaging\"](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653)).\n* Average probabilities with some transformation with back-transformation (e.g. logit with inv-logit back-transformation)\n* Average ranks (we might just as well dividing by the number of observations in the end to get back on the [0,1] scale, in case there's a check that probabilities are between 0 and 1): I first saw this in Abhishek Thakur's \"Approaching (almost) any Machine Learning Problem\" book, it's also discussed in the [Kaggle ensembling guide](https://mlwave.com/kaggle-ensembling-guide/) and it was also pointed out [on this forum before](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564), upside or downside is that it does not really matter how sure models are about cases only how the predictions are ordered\n\nSecondly, there's models, which of course would need aligned cross-validation and out-of-fold predictions for all models to implement (while the simple averages can all be done without - you just don't have much an idea of what to expect beyond what the LB tells you):\n* Weighted average (in one of the forms above) chosen so as to minimize the AUC (e.g. using `scipy.optimize`) applied to any of the averaging techniques above - sort of the simplest model and does still not account that probabilities for one class maybe should affect the probabilities for other classes\n* Logistic regression with features consisting of e.g. ranks, probabilities or some transformation of probabilities (downside loss function is not directly aligned with competition metric) - possibly with something based on some simple average as an offset term\n* Linear regression on ranks: Basically the idea is to do some modelling of the ranks to improve what we achieve via rank averaging.\n* Neural network with [AUC loss function](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204690): We probably want to keep such a neural network shallow in order to avoid overfitting, but it could be a nice attempt at combining a loss function that aligned to the competition metric with something that can take the probabilities for other categories into account\n\nPrevious competitions to look at for ensembling ideas would include those that also used AUC, in terms of medical imaging there's mostly [SIIM-ISIC Melanoma Classification](https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation). As pointed out in [another post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207775), the top solutions are [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175504), [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412), [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175324), [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/177563) and [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175633), and of course the [Histopathologic Cancer Detection playground](https://www.kaggle.com/c/histopathologic-cancer-detection/overview/evaluation). Interesting takeaways on ensembling:\n* First place used external data (previous competitions) also for validation and had two ensembles: one ensemble with all out-of-folds and one with the out-of-folds only from the current competition\n* Both first and second place seem to just have used rank averages, while the third place seemed to [average with some weights based on meta-data](https://github.com/Masdevallia/3rd-place-kaggle-siim-isic-melanoma-classification/blob/master/ensembling.py) (not quite clear to me)\nHowever, when we go beyond medical imaging, there's also e.g. [where we find this nice discussion](https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783).\n* [Discussion](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/80979) of techniques for ensembling\n* Non-imaging [Santander Customer Transaction Prediction](https://www.kaggle.com/c/santander-customer-transaction-prediction/overview/evaluation): several teams such as [3rd place solution](https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88902) and [5th place](https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88897) used a neural network for stacking; also there's a a nice [discussion thread](https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783)\n* I've probably missed a lot of useful other stuff in the long history of Kaggle.\n\nWhat does not seem to make any sense, at all: voting.\n\nI guess in the end, we might just as well also apply some monotone transformation to our predicted probabilities to improve their calibration. It's not like it will help (or harm) the competition metric, but I kind of would feel bad for doing well with a good ordering of predicted probabilities between 0.95 and 0.999 (when in fact the 0.95s should be close to zero, really)...\n\n**Update:** It was pointed out in [another discussion post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/213007) that there's a [really nice Medium post](https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e) about AUC ensembling.",
      "votes": null
    },
    {
      "id": "1182662",
      "postDate": "02/02/2021 14:32:08",
      "content": "<p>I tried ensembling (simple average) on of my model (0.954) with the public (0.965) one and I don't know why it scored worst (0.8xx). Is there any other way to overcome this ? The submission file was created with sorting the predictions by image ids but I don't think that was the problem for this abysmal score.</p>",
      "rawMarkdown": "I tried ensembling (simple average) on of my model (0.954) with the public (0.965) one and I don't know why it scored worst (0.8xx). Is there any other way to overcome this ? The submission file was created with sorting the predictions by image ids but I don't think that was the problem for this abysmal score.",
      "votes": null
    },
    {
      "id": "1182878",
      "postDate": "02/02/2021 16:22:11",
      "content": "<p>That sounds almost bizarre. I would seriously consider whether something has gone wrong in combining the model predictions. </p>\n<p>However, one possible explanation is the competition metric: Let's ignore that we have multiple classes and let's say we have just one 0/1 prediction. Example of some true labels: [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]. With that scenario, we can easily construct examples where averaging two models is worse than the original models, e.g. averaging 3 models, here's a drop from 0.96 for each  model to 0.88 for the arithmetic average of their predictions:</p>\n<pre><code>import numpy as np\nfrom sklearn import metrics\n\ny = [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]\n\npreds = [\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.34, 0.491, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.34, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.492, 0.34, 0.493, 0.494] ]\n\nfor pred in preds:\n    fpr, tpr, thresholds = metrics.roc_curve(np.array(y), np.array(pred), pos_label=1)\n    print(metrics.auc(fpr, tpr))\n\nfpr, tpr, thresholds = metrics.roc_curve(np.\nprint(metrics.auc(fpr, tpr))array(y), np.array(np.array(preds).mean(axis=0)), pos_label=1)\nprint(metrics.auc(fpr, tpr))\n</code></pre>\n<p>That's, of course, the reason why doing things other than the arithmetic mean may be better (e.g. rank averages).</p>",
      "rawMarkdown": "That sounds almost bizarre. I would seriously consider whether something has gone wrong in combining the model predictions. \n\nHowever, one possible explanation is the competition metric: Let's ignore that we have multiple classes and let's say we have just one 0/1 prediction. Example of some true labels: [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]. With that scenario, we can easily construct examples where averaging two models is worse than the original models, e.g. averaging 3 models, here's a drop from 0.96 for each  model to 0.88 for the arithmetic average of their predictions:\n```\nimport numpy as np\nfrom sklearn import metrics\n\ny = [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]\n\npreds = [\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.34, 0.491, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.34, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.492, 0.34, 0.493, 0.494] ]\n\nfor pred in preds:\n    fpr, tpr, thresholds = metrics.roc_curve(np.array(y), np.array(pred), pos_label=1)\n    print(metrics.auc(fpr, tpr))\n    \nfpr, tpr, thresholds = metrics.roc_curve(np.\nprint(metrics.auc(fpr, tpr))array(y), np.array(np.array(preds).mean(axis=0)), pos_label=1)\nprint(metrics.auc(fpr, tpr))\n\n```\n\nThat's, of course, the reason why doing things other than the arithmetic mean may be better (e.g. rank averages).",
      "votes": null
    },
    {
      "id": "1182972",
      "postDate": "02/02/2021 17:03:06",
      "content": "<p>Thanks, I will try other ensembling methods. I also noticed the same with the public notebook based on simple averages <a href=\"https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords?scriptVersionId=51477260\" target=\"_blank\">EfficientNetB7 + ResNet50 Ensemble + TFRecords</a> <em>Version 29</em>. You could see the score drop because of simple average.</p>",
      "rawMarkdown": "Thanks, I will try other ensembling methods. I also noticed the same with the public notebook based on simple averages [EfficientNetB7 + ResNet50 Ensemble + TFRecords](https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords?scriptVersionId=51477260) *Version 29*. You could see the score drop because of simple average.",
      "votes": null
    },
    {
      "id": "1235495",
      "postDate": "03/12/2021 07:35:18",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Thanks for this, really useful tbh, I didn't even realise some characteristics of AUC!</p>",
      "rawMarkdown": "bjoernholzhauer Thanks for this, really useful tbh, I didn't even realise some characteristics of AUC!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1182662,
      "author_name": "thakurudit",
      "author_url": "",
      "post_date": "02/02/2021 14:32:08",
      "content": "<p>I tried ensembling (simple average) on of my model (0.954) with the public (0.965) one and I don't know why it scored worst (0.8xx). Is there any other way to overcome this ? The submission file was created with sorting the predictions by image ids but I don't think that was the problem for this abysmal score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1182878,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "02/02/2021 16:22:11",
          "content": "<p>That sounds almost bizarre. I would seriously consider whether something has gone wrong in combining the model predictions. </p>\n<p>However, one possible explanation is the competition metric: Let's ignore that we have multiple classes and let's say we have just one 0/1 prediction. Example of some true labels: [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]. With that scenario, we can easily construct examples where averaging two models is worse than the original models, e.g. averaging 3 models, here's a drop from 0.96 for each  model to 0.88 for the arithmetic average of their predictions:</p>\n<pre><code>import numpy as np\nfrom sklearn import metrics\n\ny = [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]\n\npreds = [\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.34, 0.491, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.34, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.492, 0.34, 0.493, 0.494] ]\n\nfor pred in preds:\n    fpr, tpr, thresholds = metrics.roc_curve(np.array(y), np.array(pred), pos_label=1)\n    print(metrics.auc(fpr, tpr))\n\nfpr, tpr, thresholds = metrics.roc_curve(np.\nprint(metrics.auc(fpr, tpr))array(y), np.array(np.array(preds).mean(axis=0)), pos_label=1)\nprint(metrics.auc(fpr, tpr))\n</code></pre>\n<p>That's, of course, the reason why doing things other than the arithmetic mean may be better (e.g. rank averages).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1182972,
          "author_name": "thakurudit",
          "author_url": "",
          "post_date": "02/02/2021 17:03:06",
          "content": "<p>Thanks, I will try other ensembling methods. I also noticed the same with the public notebook based on simple averages <a href=\"https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords?scriptVersionId=51477260\" target=\"_blank\">EfficientNetB7 + ResNet50 Ensemble + TFRecords</a> <em>Version 29</em>. You could see the score drop because of simple average.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1235495,
      "author_name": "reighns",
      "author_url": "",
      "post_date": "03/12/2021 07:35:18",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Thanks for this, really useful tbh, I didn't even realise some characteristics of AUC!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1152564": "One of the question coming up as we get stuck on improving individual models further is how to best combine multiple models.\n\nFirstly, there's methods for forming averages, which we can apply to each potential diagnosis separately (after all, the competition metric is the average of the AUCs for each diagnosis - i.e. it is maximized if we maximize the AUC for each diagnosis). Of course, we may wish to use the information of what our model says on other diagnoses in order to take it into account (e.g. being pretty sure on some diagnosis that does not co-occur with another one, probably should affect what we predict for the second diagnosis - but that requires a \"proper\" model).\n\n* Simple average of probabilities such as in [this notebook](https://www.kaggle.com/tpothjuan/efficientnetb7-resnet50-ensemble-tfrecords)\n* Average probabilities with some transformation without back-transformation: e.g. averaging [square roots of probabilities](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/211194) without backtransformation or stretching from [min predicted, max predicted] to [0, 1], averaging power of 2 or higher ([\"Power Averaging\"](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653)).\n* Average probabilities with some transformation with back-transformation (e.g. logit with inv-logit back-transformation)\n* Average ranks (we might just as well dividing by the number of observations in the end to get back on the [0,1] scale, in case there's a check that probabilities are between 0 and 1): I first saw this in Abhishek Thakur's \"Approaching (almost) any Machine Learning Problem\" book, it's also discussed in the [Kaggle ensembling guide](https://mlwave.com/kaggle-ensembling-guide/) and it was also pointed out [on this forum before](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205564), upside or downside is that it does not really matter how sure models are about cases only how the predictions are ordered\n\nSecondly, there's models, which of course would need aligned cross-validation and out-of-fold predictions for all models to implement (while the simple averages can all be done without - you just don't have much an idea of what to expect beyond what the LB tells you):\n* Weighted average (in one of the forms above) chosen so as to minimize the AUC (e.g. using `scipy.optimize`) applied to any of the averaging techniques above - sort of the simplest model and does still not account that probabilities for one class maybe should affect the probabilities for other classes\n* Logistic regression with features consisting of e.g. ranks, probabilities or some transformation of probabilities (downside loss function is not directly aligned with competition metric) - possibly with something based on some simple average as an offset term\n* Linear regression on ranks: Basically the idea is to do some modelling of the ranks to improve what we achieve via rank averaging.\n* Neural network with [AUC loss function](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204690): We probably want to keep such a neural network shallow in order to avoid overfitting, but it could be a nice attempt at combining a loss function that aligned to the competition metric with something that can take the probabilities for other categories into account\n\nPrevious competitions to look at for ensembling ideas would include those that also used AUC, in terms of medical imaging there's mostly [SIIM-ISIC Melanoma Classification](https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation). As pointed out in [another post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207775), the top solutions are [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175504), [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175412), [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175324), [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/177563) and [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175633), and of course the [Histopathologic Cancer Detection playground](https://www.kaggle.com/c/histopathologic-cancer-detection/overview/evaluation). Interesting takeaways on ensembling:\n* First place used external data (previous competitions) also for validation and had two ensembles: one ensemble with all out-of-folds and one with the out-of-folds only from the current competition\n* Both first and second place seem to just have used rank averages, while the third place seemed to [average with some weights based on meta-data](https://github.com/Masdevallia/3rd-place-kaggle-siim-isic-melanoma-classification/blob/master/ensembling.py) (not quite clear to me)\nHowever, when we go beyond medical imaging, there's also e.g. [where we find this nice discussion](https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783).\n* [Discussion](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/80979) of techniques for ensembling\n* Non-imaging [Santander Customer Transaction Prediction](https://www.kaggle.com/c/santander-customer-transaction-prediction/overview/evaluation): several teams such as [3rd place solution](https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88902) and [5th place](https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/88897) used a neural network for stacking; also there's a a nice [discussion thread](https://www.kaggle.com/c/santander-customer-satisfaction/discussion/20783)\n* I've probably missed a lot of useful other stuff in the long history of Kaggle.\n\nWhat does not seem to make any sense, at all: voting.\n\nI guess in the end, we might just as well also apply some monotone transformation to our predicted probabilities to improve their calibration. It's not like it will help (or harm) the competition metric, but I kind of would feel bad for doing well with a good ordering of predicted probabilities between 0.95 and 0.999 (when in fact the 0.95s should be close to zero, really)...\n\n**Update:** It was pointed out in [another discussion post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/213007) that there's a [really nice Medium post](https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e) about AUC ensembling.",
    "1182662": "I tried ensembling (simple average) on of my model (0.954) with the public (0.965) one and I don't know why it scored worst (0.8xx). Is there any other way to overcome this ? The submission file was created with sorting the predictions by image ids but I don't think that was the problem for this abysmal score.",
    "1182878": "That sounds almost bizarre. I would seriously consider whether something has gone wrong in combining the model predictions. \n\nHowever, one possible explanation is the competition metric: Let's ignore that we have multiple classes and let's say we have just one 0/1 prediction. Example of some true labels: [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]. With that scenario, we can easily construct examples where averaging two models is worse than the original models, e.g. averaging 3 models, here's a drop from 0.96 for each  model to 0.88 for the arithmetic average of their predictions:\n```\nimport numpy as np\nfrom sklearn import metrics\n\ny = [0, 0, 0, 0, 0, 1, 1, 1, 1, 1]\n\npreds = [\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.34, 0.491, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.34, 0.492, 0.493, 0.494],\n    [0.336, 0.337, 0.338, 0.339, 0.4909, 0.491, 0.492, 0.34, 0.493, 0.494] ]\n\nfor pred in preds:\n    fpr, tpr, thresholds = metrics.roc_curve(np.array(y), np.array(pred), pos_label=1)\n    print(metrics.auc(fpr, tpr))\n    \nfpr, tpr, thresholds = metrics.roc_curve(np.\nprint(metrics.auc(fpr, tpr))array(y), np.array(np.array(preds).mean(axis=0)), pos_label=1)\nprint(metrics.auc(fpr, tpr))\n\n```\n\nThat's, of course, the reason why doing things other than the arithmetic mean may be better (e.g. rank averages).",
    "1182972": "Thanks, I will try other ensembling methods. I also noticed the same with the public notebook based on simple averages [EfficientNetB7 + ResNet50 Ensemble + TFRecords](https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords?scriptVersionId=51477260) *Version 29*. You could see the score drop because of simple average.",
    "1235495": "bjoernholzhauer Thanks for this, really useful tbh, I didn't even realise some characteristics of AUC!"
  },
  "source": "meta"
}