{
  "id": 165653,
  "title": "For AUC Metric, Ensemble using Power Averaging",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165653",
  "author_name": "Sirish Somanchi",
  "post_date": "2020-07-10T13:23:47.602000",
  "votes": 114,
  "comment_count": 20,
  "views": 0,
  "content": "<p>In Chris's excellent post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">CNN Input Size Explained</a>, he suggests ensembling models of different input sizes. For blending the predictions of all these models, see <strong>Power Averaging</strong> technique below, which was used successfully in older competitions to <strong>obtain Medals</strong> 😄 </p>\n\n<p>Various Prediction Averaging/Blending Techniques:</p>\n\n<ol>\n<li><p>Simple Averaging: Most participants are using just a simple mean of predictions generated by different models</p></li>\n<li><p>Rank Averaging: Use the \"rank\" of an input image instead it's prediction value. See public notebook <a href=\"https://www.kaggle.com/niteshx2/improve-blending-using-rankdata\">Improve blending using Rankdata</a></p></li>\n<li><p>Weighted Averaging: Specify weights, say 0.5 each in case of two models\nWeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn)\nwhere, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1</p></li>\n<li><p>Stretch Averaging: Stretch predictions using min and max values first, before averaging\nPred = (Pred - min(Pred)) / (max(Pred) - min(Pred))</p></li>\n<li><p>Power Averaging: Choose a power p = 2, 4, 8, 16\nPowerAverage(p) = (Pred1^p + Pred2^p + … + Predn^p) / n\nNote: Power Averaging to be used only when all the models are highly correlated, otherwise your score may become worse.</p></li>\n<li><p>Power Averaging with weights:\nPowerAverageWithWeights(p) = (wt1 x Pred1^p + wt2 x Pred2^p + … + wtn x Predn^p)</p></li>\n</ol>\n\n<p>Note: Ensemble your <strong>best models</strong> for final submission, typically about 2 weeks before the competition deadline.</p>\n\n<p>PS: Read this post by a Kaggle Master <a href=\"https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e\">Power/geometric ensembling when targeting the AUC metric</a> for a more detailed explanation with a lot of calculations.</p>",
  "messages": [
    {
      "id": 922999,
      "postDate": "2020-07-10T13:23:47.603Z",
      "content": "<p>In Chris's excellent post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\">CNN Input Size Explained</a>, he suggests ensembling models of different input sizes. For blending the predictions of all these models, see <strong>Power Averaging</strong> technique below, which was used successfully in older competitions to <strong>obtain Medals</strong> 😄 </p>\n\n<p>Various Prediction Averaging/Blending Techniques:</p>\n\n<ol>\n<li><p>Simple Averaging: Most participants are using just a simple mean of predictions generated by different models</p></li>\n<li><p>Rank Averaging: Use the \"rank\" of an input image instead it's prediction value. See public notebook <a href=\"https://www.kaggle.com/niteshx2/improve-blending-using-rankdata\">Improve blending using Rankdata</a></p></li>\n<li><p>Weighted Averaging: Specify weights, say 0.5 each in case of two models\nWeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn)\nwhere, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1</p></li>\n<li><p>Stretch Averaging: Stretch predictions using min and max values first, before averaging\nPred = (Pred - min(Pred)) / (max(Pred) - min(Pred))</p></li>\n<li><p>Power Averaging: Choose a power p = 2, 4, 8, 16\nPowerAverage(p) = (Pred1^p + Pred2^p + … + Predn^p) / n\nNote: Power Averaging to be used only when all the models are highly correlated, otherwise your score may become worse.</p></li>\n<li><p>Power Averaging with weights:\nPowerAverageWithWeights(p) = (wt1 x Pred1^p + wt2 x Pred2^p + … + wtn x Predn^p)</p></li>\n</ol>\n\n<p>Note: Ensemble your <strong>best models</strong> for final submission, typically about 2 weeks before the competition deadline.</p>\n\n<p>PS: Read this post by a Kaggle Master <a href=\"https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e\">Power/geometric ensembling when targeting the AUC metric</a> for a more detailed explanation with a lot of calculations.</p>",
      "rawMarkdown": "In Chris's excellent post [CNN Input Size Explained](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147), he suggests ensembling models of different input sizes. For blending the predictions of all these models, see **Power Averaging** technique below, which was used successfully in older competitions to **obtain Medals** 😄 \n\nVarious Prediction Averaging/Blending Techniques:\n\n1. Simple Averaging: Most participants are using just a simple mean of predictions generated by different models\n\n2. Rank Averaging: Use the \"rank\" of an input image instead it's prediction value. See public notebook [Improve blending using Rankdata](https://www.kaggle.com/niteshx2/improve-blending-using-rankdata)\n\n3. Weighted Averaging: Specify weights, say 0.5 each in case of two models\n   WeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn)\n   where, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1\n\n4. Stretch Averaging: Stretch predictions using min and max values first, before averaging\n   Pred = (Pred - min(Pred)) / (max(Pred) - min(Pred))\n\n5. Power Averaging: Choose a power p = 2, 4, 8, 16\n   PowerAverage(p) = (Pred1^p + Pred2^p + … + Predn^p) / n\n   Note: Power Averaging to be used only when all the models are highly correlated, otherwise your score may become worse.\n\n6. Power Averaging with weights:\n   PowerAverageWithWeights(p) = (wt1 x Pred1^p + wt2 x Pred2^p + … + wtn x Predn^p)\n\nNote: Ensemble your **best models** for final submission, typically about 2 weeks before the competition deadline.\n\nPS: Read this post by a Kaggle Master [Power/geometric ensembling when targeting the AUC metric](https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e) for a more detailed explanation with a lot of calculations.",
      "votes": 113
    },
    {
      "id": 961196,
      "postDate": "2020-08-07T01:54:15.857Z",
      "content": "<p>wow.. amazing technique..</p>\n<p>Actually I was looking for something like this for a long time. </p>",
      "rawMarkdown": "wow.. amazing technique..\n\nActually I was looking for something like this for a long time. \n",
      "votes": 1
    },
    {
      "id": 954676,
      "postDate": "2020-08-02T01:53:49.433Z",
      "content": "<p>Thank you for sharing this, I have to take notes and probably make a notebook on all this.</p>",
      "rawMarkdown": "Thank you for sharing this, I have to take notes and probably make a notebook on all this.",
      "votes": 1
    },
    {
      "id": 925894,
      "postDate": "2020-07-12T11:20:45.713Z",
      "content": "<p>Thanks for sharing this <a href=\"/sirishks\">@sirishks</a>.</p>\n\n<p>I guess the theoretical validation of using rank averaging is a bit stronger than simple averaging since we are choosing the best epoch on AUC where probabilities by themselves are disregarded (only the order of probabilities is considered). However, theory hardly agrees with practice.</p>\n\n<p>How would you apply advanced concepts like weighted averaging or power averaging to rank averaging? Using power averaging on rank averaging by itself would potentially lead to disastrous effects since it would scale images with values closer to <code>1</code> much lesser than those close to <code>N</code>, whereas the metric confers equal weight to each image whether the output is on one end of the spectrum or the other. However, to think about it, even in the case of power averaging on the probabilities, the probabilities closer to <code>1.00</code> are suppressed much lesser than those closer to <code>0.00</code>. But, probably that's a desirable quality. And since the higher p's provide better results, probably power averaging on rank averaging doesn't sound that bad an idea after all.</p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "Thanks for sharing this @sirishks.\n\nI guess the theoretical validation of using rank averaging is a bit stronger than simple averaging since we are choosing the best epoch on AUC where probabilities by themselves are disregarded (only the order of probabilities is considered). However, theory hardly agrees with practice.\n\nHow would you apply advanced concepts like weighted averaging or power averaging to rank averaging? Using power averaging on rank averaging by itself would potentially lead to disastrous effects since it would scale images with values closer to `1` much lesser than those close to `N`, whereas the metric confers equal weight to each image whether the output is on one end of the spectrum or the other. However, to think about it, even in the case of power averaging on the probabilities, the probabilities closer to `1.00` are suppressed much lesser than those closer to `0.00`. But, probably that's a desirable quality. And since the higher p's provide better results, probably power averaging on rank averaging doesn't sound that bad an idea after all.\n\nWhat do you think?",
      "votes": 1
    },
    {
      "id": 923129,
      "postDate": "2020-07-10T14:52:04.553Z",
      "content": "<p>Thanks a lot for this <a href=\"/sirishks\">@sirishks</a>, very helpful!</p>\n\n<p>But we need to be <strong>extremely</strong> careful about choosing the models for power averaging, one bad model and everything might go for a toss.</p>\n\n<p>Also, in point 3 you mention:</p>\n\n<blockquote>\n  <p>Weighted Averaging: Specify weights, say 0.5 each in case of two models\n  WeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn) / n\n  where, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1</p>\n</blockquote>\n\n<p>Shouldn't we have <code>wt1+...+wtn = n</code>, since we are dividing the equation by <code>n</code> OR just remove <code>n</code> from the denominator and have <code>wt1+...+wtn = 1</code>? Correct me if I am wrong.</p>",
      "rawMarkdown": "Thanks a lot for this @sirishks, very helpful!\n\nBut we need to be **extremely** careful about choosing the models for power averaging, one bad model and everything might go for a toss.\n\nAlso, in point 3 you mention:\n&gt;Weighted Averaging: Specify weights, say 0.5 each in case of two models\nWeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn) / n\nwhere, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1\n\nShouldn't we have `wt1+...+wtn = n`, since we are dividing the equation by `n` OR just remove `n` from the denominator and have `wt1+...+wtn = 1`? Correct me if I am wrong.",
      "votes": 2,
      "replies": [
        {
          "id": 923160,
          "postDate": "2020-07-10T15:14:18.473Z",
          "content": "<p><a href=\"/abhishekgbhat\">@abhishekgbhat</a> yes, it was a typo. You are correct👍 </p>\n\n<p>I have edited the answer and removed \"/n\" when using weights in the range [0,1] such that sum of weights = 1.</p>",
          "rawMarkdown": "@abhishekgbhat yes, it was a typo. You are correct👍 \n\nI have edited the answer and removed \"/n\" when using weights in the range [0,1] such that sum of weights = 1.",
          "votes": 1
        },
        {
          "id": 923295,
          "postDate": "2020-07-10T17:42:22.710Z",
          "content": "<p><a href=\"/abhishekgbhat\">@abhishekgbhat</a> AUC is an interesting metric -- all it cares about is the order of your predictions. Dividing the predictions by any common number (like n) is not going to change the values of the metric, so, from practical standpoint, it doesn't make a difference.</p>",
          "rawMarkdown": "@abhishekgbhat AUC is an interesting metric -- all it cares about is the order of your predictions. Dividing the predictions by any common number (like n) is not going to change the values of the metric, so, from practical standpoint, it doesn't make a difference.",
          "votes": 1
        },
        {
          "id": 923314,
          "postDate": "2020-07-10T18:04:44.183Z",
          "content": "<p><a href=\"/graf10a\">@graf10a</a> I totally agree with what you are saying about the AUC.. I was making a general comment about the weighted average definition. But as you mentioned, when the metric is AUC all that matters is the order.</p>",
          "rawMarkdown": "@graf10a I totally agree with what you are saying about the AUC.. I was making a general comment about the weighted average definition. But as you mentioned, when the metric is AUC all that matters is the order.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3361837,
      "postDate": "2025-12-03T21:29:27.237Z",
      "content": "<p>Well Written!</p>\n<p>Have a look at my notebook <a href=\"https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble\" target=\"_blank\">https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble</a></p>",
      "rawMarkdown": "Well Written!\n\n Have a look at my notebook https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble"
    },
    {
      "id": 969206,
      "postDate": "2020-08-13T14:46:43.310Z",
      "content": "<p>Nice Explanation </p>",
      "rawMarkdown": "Nice Explanation "
    },
    {
      "id": 951748,
      "postDate": "2020-07-30T11:28:35.280Z",
      "content": "<p><a href=\"/sirishks\">@sirishks</a> thanks for the post. In the mentioned blog, there is a term <code>weight_initializer</code>in part 5 of blog which I could not understand. Can you please clarify what is <code>weight_initializer</code>and how one can calculate it ?</p>",
      "rawMarkdown": "@sirishks thanks for the post. In the mentioned blog, there is a term `weight_initializer `in part 5 of blog which I could not understand. Can you please clarify what is `weight_initializer `and how one can calculate it ?"
    },
    {
      "id": 949974,
      "postDate": "2020-07-29T05:19:57.040Z",
      "content": "<p>Thanks a lot for sharing the article\nHow do you define  highly correlation between models .Is  it CV or Model architecture or something else?</p>",
      "rawMarkdown": "Thanks a lot for sharing the article\nHow do you define  highly correlation between models .Is  it CV or Model architecture or something else?"
    },
    {
      "id": 928929,
      "postDate": "2020-07-14T10:30:17.527Z",
      "content": "<p>Your 5th method is wrong. Don't you think it should be  ((pred1^p + pred2^p + .... predn^p)/n)^(1/n)\nIf it is not then you are basically changing the scale.</p>",
      "rawMarkdown": "Your 5th method is wrong. Don't you think it should be  ((pred1^p + pred2^p + .... predn^p)/n)^(1/n)\nIf it is not then you are basically changing the scale.",
      "replies": [
        {
          "id": 929171,
          "postDate": "2020-07-14T13:58:24.263Z",
          "content": "<p>Power averaging ensures that the top values are pushed further away from the rest, thereby <strong>separating</strong> Malignant images from Benign, resulting in a higher AUC/LB score.</p>\n\n<p>However, that does not mean that you cannot apply a square-root. In Data Science, it is all about the data i.e. depending on your models and predicted probabilities, you may get better results with a square-root vs someone else whose models work better when directly used with power values.</p>",
          "rawMarkdown": "Power averaging ensures that the top values are pushed further away from the rest, thereby **separating** Malignant images from Benign, resulting in a higher AUC/LB score.\n\nHowever, that does not mean that you cannot apply a square-root. In Data Science, it is all about the data i.e. depending on your models and predicted probabilities, you may get better results with a square-root vs someone else whose models work better when directly used with power values.",
          "votes": 2
        },
        {
          "id": 929832,
          "postDate": "2020-07-15T02:20:07.337Z",
          "content": "<p>For a metric like AUC where only the rank matters, applying functions that don't change the rank but only the scale like power or multiplications or divisions have absolutely no impact on the results. However, care must be taken while ensembling outputs from multiple models - a model with higher variation in output is most likely to supersede the other outputs and hence the outputs must be normalised before ensembling.</p>",
          "rawMarkdown": "For a metric like AUC where only the rank matters, applying functions that don't change the rank but only the scale like power or multiplications or divisions have absolutely no impact on the results. However, care must be taken while ensembling outputs from multiple models - a model with higher variation in output is most likely to supersede the other outputs and hence the outputs must be normalised before ensembling."
        },
        {
          "id": 961628,
          "postDate": "2020-08-07T10:56:50.290Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        },
        {
          "id": 961860,
          "postDate": "2020-08-07T15:07:40.583Z",
          "content": "<p><a href=\"/krisho007\">@krisho007</a> Can you show a small reproducible example illustrating your statement?</p>",
          "rawMarkdown": "@krisho007 Can you show a small reproducible example illustrating your statement?"
        }
      ]
    },
    {
      "id": 926101,
      "postDate": "2020-07-12T13:52:39.277Z",
      "content": "<p>Nice insight</p>",
      "rawMarkdown": "Nice insight"
    },
    {
      "id": 923270,
      "postDate": "2020-07-10T17:06:40.747Z",
      "content": "<p>Cool, thanks for sharing! 👍 Let me try this power averaging method! </p>",
      "rawMarkdown": "Cool, thanks for sharing! 👍 Let me try this power averaging method! "
    },
    {
      "id": 961622,
      "postDate": "2020-08-07T10:49:59.050Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 930093,
      "postDate": "2020-07-15T07:35:07.173Z",
      "content": "<p>thanks for your contibution</p>",
      "rawMarkdown": "thanks for your contibution"
    }
  ],
  "comments": [
    {
      "id": 961196,
      "author_name": "Redwan Sony",
      "author_url": "",
      "post_date": "2020-08-07T01:54:15.857000",
      "content": "<p>wow.. amazing technique..</p>\n<p>Actually I was looking for something like this for a long time. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 954676,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-08-02T01:53:49.433000",
      "content": "<p>Thank you for sharing this, I have to take notes and probably make a notebook on all this.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 925894,
      "author_name": "Rohit Agarwal",
      "author_url": "",
      "post_date": "2020-07-12T11:20:45.713000",
      "content": "<p>Thanks for sharing this <a href=\"/sirishks\">@sirishks</a>.</p>\n\n<p>I guess the theoretical validation of using rank averaging is a bit stronger than simple averaging since we are choosing the best epoch on AUC where probabilities by themselves are disregarded (only the order of probabilities is considered). However, theory hardly agrees with practice.</p>\n\n<p>How would you apply advanced concepts like weighted averaging or power averaging to rank averaging? Using power averaging on rank averaging by itself would potentially lead to disastrous effects since it would scale images with values closer to <code>1</code> much lesser than those close to <code>N</code>, whereas the metric confers equal weight to each image whether the output is on one end of the spectrum or the other. However, to think about it, even in the case of power averaging on the probabilities, the probabilities closer to <code>1.00</code> are suppressed much lesser than those closer to <code>0.00</code>. But, probably that's a desirable quality. And since the higher p's provide better results, probably power averaging on rank averaging doesn't sound that bad an idea after all.</p>\n\n<p>What do you think?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 923129,
      "author_name": "Abhishek Bhat",
      "author_url": "",
      "post_date": "2020-07-10T14:52:04.553000",
      "content": "<p>Thanks a lot for this <a href=\"/sirishks\">@sirishks</a>, very helpful!</p>\n\n<p>But we need to be <strong>extremely</strong> careful about choosing the models for power averaging, one bad model and everything might go for a toss.</p>\n\n<p>Also, in point 3 you mention:</p>\n\n<blockquote>\n  <p>Weighted Averaging: Specify weights, say 0.5 each in case of two models\n  WeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn) / n\n  where, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1</p>\n</blockquote>\n\n<p>Shouldn't we have <code>wt1+...+wtn = n</code>, since we are dividing the equation by <code>n</code> OR just remove <code>n</code> from the denominator and have <code>wt1+...+wtn = 1</code>? Correct me if I am wrong.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 923160,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-07-10T15:14:18.473000",
          "content": "<p><a href=\"/abhishekgbhat\">@abhishekgbhat</a> yes, it was a typo. You are correct👍 </p>\n\n<p>I have edited the answer and removed \"/n\" when using weights in the range [0,1] such that sum of weights = 1.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 923295,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-10T17:42:22.710000",
          "content": "<p><a href=\"/abhishekgbhat\">@abhishekgbhat</a> AUC is an interesting metric -- all it cares about is the order of your predictions. Dividing the predictions by any common number (like n) is not going to change the values of the metric, so, from practical standpoint, it doesn't make a difference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 923314,
          "author_name": "Abhishek Bhat",
          "author_url": "",
          "post_date": "2020-07-10T18:04:44.183000",
          "content": "<p><a href=\"/graf10a\">@graf10a</a> I totally agree with what you are saying about the AUC.. I was making a general comment about the weighted average definition. But as you mentioned, when the metric is AUC all that matters is the order.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3361837,
      "author_name": "Shivani Sharma",
      "author_url": "",
      "post_date": "2025-12-03T21:29:27.237000",
      "content": "<p>Well Written!</p>\n<p>Have a look at my notebook <a href=\"https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble\" target=\"_blank\">https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 969206,
      "author_name": "Haider Ali Shuvo",
      "author_url": "",
      "post_date": "2020-08-13T14:46:43.310000",
      "content": "<p>Nice Explanation </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951748,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2020-07-30T11:28:35.280000",
      "content": "<p><a href=\"/sirishks\">@sirishks</a> thanks for the post. In the mentioned blog, there is a term <code>weight_initializer</code>in part 5 of blog which I could not understand. Can you please clarify what is <code>weight_initializer</code>and how one can calculate it ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949974,
      "author_name": "sagarkarki142",
      "author_url": "",
      "post_date": "2020-07-29T05:19:57.040000",
      "content": "<p>Thanks a lot for sharing the article\nHow do you define  highly correlation between models .Is  it CV or Model architecture or something else?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 928929,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T10:30:17.527000",
      "content": "<p>Your 5th method is wrong. Don't you think it should be  ((pred1^p + pred2^p + .... predn^p)/n)^(1/n)\nIf it is not then you are basically changing the scale.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 929171,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-07-14T13:58:24.263000",
          "content": "<p>Power averaging ensures that the top values are pushed further away from the rest, thereby <strong>separating</strong> Malignant images from Benign, resulting in a higher AUC/LB score.</p>\n\n<p>However, that does not mean that you cannot apply a square-root. In Data Science, it is all about the data i.e. depending on your models and predicted probabilities, you may get better results with a square-root vs someone else whose models work better when directly used with power values.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 929832,
          "author_name": "Rohit Agarwal",
          "author_url": "",
          "post_date": "2020-07-15T02:20:07.337000",
          "content": "<p>For a metric like AUC where only the rank matters, applying functions that don't change the rank but only the scale like power or multiplications or divisions have absolutely no impact on the results. However, care must be taken while ensembling outputs from multiple models - a model with higher variation in output is most likely to supersede the other outputs and hence the outputs must be normalised before ensembling.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 961628,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-07T10:56:50.290000",
          "content": "",
          "votes": -2,
          "replies": []
        },
        {
          "id": 961860,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-08-07T15:07:40.583000",
          "content": "<p><a href=\"/krisho007\">@krisho007</a> Can you show a small reproducible example illustrating your statement?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 926101,
      "author_name": "virendra yadav0207",
      "author_url": "",
      "post_date": "2020-07-12T13:52:39.277000",
      "content": "<p>Nice insight</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 923270,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-07-10T17:06:40.747000",
      "content": "<p>Cool, thanks for sharing! 👍 Let me try this power averaging method! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 961622,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-07T10:49:59.050000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 930093,
      "author_name": "SURJEET KUMAR",
      "author_url": "",
      "post_date": "2020-07-15T07:35:07.173000",
      "content": "<p>thanks for your contibution</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "922999": "In Chris's excellent post [CNN Input Size Explained](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147), he suggests ensembling models of different input sizes. For blending the predictions of all these models, see **Power Averaging** technique below, which was used successfully in older competitions to **obtain Medals** 😄 \n\nVarious Prediction Averaging/Blending Techniques:\n\n1. Simple Averaging: Most participants are using just a simple mean of predictions generated by different models\n\n2. Rank Averaging: Use the \"rank\" of an input image instead it's prediction value. See public notebook [Improve blending using Rankdata](https://www.kaggle.com/niteshx2/improve-blending-using-rankdata)\n\n3. Weighted Averaging: Specify weights, say 0.5 each in case of two models\n   WeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn)\n   where, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1\n\n4. Stretch Averaging: Stretch predictions using min and max values first, before averaging\n   Pred = (Pred - min(Pred)) / (max(Pred) - min(Pred))\n\n5. Power Averaging: Choose a power p = 2, 4, 8, 16\n   PowerAverage(p) = (Pred1^p + Pred2^p + … + Predn^p) / n\n   Note: Power Averaging to be used only when all the models are highly correlated, otherwise your score may become worse.\n\n6. Power Averaging with weights:\n   PowerAverageWithWeights(p) = (wt1 x Pred1^p + wt2 x Pred2^p + … + wtn x Predn^p)\n\nNote: Ensemble your **best models** for final submission, typically about 2 weeks before the competition deadline.\n\nPS: Read this post by a Kaggle Master [Power/geometric ensembling when targeting the AUC metric](https://medium.com/data-design/reaching-the-depths-of-power-geometric-ensembling-when-targeting-the-auc-metric-2f356ea3250e) for a more detailed explanation with a lot of calculations.",
    "961196": "wow.. amazing technique..\n\nActually I was looking for something like this for a long time. \n",
    "954676": "Thank you for sharing this, I have to take notes and probably make a notebook on all this.",
    "925894": "Thanks for sharing this @sirishks.\n\nI guess the theoretical validation of using rank averaging is a bit stronger than simple averaging since we are choosing the best epoch on AUC where probabilities by themselves are disregarded (only the order of probabilities is considered). However, theory hardly agrees with practice.\n\nHow would you apply advanced concepts like weighted averaging or power averaging to rank averaging? Using power averaging on rank averaging by itself would potentially lead to disastrous effects since it would scale images with values closer to `1` much lesser than those close to `N`, whereas the metric confers equal weight to each image whether the output is on one end of the spectrum or the other. However, to think about it, even in the case of power averaging on the probabilities, the probabilities closer to `1.00` are suppressed much lesser than those closer to `0.00`. But, probably that's a desirable quality. And since the higher p's provide better results, probably power averaging on rank averaging doesn't sound that bad an idea after all.\n\nWhat do you think?",
    "923129": "Thanks a lot for this @sirishks, very helpful!\n\nBut we need to be **extremely** careful about choosing the models for power averaging, one bad model and everything might go for a toss.\n\nAlso, in point 3 you mention:\n&gt;Weighted Averaging: Specify weights, say 0.5 each in case of two models\nWeightedAverage(p) = (wt1 x Pred1 + wt2 x Pred2 + … + wtn x Predn) / n\nwhere, n is the number of models, and sum of weights wt1+wt2+…+wtn = 1\n\nShouldn't we have `wt1+...+wtn = n`, since we are dividing the equation by `n` OR just remove `n` from the denominator and have `wt1+...+wtn = 1`? Correct me if I am wrong.",
    "3361837": "Well Written!\n\n Have a look at my notebook https://www.kaggle.com/code/shivanisharma1297/ai-text-detection-myensemble",
    "969206": "Nice Explanation ",
    "951748": "@sirishks thanks for the post. In the mentioned blog, there is a term `weight_initializer `in part 5 of blog which I could not understand. Can you please clarify what is `weight_initializer `and how one can calculate it ?",
    "949974": "Thanks a lot for sharing the article\nHow do you define  highly correlation between models .Is  it CV or Model architecture or something else?",
    "928929": "Your 5th method is wrong. Don't you think it should be  ((pred1^p + pred2^p + .... predn^p)/n)^(1/n)\nIf it is not then you are basically changing the scale.",
    "926101": "Nice insight",
    "923270": "Cool, thanks for sharing! 👍 Let me try this power averaging method! ",
    "961622": "",
    "930093": "thanks for your contibution"
  }
}