{
  "id": 171331,
  "title": "Simple average may be better than rank average",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171331",
  "author_name": "",
  "post_date": "2020-07-31T11:34:14.088321100Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>When ensembling some models, I tried both of simple average and rank average. Mysteriously, although rank average almost always improves local cv better than a simple average, the LB of simple average is always higher than rank average. I carefully checked the distributions and concluded that the simple average is better for the following reason.</p>\n\n<p>Rank average treats every model equally because it only cares about the order of prediction. However, given that test data is different from the train data which is highly imbalanced, we should focus on the models which are \"aggresive\".  Thus, I think we should use a simple average.</p>\n\n<p>I really want to hear your opinion!</p>\n\n<p>Thanks !</p>",
  "messages": [
    {
      "id": "952968",
      "postDate": "07/31/2020 11:34:14",
      "content": "<p>When ensembling some models, I tried both of simple average and rank average. Mysteriously, although rank average almost always improves local cv better than a simple average, the LB of simple average is always higher than rank average. I carefully checked the distributions and concluded that the simple average is better for the following reason.</p>\n\n<p>Rank average treats every model equally because it only cares about the order of prediction. However, given that test data is different from the train data which is highly imbalanced, we should focus on the models which are \"aggresive\".  Thus, I think we should use a simple average.</p>\n\n<p>I really want to hear your opinion!</p>\n\n<p>Thanks !</p>",
      "rawMarkdown": "When ensembling some models, I tried both of simple average and rank average. Mysteriously, although rank average almost always improves local cv better than a simple average, the LB of simple average is always higher than rank average. I carefully checked the distributions and concluded that the simple average is better for the following reason.\n\nRank average treats every model equally because it only cares about the order of prediction. However, given that test data is different from the train data which is highly imbalanced, we should focus on the models which are \"aggresive\".  Thus, I think we should use a simple average.\n\nI really want to hear your opinion!\n\nThanks !",
      "votes": null
    },
    {
      "id": "953205",
      "postDate": "07/31/2020 15:45:48",
      "content": "<p>In my experiments I found that using XGBoost trained on out of fold predictions give the best result, it is also an easy way to include meta-data</p>",
      "rawMarkdown": "In my experiments I found that using XGBoost trained on out of fold predictions give the best result, it is also an easy way to include meta-data",
      "votes": null
    },
    {
      "id": "953270",
      "postDate": "07/31/2020 16:36:51",
      "content": "<p>I can confirm. Had the same results, just play with the weights of image recognition and metadata for the result. I have the best results with 0.7 image recognition weight and 0.3 as metadata weight.</p>",
      "rawMarkdown": "I can confirm. Had the same results, just play with the weights of image recognition and metadata for the result. I have the best results with 0.7 image recognition weight and 0.3 as metadata weight.",
      "votes": null
    },
    {
      "id": "953274",
      "postDate": "07/31/2020 16:38:49",
      "content": "<p>I tried simple average aswell as power average (p=2, 4, 16), simple average is still my best result.</p>",
      "rawMarkdown": "I tried simple average aswell as power average (p=2, 4, 16), simple average is still my best result.",
      "votes": null
    },
    {
      "id": "953396",
      "postDate": "07/31/2020 18:49:00",
      "content": "<p>Same observation here !</p>",
      "rawMarkdown": "Same observation here !",
      "votes": null
    },
    {
      "id": "953508",
      "postDate": "07/31/2020 21:01:20",
      "content": "<p>I just thought about it and it also makes sense to use p=0.5 (or sqrt()) since the data is so imbalanced all the prediction values are stacked near 0, sqrt() spreads everything out. It increased my lb by 0.0004 which isn't much. I suggest you try it but it also depends on your model i think! let me know what your results are :)</p>",
      "rawMarkdown": "I just thought about it and it also makes sense to use p=0.5 (or sqrt()) since the data is so imbalanced all the prediction values are stacked near 0, sqrt() spreads everything out. It increased my lb by 0.0004 which isn't much. I suggest you try it but it also depends on your model i think! let me know what your results are :)",
      "votes": null
    },
    {
      "id": "953785",
      "postDate": "08/01/2020 05:18:15",
      "content": "<p>Thanks for your suggestion! I tried sqrt() and got the same results.</p>",
      "rawMarkdown": "Thanks for your suggestion! I tried sqrt() and got the same results.",
      "votes": null
    },
    {
      "id": "954397",
      "postDate": "08/01/2020 17:46:41",
      "content": "<p>rank is working better for me in public LB. Although, the shakeup can be great, so I recommend a simple average only.</p>",
      "rawMarkdown": "rank is working better for me in public LB. Although, the shakeup can be great, so I recommend a simple average only.",
      "votes": null
    },
    {
      "id": "956165",
      "postDate": "08/03/2020 09:56:52",
      "content": "<p>same</p>",
      "rawMarkdown": "same",
      "votes": null
    },
    {
      "id": "956236",
      "postDate": "08/03/2020 10:39:36",
      "content": "<p>Can you develop further more about that ? Are you saying you are using your deep learning models predictions as an input for a XGBoost classifier ? Or just blending metadata prediction using XGBoost and DL models predictions ?</p>",
      "rawMarkdown": "Can you develop further more about that ? Are you saying you are using your deep learning models predictions as an input for a XGBoost classifier ? Or just blending metadata prediction using XGBoost and DL models predictions ?",
      "votes": null
    },
    {
      "id": "956361",
      "postDate": "08/03/2020 12:43:11",
      "content": "<p>The second option is right. I train a XGBClassifier on Metadata let it predict the test images, let my image recognition model predict the test images and blend both results for each image together. It increases my LB Score significantly.</p>",
      "rawMarkdown": "The second option is right. I train a XGBClassifier on Metadata let it predict the test images, let my image recognition model predict the test images and blend both results for each image together. It increases my LB Score significantly.",
      "votes": null
    },
    {
      "id": "957385",
      "postDate": "08/04/2020 09:02:59",
      "content": "<p>My results a little better but when using Metadata I lost 0.20 on LB. I try to explain it the following way:</p>\n\n<p>Lets assume we have a target where we know the true_label is: 1 \nWe have a 5-Fold-Cross-Validation of 5 different models which predict the following:\nModel-1: 0.6\nModel-2: 0.65\nModel-3: 0.72\nModel-4: 0.85\nModel-5: 0.91</p>\n\n<p>Power-Averaging with a pow of 2 would deliever:\n(0.6^2+0.65^2+0.72^2+0.85^2+0.91^2)/5 = 0.5703</p>\n\n<p>With a pow of 0.5 it would be:\n(0.6^0.5+0.65^0.5+0.72^0.5+0.85^0.5+0.91^0.5)/5 = 0.8610</p>\n\n<p>Diff of both results:\n0.8610-0.5703 = 0.2907</p>\n\n<p>Here the pow of 0.5 is better. Lets assume now the true_label is: 0</p>\n\n<p>Same as above, we have a 5-Fold-Cross-Validation of 5 different models which predict the following: \nModel-1: 0.3\nModel-2: 0.25\nModel-3: 0.20\nModel-4: 0.17\nModel-5: 0.12</p>\n\n<p>Power-Averaging with a pow of 2 results in:\n(0.3^2+0.25^2+0.20^2+0.17^2+0.12^2)/5 = 0.04716</p>\n\n<p>With a pow of 0.5 it would be:\n(0.3^0.5+0.25^0.5+0.20^0.5+0.17^0.5+0.12^0.5)/5 = 0.45073137541</p>\n\n<p>Diff of both results:\n0.8610-0.5703 = ~0.40357</p>\n\n<p>In this case, pow of 2 is better. </p>\n\n<p>Is it known that the test-set has more benign images than maligant? We only know from the train-set that benign is way more represented than maligant, maybe focusing on the ones we can classify easily (benign images) is better than trying to reach higher scores for maligant images. Did you try it with metadata?</p>",
      "rawMarkdown": "My results a little better but when using Metadata I lost 0.20 on LB. I try to explain it the following way:\n\nLets assume we have a target where we know the true_label is: 1 \nWe have a 5-Fold-Cross-Validation of 5 different models which predict the following:\nModel-1: 0.6\nModel-2: 0.65\nModel-3: 0.72\nModel-4: 0.85\nModel-5: 0.91\n\nPower-Averaging with a pow of 2 would deliever:\n(0.6^2+0.65^2+0.72^2+0.85^2+0.91^2)/5 = 0.5703\n\nWith a pow of 0.5 it would be:\n(0.6^0.5+0.65^0.5+0.72^0.5+0.85^0.5+0.91^0.5)/5 = 0.8610\n\nDiff of both results:\n0.8610-0.5703 = 0.2907\n\nHere the pow of 0.5 is better. Lets assume now the true_label is: 0\n\n\nSame as above, we have a 5-Fold-Cross-Validation of 5 different models which predict the following: \nModel-1: 0.3\nModel-2: 0.25\nModel-3: 0.20\nModel-4: 0.17\nModel-5: 0.12\n\nPower-Averaging with a pow of 2 results in:\n(0.3^2+0.25^2+0.20^2+0.17^2+0.12^2)/5 = 0.04716\n\nWith a pow of 0.5 it would be:\n(0.3^0.5+0.25^0.5+0.20^0.5+0.17^0.5+0.12^0.5)/5 = 0.45073137541\n\nDiff of both results:\n0.8610-0.5703 = ~0.40357\n\nIn this case, pow of 2 is better. \n\nIs it known that the test-set has more benign images than maligant? We only know from the train-set that benign is way more represented than maligant, maybe focusing on the ones we can classify easily (benign images) is better than trying to reach higher scores for maligant images. Did you try it with metadata?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 953205,
      "author_name": "abiolatti",
      "author_url": "",
      "post_date": "07/31/2020 15:45:48",
      "content": "<p>In my experiments I found that using XGBoost trained on out of fold predictions give the best result, it is also an easy way to include meta-data</p>",
      "votes": null,
      "replies": [
        {
          "id": 953270,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "07/31/2020 16:36:51",
          "content": "<p>I can confirm. Had the same results, just play with the weights of image recognition and metadata for the result. I have the best results with 0.7 image recognition weight and 0.3 as metadata weight.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956165,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "08/03/2020 09:56:52",
          "content": "<p>same</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956236,
          "author_name": "romainfabre",
          "author_url": "",
          "post_date": "08/03/2020 10:39:36",
          "content": "<p>Can you develop further more about that ? Are you saying you are using your deep learning models predictions as an input for a XGBoost classifier ? Or just blending metadata prediction using XGBoost and DL models predictions ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956361,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "08/03/2020 12:43:11",
          "content": "<p>The second option is right. I train a XGBClassifier on Metadata let it predict the test images, let my image recognition model predict the test images and blend both results for each image together. It increases my LB Score significantly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 953274,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "07/31/2020 16:38:49",
      "content": "<p>I tried simple average aswell as power average (p=2, 4, 16), simple average is still my best result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 953508,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "07/31/2020 21:01:20",
          "content": "<p>I just thought about it and it also makes sense to use p=0.5 (or sqrt()) since the data is so imbalanced all the prediction values are stacked near 0, sqrt() spreads everything out. It increased my lb by 0.0004 which isn't much. I suggest you try it but it also depends on your model i think! let me know what your results are :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 953785,
          "author_name": "syumei",
          "author_url": "",
          "post_date": "08/01/2020 05:18:15",
          "content": "<p>Thanks for your suggestion! I tried sqrt() and got the same results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 957385,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "08/04/2020 09:02:59",
          "content": "<p>My results a little better but when using Metadata I lost 0.20 on LB. I try to explain it the following way:</p>\n\n<p>Lets assume we have a target where we know the true_label is: 1 \nWe have a 5-Fold-Cross-Validation of 5 different models which predict the following:\nModel-1: 0.6\nModel-2: 0.65\nModel-3: 0.72\nModel-4: 0.85\nModel-5: 0.91</p>\n\n<p>Power-Averaging with a pow of 2 would deliever:\n(0.6^2+0.65^2+0.72^2+0.85^2+0.91^2)/5 = 0.5703</p>\n\n<p>With a pow of 0.5 it would be:\n(0.6^0.5+0.65^0.5+0.72^0.5+0.85^0.5+0.91^0.5)/5 = 0.8610</p>\n\n<p>Diff of both results:\n0.8610-0.5703 = 0.2907</p>\n\n<p>Here the pow of 0.5 is better. Lets assume now the true_label is: 0</p>\n\n<p>Same as above, we have a 5-Fold-Cross-Validation of 5 different models which predict the following: \nModel-1: 0.3\nModel-2: 0.25\nModel-3: 0.20\nModel-4: 0.17\nModel-5: 0.12</p>\n\n<p>Power-Averaging with a pow of 2 results in:\n(0.3^2+0.25^2+0.20^2+0.17^2+0.12^2)/5 = 0.04716</p>\n\n<p>With a pow of 0.5 it would be:\n(0.3^0.5+0.25^0.5+0.20^0.5+0.17^0.5+0.12^0.5)/5 = 0.45073137541</p>\n\n<p>Diff of both results:\n0.8610-0.5703 = ~0.40357</p>\n\n<p>In this case, pow of 2 is better. </p>\n\n<p>Is it known that the test-set has more benign images than maligant? We only know from the train-set that benign is way more represented than maligant, maybe focusing on the ones we can classify easily (benign images) is better than trying to reach higher scores for maligant images. Did you try it with metadata?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 953396,
      "author_name": "romainfabre",
      "author_url": "",
      "post_date": "07/31/2020 18:49:00",
      "content": "<p>Same observation here !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 954397,
      "author_name": "yash612",
      "author_url": "",
      "post_date": "08/01/2020 17:46:41",
      "content": "<p>rank is working better for me in public LB. Although, the shakeup can be great, so I recommend a simple average only.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "952968": "When ensembling some models, I tried both of simple average and rank average. Mysteriously, although rank average almost always improves local cv better than a simple average, the LB of simple average is always higher than rank average. I carefully checked the distributions and concluded that the simple average is better for the following reason.\n\nRank average treats every model equally because it only cares about the order of prediction. However, given that test data is different from the train data which is highly imbalanced, we should focus on the models which are \"aggresive\".  Thus, I think we should use a simple average.\n\nI really want to hear your opinion!\n\nThanks !",
    "953205": "In my experiments I found that using XGBoost trained on out of fold predictions give the best result, it is also an easy way to include meta-data",
    "953270": "I can confirm. Had the same results, just play with the weights of image recognition and metadata for the result. I have the best results with 0.7 image recognition weight and 0.3 as metadata weight.",
    "953274": "I tried simple average aswell as power average (p=2, 4, 16), simple average is still my best result.",
    "953396": "Same observation here !",
    "953508": "I just thought about it and it also makes sense to use p=0.5 (or sqrt()) since the data is so imbalanced all the prediction values are stacked near 0, sqrt() spreads everything out. It increased my lb by 0.0004 which isn't much. I suggest you try it but it also depends on your model i think! let me know what your results are :)",
    "953785": "Thanks for your suggestion! I tried sqrt() and got the same results.",
    "954397": "rank is working better for me in public LB. Although, the shakeup can be great, so I recommend a simple average only.",
    "956165": "same",
    "956236": "Can you develop further more about that ? Are you saying you are using your deep learning models predictions as an input for a XGBoost classifier ? Or just blending metadata prediction using XGBoost and DL models predictions ?",
    "956361": "The second option is right. I train a XGBClassifier on Metadata let it predict the test images, let my image recognition model predict the test images and blend both results for each image together. It increases my LB Score significantly.",
    "957385": "My results a little better but when using Metadata I lost 0.20 on LB. I try to explain it the following way:\n\nLets assume we have a target where we know the true_label is: 1 \nWe have a 5-Fold-Cross-Validation of 5 different models which predict the following:\nModel-1: 0.6\nModel-2: 0.65\nModel-3: 0.72\nModel-4: 0.85\nModel-5: 0.91\n\nPower-Averaging with a pow of 2 would deliever:\n(0.6^2+0.65^2+0.72^2+0.85^2+0.91^2)/5 = 0.5703\n\nWith a pow of 0.5 it would be:\n(0.6^0.5+0.65^0.5+0.72^0.5+0.85^0.5+0.91^0.5)/5 = 0.8610\n\nDiff of both results:\n0.8610-0.5703 = 0.2907\n\nHere the pow of 0.5 is better. Lets assume now the true_label is: 0\n\n\nSame as above, we have a 5-Fold-Cross-Validation of 5 different models which predict the following: \nModel-1: 0.3\nModel-2: 0.25\nModel-3: 0.20\nModel-4: 0.17\nModel-5: 0.12\n\nPower-Averaging with a pow of 2 results in:\n(0.3^2+0.25^2+0.20^2+0.17^2+0.12^2)/5 = 0.04716\n\nWith a pow of 0.5 it would be:\n(0.3^0.5+0.25^0.5+0.20^0.5+0.17^0.5+0.12^0.5)/5 = 0.45073137541\n\nDiff of both results:\n0.8610-0.5703 = ~0.40357\n\nIn this case, pow of 2 is better. \n\nIs it known that the test-set has more benign images than maligant? We only know from the train-set that benign is way more represented than maligant, maybe focusing on the ones we can classify easily (benign images) is better than trying to reach higher scores for maligant images. Did you try it with metadata?"
  },
  "source": "meta"
}