{
  "id": 171124,
  "title": "ensemble for AUC?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171124",
  "author_name": "",
  "post_date": "2020-07-30T14:18:00.732628600Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>i haven't implement the method below, but it is just a \"though experiment\". you may want to think about it and see if it would help ...</p>\n\n<ol>\n<li><p>AUC is ranking. so absolute predictions scores are sometimes not useful?</p></li>\n<li><p>imagine we have trained a model. then we compute all pairwise distances.</p></li>\n<li><p>instance of averaging the score, we average the pairwise distances from different models. finally think of a way to map pairwise distances back to rank (e.g. quick sort).</p></li>\n</ol>\n\n<p>i wonder if it would help?</p>",
  "messages": [
    {
      "id": "951960",
      "postDate": "07/30/2020 14:18:00",
      "content": "<p>i haven't implement the method below, but it is just a \"though experiment\". you may want to think about it and see if it would help ...</p>\n\n<ol>\n<li><p>AUC is ranking. so absolute predictions scores are sometimes not useful?</p></li>\n<li><p>imagine we have trained a model. then we compute all pairwise distances.</p></li>\n<li><p>instance of averaging the score, we average the pairwise distances from different models. finally think of a way to map pairwise distances back to rank (e.g. quick sort).</p></li>\n</ol>\n\n<p>i wonder if it would help?</p>",
      "rawMarkdown": "i haven't implement the method below, but it is just a \"though experiment\". you may want to think about it and see if it would help ...\n\n1. AUC is ranking. so absolute predictions scores are sometimes not useful?\n\n2. imagine we have trained a model. then we compute all pairwise distances.\n\n3. instance of averaging the score, we average the pairwise distances from different models. finally think of a way to map pairwise distances back to rank (e.g. quick sort).\n\ni wonder if it would help?",
      "votes": null
    },
    {
      "id": "951961",
      "postDate": "07/30/2020 14:20:08",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> did you check my post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">For AUC Metric, Ensemble using Power Averaging</a></p>",
      "rawMarkdown": "hengck23 did you check my post [For AUC Metric, Ensemble using Power Averaging](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653)",
      "votes": null
    },
    {
      "id": "952162",
      "postDate": "07/30/2020 17:07:52",
      "content": "<p>It seems that such an approach can be really tricky, because different models and especially losses may produce different distributions. So, for example, one model can be trained with BCE and another one with BCE and label smotthing. They may produce good ranking but the first will have bigger pairwise distances. So after averaging of distances they will be messed up.\nAll in all, using BCE and Sigmoid we can introduce probabilities like distances to ONE or ZERO class. And using ranked predictions in blends seems to produce better results (according to high score public blends). So we can make a conclusion that distances could be not the best option for AUC metric.</p>",
      "rawMarkdown": "It seems that such an approach can be really tricky, because different models and especially losses may produce different distributions. So, for example, one model can be trained with BCE and another one with BCE and label smotthing. They may produce good ranking but the first will have bigger pairwise distances. So after averaging of distances they will be messed up.\nAll in all, using BCE and Sigmoid we can introduce probabilities like distances to ONE or ZERO class. And using ranked predictions in blends seems to produce better results (according to high score public blends). So we can make a conclusion that distances could be not the best option for AUC metric.",
      "votes": null
    },
    {
      "id": "952483",
      "postDate": "07/31/2020 00:57:25",
      "content": "<p><a href=\"/vladimirsydor\">@vladimirsydor</a>  and <a href=\"/sirishks\">@sirishks</a> </p>\n\n<p>Thanks for the post. </p>\n\n<p>\"so after averaging of distances they will be messed up.\" ... that could be true.</p>\n\n<p>\"Sigmoid we can introduce probabilities like distances to ONE or ZERO class\" ... i wonder if distance to more \"prototype\" would help? </p>\n\n<p>\"And using ranked predictions in blends seems to produce better results (according to high score public blends)\" I tried using rankdata., but it gives mixed results for me. Some distance information is lost (or corrected) when converting probability score to rank.</p>\n\n<p>\"my post For AUC Metric, Ensemble using Power Averaging\"  I have been using temperature scaling (raise power to probability score before, it is actually a trick in classification challenge as well). Again results are mixed as it really depends on the data and quality of the models.</p>\n\n<hr>\n\n<p>I think the best solution is to trained another stacking model to combine the base models. But this has another problem ... lack of new train data for learning stacking model</p>",
      "rawMarkdown": "vladimirsydor  and @sirishks \n\nThanks for the post. \n\n\"so after averaging of distances they will be messed up.\" ... that could be true.\n\n\"Sigmoid we can introduce probabilities like distances to ONE or ZERO class\" ... i wonder if distance to more \"prototype\" would help? \n\n\"And using ranked predictions in blends seems to produce better results (according to high score public blends)\" I tried using rankdata., but it gives mixed results for me. Some distance information is lost (or corrected) when converting probability score to rank.\n\n \n\"my post For AUC Metric, Ensemble using Power Averaging\"  I have been using temperature scaling (raise power to probability score before, it is actually a trick in classification challenge as well). Again results are mixed as it really depends on the data and quality of the models.\n\n---\n\nI think the best solution is to trained another stacking model to combine the base models. But this has another problem ... lack of new train data for learning stacking model",
      "votes": null
    },
    {
      "id": "952515",
      "postDate": "07/31/2020 02:16:31",
      "content": "<p>and one comment on label smoothing and temperature scaling (raise power to probability score):</p>\n\n<p>```\nnormal binary cross entropy <br>\n= -truth*log_p - (1-truth)*(1-log_p)\n= -truth*log_pos - (1-truth)*log_neg\n= -log_pos**(truth) - log_neg**(1-truth)</p>\n\n<p>smooth binary cross entropy \n= -(0.9*truth)*log_p - (1-truth*0.9)*(1-log_p)\n= -log_pos**(0.9*truth) - log_neg**(1-truth*0.9)</p>\n\n<p>```</p>",
      "rawMarkdown": "and one comment on label smoothing and temperature scaling (raise power to probability score):\n\n```\nnormal binary cross entropy  \n= -truth*log_p - (1-truth)*(1-log_p)\n= -truth*log_pos - (1-truth)*log_neg\n= -log_pos**(truth) - log_neg**(1-truth)\n\n\nsmooth binary cross entropy \n= -(0.9*truth)*log_p - (1-truth*0.9)*(1-log_p)\n= -log_pos**(0.9*truth) - log_neg**(1-truth*0.9)\n\n```",
      "votes": null
    },
    {
      "id": "952579",
      "postDate": "07/31/2020 04:06:56",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> Instead of using predictions/rankings from different models directly, perhaps we can use t-SNE (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">see Chris' post</a>) and compute 2-D pairwise distances.</p>\n\n<p><a href=\"/vladimirsydor\">@vladimirsydor</a>  This approach (t-SNE) preserves the pairwise distances regardless of the model we use i.e. say with/without label-smoothing, or even different models trained with different backbones.</p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "hengck23 Instead of using predictions/rankings from different models directly, perhaps we can use t-SNE ([see Chris' post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028)) and compute 2-D pairwise distances.\n\n@vladimirsydor  This approach (t-SNE) preserves the pairwise distances regardless of the model we use i.e. say with/without label-smoothing, or even different models trained with different backbones.\n\nWhat do you think?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 951961,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "07/30/2020 14:20:08",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> did you check my post <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">For AUC Metric, Ensemble using Power Averaging</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 952162,
      "author_name": "vladimirsydor",
      "author_url": "",
      "post_date": "07/30/2020 17:07:52",
      "content": "<p>It seems that such an approach can be really tricky, because different models and especially losses may produce different distributions. So, for example, one model can be trained with BCE and another one with BCE and label smotthing. They may produce good ranking but the first will have bigger pairwise distances. So after averaging of distances they will be messed up.\nAll in all, using BCE and Sigmoid we can introduce probabilities like distances to ONE or ZERO class. And using ranked predictions in blends seems to produce better results (according to high score public blends). So we can make a conclusion that distances could be not the best option for AUC metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 952579,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/31/2020 04:06:56",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Instead of using predictions/rankings from different models directly, perhaps we can use t-SNE (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028\">see Chris' post</a>) and compute 2-D pairwise distances.</p>\n\n<p><a href=\"/vladimirsydor\">@vladimirsydor</a>  This approach (t-SNE) preserves the pairwise distances regardless of the model we use i.e. say with/without label-smoothing, or even different models trained with different backbones.</p>\n\n<p>What do you think?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 952483,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2020 00:57:25",
      "content": "<p><a href=\"/vladimirsydor\">@vladimirsydor</a>  and <a href=\"/sirishks\">@sirishks</a> </p>\n\n<p>Thanks for the post. </p>\n\n<p>\"so after averaging of distances they will be messed up.\" ... that could be true.</p>\n\n<p>\"Sigmoid we can introduce probabilities like distances to ONE or ZERO class\" ... i wonder if distance to more \"prototype\" would help? </p>\n\n<p>\"And using ranked predictions in blends seems to produce better results (according to high score public blends)\" I tried using rankdata., but it gives mixed results for me. Some distance information is lost (or corrected) when converting probability score to rank.</p>\n\n<p>\"my post For AUC Metric, Ensemble using Power Averaging\"  I have been using temperature scaling (raise power to probability score before, it is actually a trick in classification challenge as well). Again results are mixed as it really depends on the data and quality of the models.</p>\n\n<hr>\n\n<p>I think the best solution is to trained another stacking model to combine the base models. But this has another problem ... lack of new train data for learning stacking model</p>",
      "votes": null,
      "replies": [
        {
          "id": 952515,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2020 02:16:31",
          "content": "<p>and one comment on label smoothing and temperature scaling (raise power to probability score):</p>\n\n<p>```\nnormal binary cross entropy <br>\n= -truth*log_p - (1-truth)*(1-log_p)\n= -truth*log_pos - (1-truth)*log_neg\n= -log_pos**(truth) - log_neg**(1-truth)</p>\n\n<p>smooth binary cross entropy \n= -(0.9*truth)*log_p - (1-truth*0.9)*(1-log_p)\n= -log_pos**(0.9*truth) - log_neg**(1-truth*0.9)</p>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "951960": "i haven't implement the method below, but it is just a \"though experiment\". you may want to think about it and see if it would help ...\n\n1. AUC is ranking. so absolute predictions scores are sometimes not useful?\n\n2. imagine we have trained a model. then we compute all pairwise distances.\n\n3. instance of averaging the score, we average the pairwise distances from different models. finally think of a way to map pairwise distances back to rank (e.g. quick sort).\n\ni wonder if it would help?",
    "951961": "hengck23 did you check my post [For AUC Metric, Ensemble using Power Averaging](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653)",
    "952162": "It seems that such an approach can be really tricky, because different models and especially losses may produce different distributions. So, for example, one model can be trained with BCE and another one with BCE and label smotthing. They may produce good ranking but the first will have bigger pairwise distances. So after averaging of distances they will be messed up.\nAll in all, using BCE and Sigmoid we can introduce probabilities like distances to ONE or ZERO class. And using ranked predictions in blends seems to produce better results (according to high score public blends). So we can make a conclusion that distances could be not the best option for AUC metric.",
    "952483": "vladimirsydor  and @sirishks \n\nThanks for the post. \n\n\"so after averaging of distances they will be messed up.\" ... that could be true.\n\n\"Sigmoid we can introduce probabilities like distances to ONE or ZERO class\" ... i wonder if distance to more \"prototype\" would help? \n\n\"And using ranked predictions in blends seems to produce better results (according to high score public blends)\" I tried using rankdata., but it gives mixed results for me. Some distance information is lost (or corrected) when converting probability score to rank.\n\n \n\"my post For AUC Metric, Ensemble using Power Averaging\"  I have been using temperature scaling (raise power to probability score before, it is actually a trick in classification challenge as well). Again results are mixed as it really depends on the data and quality of the models.\n\n---\n\nI think the best solution is to trained another stacking model to combine the base models. But this has another problem ... lack of new train data for learning stacking model",
    "952515": "and one comment on label smoothing and temperature scaling (raise power to probability score):\n\n```\nnormal binary cross entropy  \n= -truth*log_p - (1-truth)*(1-log_p)\n= -truth*log_pos - (1-truth)*log_neg\n= -log_pos**(truth) - log_neg**(1-truth)\n\n\nsmooth binary cross entropy \n= -(0.9*truth)*log_p - (1-truth*0.9)*(1-log_p)\n= -log_pos**(0.9*truth) - log_neg**(1-truth*0.9)\n\n```",
    "952579": "hengck23 Instead of using predictions/rankings from different models directly, perhaps we can use t-SNE ([see Chris' post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168028)) and compute 2-D pairwise distances.\n\n@vladimirsydor  This approach (t-SNE) preserves the pairwise distances regardless of the model we use i.e. say with/without label-smoothing, or even different models trained with different backbones.\n\nWhat do you think?"
  },
  "source": "meta"
}