{
  "id": 251004,
  "title": "ROC-Star loss, maximizing AUC directly",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/251004",
  "author_name": "",
  "post_date": "2021-07-05T12:37:45.326298500Z",
  "votes": 59,
  "comment_count": 12,
  "views": 0,
  "content": "<p><img src=\"https://raw.githubusercontent.com/iridiumblue/about/master/newplot.png\" alt=\"\"><br>\nPaper: <a href=\"https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf\" target=\"_blank\">https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf</a><br>\nGithub: <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\">https://github.com/iridiumblue/roc-star</a></p>\n<p>ROC-Star uses an approximation of the Wilcoxon-Mann-Whitney Statistic in order to train on AUC more directly than BCE. The paper achieved better AUC when comparing to BCE on average, that seems to be also the case from my experiment:</p>\n<p>I've adapted the github repo above to my notebook below to run on my training loop on Pytorch. It also performed better than BCE as the paper suggested with the comparison below:</p>\n<p>Same Params:<br>\nBCE: LB 0.829 (<a href=\"https://www.kaggle.com/capiru/g2net-starter-easy-training-model-gpu-tpu?scriptVersionId=67471237\" target=\"_blank\">https://www.kaggle.com/capiru/g2net-starter-easy-training-model-gpu-tpu?scriptVersionId=67471237</a>)<br>\nROC-Star: LB 0.836 (<a href=\"https://www.kaggle.com/capiru/g2net-roc-star-maximize-auc\" target=\"_blank\">https://www.kaggle.com/capiru/g2net-roc-star-maximize-auc</a>)</p>",
  "messages": [
    {
      "id": "1376920",
      "postDate": "07/05/2021 12:37:45",
      "content": "<p><img src=\"https://raw.githubusercontent.com/iridiumblue/about/master/newplot.png\" alt=\"\"><br>\nPaper: <a href=\"https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf\" target=\"_blank\">https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf</a><br>\nGithub: <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\">https://github.com/iridiumblue/roc-star</a></p>\n<p>ROC-Star uses an approximation of the Wilcoxon-Mann-Whitney Statistic in order to train on AUC more directly than BCE. The paper achieved better AUC when comparing to BCE on average, that seems to be also the case from my experiment:</p>\n<p>I've adapted the github repo above to my notebook below to run on my training loop on Pytorch. It also performed better than BCE as the paper suggested with the comparison below:</p>\n<p>Same Params:<br>\nBCE: LB 0.829 (<a href=\"https://www.kaggle.com/capiru/g2net-starter-easy-training-model-gpu-tpu?scriptVersionId=67471237\" target=\"_blank\">https://www.kaggle.com/capiru/g2net-starter-easy-training-model-gpu-tpu?scriptVersionId=67471237</a>)<br>\nROC-Star: LB 0.836 (<a href=\"https://www.kaggle.com/capiru/g2net-roc-star-maximize-auc\" target=\"_blank\">https://www.kaggle.com/capiru/g2net-roc-star-maximize-auc</a>)</p>",
      "rawMarkdown": "![](https://raw.githubusercontent.com/iridiumblue/about/master/newplot.png)\nPaper: https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf\nGithub: https://github.com/iridiumblue/roc-star\n\nROC-Star uses an approximation of the Wilcoxon-Mann-Whitney Statistic in order to train on AUC more directly than BCE. The paper achieved better AUC when comparing to BCE on average, that seems to be also the case from my experiment:\n\nI've adapted the github repo above to my notebook below to run on my training loop on Pytorch. It also performed better than BCE as the paper suggested with the comparison below:\n\nSame Params:\nBCE: LB 0.829 (https://www.kaggle.com/capiru/g2net-starter-easy-training-model-gpu-tpu?scriptVersionId=67471237)\nROC-Star: LB 0.836 (https://www.kaggle.com/capiru/g2net-roc-star-maximize-auc)",
      "votes": null
    },
    {
      "id": "1379514",
      "postDate": "07/07/2021 12:28:21",
      "content": "<p>Does it take about twice longer to run, as it seems from your notebooks?</p>",
      "rawMarkdown": "Does it take about twice longer to run, as it seems from your notebooks?",
      "votes": null
    },
    {
      "id": "1379532",
      "postDate": "07/07/2021 12:47:18",
      "content": "<p>The main difference i found was that it could train for more epochs, after the first minimum with BCE it wouldn't improve anymore, but with ROC-Star, it manages to keep finding new minima.<br>\nEpoch by Epoch comparison is:<br>\nBCE 14 min per epoch<br>\nROC-Star 16 min per epoch (14% slower)</p>",
      "rawMarkdown": "The main difference i found was that it could train for more epochs, after the first minimum with BCE it wouldn't improve anymore, but with ROC-Star, it manages to keep finding new minima.\nEpoch by Epoch comparison is:\nBCE 14 min per epoch\nROC-Star 16 min per epoch (14% slower)",
      "votes": null
    },
    {
      "id": "1379688",
      "postDate": "07/07/2021 14:37:20",
      "content": "<p>Got it, thanks!</p>",
      "rawMarkdown": "Got it, thanks!",
      "votes": null
    },
    {
      "id": "1382207",
      "postDate": "07/09/2021 15:31:15",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> , is the ensemble between BCE and ROC Star working better that either or worse?</p>",
      "rawMarkdown": "Thanks for sharing @capiru , is the ensemble between BCE and ROC Star working better that either or worse?",
      "votes": null
    },
    {
      "id": "1384761",
      "postDate": "07/12/2021 07:37:48",
      "content": "<p>good job! <br>\nimage_size=640<br>\nROC-Star loss<br>\nepochs=4<br>\n1fold(5fold)<br>\ncv0.8685<br>\nlb0.871</p>",
      "rawMarkdown": "good job! \nimage_size=640\nROC-Star loss\nepochs=4\n1fold(5fold)\ncv0.8685\nlb0.871",
      "votes": null
    },
    {
      "id": "1384763",
      "postDate": "07/12/2021 07:39:50",
      "content": "<p>thanks! seems that I can train 1-2 epochs more…</p>",
      "rawMarkdown": "thanks! seems that I can train 1-2 epochs more...",
      "votes": null
    },
    {
      "id": "1384981",
      "postDate": "07/12/2021 11:29:27",
      "content": "<p>Good compared to what?  How much boost do you get form roc star loss?</p>",
      "rawMarkdown": "Good compared to what?  How much boost do you get form roc star loss?",
      "votes": null
    },
    {
      "id": "1384993",
      "postDate": "07/12/2021 11:40:52",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> want to compare with my image (640) of the same size. They are all 5fold. One uses ROC star loss, and the other is not used. Unfortunately, there is a small problem in the code. After running 1fold, CUDA out of memory. The result needs to wait until tomorrow. </p>",
      "rawMarkdown": "cpmpml want to compare with my image (640) of the same size. They are all 5fold. One uses ROC star loss, and the other is not used. Unfortunately, there is a small problem in the code. After running 1fold, CUDA out of memory. The result needs to wait until tomorrow.",
      "votes": null
    },
    {
      "id": "1385900",
      "postDate": "07/13/2021 05:12:31",
      "content": "<p>After modifying the code, the result is 0.872. It seems that this loss function is not very obvious<br>\nimage_size=640<br>\nROC-Star loss<br>\nepochs=4<br>\n5fold(5fold)<br>\ncv0.86828<br>\nlb0.872</p>\n<p>When this loss function is used, the larger the image is, the smaller the improvement of CV is. Until 640, the improvement of CV is too small to change lb</p>",
      "rawMarkdown": "After modifying the code, the result is 0.872. It seems that this loss function is not very obvious\nimage_size=640\nROC-Star loss\nepochs=4\n5fold(5fold)\ncv0.86828\nlb0.872\n\nWhen this loss function is used, the larger the image is, the smaller the improvement of CV is. Until 640, the improvement of CV is too small to change lb",
      "votes": null
    },
    {
      "id": "1391783",
      "postDate": "07/18/2021 04:45:53",
      "content": "<p>Thanks a lot for sharing! How much longer do you need to train with this loss function? </p>",
      "rawMarkdown": "Thanks a lot for sharing! How much longer do you need to train with this loss function?",
      "votes": null
    },
    {
      "id": "1400475",
      "postDate": "07/26/2021 10:16:09",
      "content": "<p>Thanks for sharing! It worked for me! I didn't need longer time per epoch for learning.</p>",
      "rawMarkdown": "Thanks for sharing! It worked for me! I didn't need longer time per epoch for learning.",
      "votes": null
    },
    {
      "id": "1561281",
      "postDate": "10/27/2021 13:12:06",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1379514,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "07/07/2021 12:28:21",
      "content": "<p>Does it take about twice longer to run, as it seems from your notebooks?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1379532,
          "author_name": "capiru",
          "author_url": "",
          "post_date": "07/07/2021 12:47:18",
          "content": "<p>The main difference i found was that it could train for more epochs, after the first minimum with BCE it wouldn't improve anymore, but with ROC-Star, it manages to keep finding new minima.<br>\nEpoch by Epoch comparison is:<br>\nBCE 14 min per epoch<br>\nROC-Star 16 min per epoch (14% slower)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1379688,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "07/07/2021 14:37:20",
          "content": "<p>Got it, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1384763,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "07/12/2021 07:39:50",
          "content": "<p>thanks! seems that I can train 1-2 epochs more…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1382207,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "07/09/2021 15:31:15",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> , is the ensemble between BCE and ROC Star working better that either or worse?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1384761,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "07/12/2021 07:37:48",
      "content": "<p>good job! <br>\nimage_size=640<br>\nROC-Star loss<br>\nepochs=4<br>\n1fold(5fold)<br>\ncv0.8685<br>\nlb0.871</p>",
      "votes": null,
      "replies": [
        {
          "id": 1384981,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/12/2021 11:29:27",
          "content": "<p>Good compared to what?  How much boost do you get form roc star loss?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1384993,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "07/12/2021 11:40:52",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> want to compare with my image (640) of the same size. They are all 5fold. One uses ROC star loss, and the other is not used. Unfortunately, there is a small problem in the code. After running 1fold, CUDA out of memory. The result needs to wait until tomorrow. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1385900,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "07/13/2021 05:12:31",
          "content": "<p>After modifying the code, the result is 0.872. It seems that this loss function is not very obvious<br>\nimage_size=640<br>\nROC-Star loss<br>\nepochs=4<br>\n5fold(5fold)<br>\ncv0.86828<br>\nlb0.872</p>\n<p>When this loss function is used, the larger the image is, the smaller the improvement of CV is. Until 640, the improvement of CV is too small to change lb</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1391783,
      "author_name": "richx86",
      "author_url": "",
      "post_date": "07/18/2021 04:45:53",
      "content": "<p>Thanks a lot for sharing! How much longer do you need to train with this loss function? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1400475,
      "author_name": "yoshimasaizaki",
      "author_url": "",
      "post_date": "07/26/2021 10:16:09",
      "content": "<p>Thanks for sharing! It worked for me! I didn't need longer time per epoch for learning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1561281,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 13:12:06",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1376920": "![](https://raw.githubusercontent.com/iridiumblue/about/master/newplot.png)\nPaper: https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf\nGithub: https://github.com/iridiumblue/roc-star\n\nROC-Star uses an approximation of the Wilcoxon-Mann-Whitney Statistic in order to train on AUC more directly than BCE. The paper achieved better AUC when comparing to BCE on average, that seems to be also the case from my experiment:\n\nI've adapted the github repo above to my notebook below to run on my training loop on Pytorch. It also performed better than BCE as the paper suggested with the comparison below:\n\nSame Params:\nBCE: LB 0.829 (https://www.kaggle.com/capiru/g2net-starter-easy-training-model-gpu-tpu?scriptVersionId=67471237)\nROC-Star: LB 0.836 (https://www.kaggle.com/capiru/g2net-roc-star-maximize-auc)",
    "1379514": "Does it take about twice longer to run, as it seems from your notebooks?",
    "1379532": "The main difference i found was that it could train for more epochs, after the first minimum with BCE it wouldn't improve anymore, but with ROC-Star, it manages to keep finding new minima.\nEpoch by Epoch comparison is:\nBCE 14 min per epoch\nROC-Star 16 min per epoch (14% slower)",
    "1379688": "Got it, thanks!",
    "1382207": "Thanks for sharing @capiru , is the ensemble between BCE and ROC Star working better that either or worse?",
    "1384761": "good job! \nimage_size=640\nROC-Star loss\nepochs=4\n1fold(5fold)\ncv0.8685\nlb0.871",
    "1384763": "thanks! seems that I can train 1-2 epochs more...",
    "1384981": "Good compared to what?  How much boost do you get form roc star loss?",
    "1384993": "cpmpml want to compare with my image (640) of the same size. They are all 5fold. One uses ROC star loss, and the other is not used. Unfortunately, there is a small problem in the code. After running 1fold, CUDA out of memory. The result needs to wait until tomorrow.",
    "1385900": "After modifying the code, the result is 0.872. It seems that this loss function is not very obvious\nimage_size=640\nROC-Star loss\nepochs=4\n5fold(5fold)\ncv0.86828\nlb0.872\n\nWhen this loss function is used, the larger the image is, the smaller the improvement of CV is. Until 640, the improvement of CV is too small to change lb",
    "1391783": "Thanks a lot for sharing! How much longer do you need to train with this loss function?",
    "1400475": "Thanks for sharing! It worked for me! I didn't need longer time per epoch for learning.",
    "1561281": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}