{
  "id": 166754,
  "title": "roc_auc_score and sigmoid",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/166754",
  "author_name": "",
  "post_date": "2020-07-13T20:41:06.646109Z",
  "votes": 3,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Last layer in my model is nn.Linear.\nSo there is no sigmoid or anything else after Linear, the values can be negative.</p>\n\n<p>Then I use following loss function:\n```\nclass WeightedFocalLoss(nn.Module):\n    def <strong>init</strong>(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).<strong>init</strong>()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma</p>\n\n<pre><code>def forward(self, inputs, targets):\n    BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n    targets = targets.type(torch.long)\n    at = self.alpha.gather(0, targets.data.view(-1))\n    pt = torch.exp(-BCE_loss)\n    F_loss = at*(1-pt)**self.gamma * BCE_loss\n    return F_loss.mean()\n</code></pre>\n\n<p>```</p>\n\n<p>and then I calculate score with <code>roc_auc_score</code></p>\n\n<p>Question: how should I prepare submission?\nMy file looks like that:\n<code>\nimage_name,target\nISIC_0052060,0.012927517\nISIC_0052349,0.018632364\nISIC_0058510,0.029537585\nISIC_0073313,-0.00839683\nISIC_0073502,-0.02732876\n</code></p>\n\n<p>According to competition description \"For each image_name in the test set, you must predict the probability (target) that the sample is malignant.\"</p>\n\n<p>So probability should not be negative.</p>\n\n<p>Do you:\n- scale the prediction to fit inside 0.0 and 1.0?\n- process with sigmoid or something else?\n- use different network architecture so last layer is not just Linear?</p>",
  "messages": [
    {
      "id": "928269",
      "postDate": "07/13/2020 20:41:06",
      "content": "<p>Last layer in my model is nn.Linear.\nSo there is no sigmoid or anything else after Linear, the values can be negative.</p>\n\n<p>Then I use following loss function:\n```\nclass WeightedFocalLoss(nn.Module):\n    def <strong>init</strong>(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).<strong>init</strong>()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma</p>\n\n<pre><code>def forward(self, inputs, targets):\n    BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n    targets = targets.type(torch.long)\n    at = self.alpha.gather(0, targets.data.view(-1))\n    pt = torch.exp(-BCE_loss)\n    F_loss = at*(1-pt)**self.gamma * BCE_loss\n    return F_loss.mean()\n</code></pre>\n\n<p>```</p>\n\n<p>and then I calculate score with <code>roc_auc_score</code></p>\n\n<p>Question: how should I prepare submission?\nMy file looks like that:\n<code>\nimage_name,target\nISIC_0052060,0.012927517\nISIC_0052349,0.018632364\nISIC_0058510,0.029537585\nISIC_0073313,-0.00839683\nISIC_0073502,-0.02732876\n</code></p>\n\n<p>According to competition description \"For each image_name in the test set, you must predict the probability (target) that the sample is malignant.\"</p>\n\n<p>So probability should not be negative.</p>\n\n<p>Do you:\n- scale the prediction to fit inside 0.0 and 1.0?\n- process with sigmoid or something else?\n- use different network architecture so last layer is not just Linear?</p>",
      "rawMarkdown": "Last layer in my model is nn.Linear.\nSo there is no sigmoid or anything else after Linear, the values can be negative.\n\nThen I use following loss function:\n```\nclass WeightedFocalLoss(nn.Module):\n    def __init__(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).__init__()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma\n\n    def forward(self, inputs, targets):\n        BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        targets = targets.type(torch.long)\n        at = self.alpha.gather(0, targets.data.view(-1))\n        pt = torch.exp(-BCE_loss)\n        F_loss = at*(1-pt)**self.gamma * BCE_loss\n        return F_loss.mean()\n```\n\nand then I calculate score with `roc_auc_score`\n\nQuestion: how should I prepare submission?\nMy file looks like that:\n```\nimage_name,target\nISIC_0052060,0.012927517\nISIC_0052349,0.018632364\nISIC_0058510,0.029537585\nISIC_0073313,-0.00839683\nISIC_0073502,-0.02732876\n```\n\nAccording to competition description \"For each image_name in the test set, you must predict the probability (target) that the sample is malignant.\"\n\nSo probability should not be negative.\n\nDo you:\n- scale the prediction to fit inside 0.0 and 1.0?\n- process with sigmoid or something else?\n- use different network architecture so last layer is not just Linear?",
      "votes": null
    },
    {
      "id": "928291",
      "postDate": "07/13/2020 21:03:02",
      "content": "<p>Processing with Sigmoid seems to work pretty well for me. \nThis is a great question! I’m also interested to hear about other options. </p>",
      "rawMarkdown": "Processing with Sigmoid seems to work pretty well for me. \nThis is a great question! I’m also interested to hear about other options.",
      "votes": null
    },
    {
      "id": "928296",
      "postDate": "07/13/2020 21:11:34",
      "content": "<p>can you show beginning of your submission.csv? </p>\n\n<p>I think for ROC only order matters (not absolute value), that's why <code>roc_auc_score</code> works even for negative numbers but I am not sure what's the implementation in the competition</p>\n\n<p>I am still in the research phase when I check how different things work and now I want to compare some simple predictions so at this point I need to prepare submission code</p>",
      "rawMarkdown": "can you show beginning of your submission.csv? \n\nI think for ROC only order matters (not absolute value), that's why `roc_auc_score` works even for negative numbers but I am not sure what's the implementation in the competition\n\nI am still in the research phase when I check how different things work and now I want to compare some simple predictions so at this point I need to prepare submission code",
      "votes": null
    },
    {
      "id": "928338",
      "postDate": "07/13/2020 22:25:46",
      "content": "<p>The values are just between 0-1 like 0.021, 0.043, 0.53... and so on. </p>\n\n<p>Also, would be great if you add some credits for the loss function code you have borrowed. :) </p>",
      "rawMarkdown": "The values are just between 0-1 like 0.021, 0.043, 0.53... and so on. \n\nAlso, would be great if you add some credits for the loss function code you have borrowed. :)",
      "votes": null
    },
    {
      "id": "928341",
      "postDate": "07/13/2020 22:32:01",
      "content": "<p>:)</p>\n\n<p>```\nimport torch.nn as nn\nimport torch\nimport torch.nn.functional as F</p>\n\n<p>\"\"\"\nby Aman Arora\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/162035\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/162035</a>\n\"\"\"\nclass WeightedFocalLoss(nn.Module):\n    def <strong>init</strong>(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).<strong>init</strong>()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma</p>\n\n<pre><code>def forward(self, inputs, targets):\n    BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n    targets = targets.type(torch.long)\n    at = self.alpha.gather(0, targets.data.view(-1))\n    pt = torch.exp(-BCE_loss)\n    F_loss = at*(1-pt)**self.gamma * BCE_loss\n    return F_loss.mean()\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": ":)\n\n```\nimport torch.nn as nn\nimport torch\nimport torch.nn.functional as F\n\n\"\"\"\nby Aman Arora\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/162035\n\"\"\"\nclass WeightedFocalLoss(nn.Module):\n    def __init__(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).__init__()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma\n\n    def forward(self, inputs, targets):\n        BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        targets = targets.type(torch.long)\n        at = self.alpha.gather(0, targets.data.view(-1))\n        pt = torch.exp(-BCE_loss)\n        F_loss = at*(1-pt)**self.gamma * BCE_loss\n        return F_loss.mean()\n```",
      "votes": null
    },
    {
      "id": "928345",
      "postDate": "07/13/2020 22:36:03",
      "content": "<p>Thank you, now I feel guilty for asking - haha!</p>",
      "rawMarkdown": "Thank you, now I feel guilty for asking - haha!",
      "votes": null
    },
    {
      "id": "928348",
      "postDate": "07/13/2020 22:45:51",
      "content": "<p>if I win you will be in the published code, of course assuming focal loss will survive :)\nI switched from BCE to focal loss and now I research other areas, will back to losses later</p>",
      "rawMarkdown": "if I win you will be in the published code, of course assuming focal loss will survive :)\nI switched from BCE to focal loss and now I research other areas, will back to losses later",
      "votes": null
    },
    {
      "id": "928352",
      "postDate": "07/13/2020 22:57:39",
      "content": "<p>Thank you - appreciate it :) </p>\n\n<p>Just to mention, some have reported BCE with label smoothing 0.5 works better than FL. I wasn't able to find the exact post where this was discussed so might be worth trying that too :) </p>",
      "rawMarkdown": "Thank you - appreciate it :) \n\nJust to mention, some have reported BCE with label smoothing 0.5 works better than FL. I wasn't able to find the exact post where this was discussed so might be worth trying that too :)",
      "votes": null
    },
    {
      "id": "928366",
      "postDate": "07/13/2020 23:37:22",
      "content": "<p>I think all people use sigmoid to scale output 0 to 1. That works for me to get 0.951</p>",
      "rawMarkdown": "I think all people use sigmoid to scale output 0 to 1. That works for me to get 0.951",
      "votes": null
    },
    {
      "id": "928373",
      "postDate": "07/13/2020 23:52:45",
      "content": "<p>Yes I have seen these discussions but I will focus on it later, there is a lot to explore in this competition.</p>",
      "rawMarkdown": "Yes I have seen these discussions but I will focus on it later, there is a lot to explore in this competition.",
      "votes": null
    },
    {
      "id": "928431",
      "postDate": "07/14/2020 01:39:21",
      "content": "<p>It doesn't matter for AUC. As long as the order (ranking) of all the preds stays the same, the AUC will stay the same. (And i'm pretty sure you can submit the negative values to Kaggle). Just do the same thing to all your models before you ensemble them.</p>",
      "rawMarkdown": "It doesn't matter for AUC. As long as the order (ranking) of all the preds stays the same, the AUC will stay the same. (And i'm pretty sure you can submit the negative values to Kaggle). Just do the same thing to all your models before you ensemble them.",
      "votes": null
    },
    {
      "id": "928529",
      "postDate": "07/14/2020 04:16:10",
      "content": "<p>confirmed, submission with or without sigmoid gives same score, so even negative numbers are ok</p>\n\n<p>however there must be some bug in my code, I am doing some simple model test and single fold score was 0.906 and LB score is 0.858</p>",
      "rawMarkdown": "confirmed, submission with or without sigmoid gives same score, so even negative numbers are ok\n\nhowever there must be some bug in my code, I am doing some simple model test and single fold score was 0.906 and LB score is 0.858",
      "votes": null
    },
    {
      "id": "928565",
      "postDate": "07/14/2020 04:49:52",
      "content": "<p>It shouldn't make a difference. I'm submitting raw logits. If you plot a histogram of your out-of-fold logits and test prediction logits, you can easily see how the two distributions differ (at least in prediction)</p>",
      "rawMarkdown": "It shouldn't make a difference. I'm submitting raw logits. If you plot a histogram of your out-of-fold logits and test prediction logits, you can easily see how the two distributions differ (at least in prediction)",
      "votes": null
    },
    {
      "id": "928594",
      "postDate": "07/14/2020 05:03:34",
      "content": "<p>There a lot of duplicates in the dataset and also best to create folds split by patient_id too. </p>",
      "rawMarkdown": "There a lot of duplicates in the dataset and also best to create folds split by patient_id too.",
      "votes": null
    },
    {
      "id": "928615",
      "postDate": "07/14/2020 05:22:56",
      "content": "<p>this is split I use:</p>\n\n<p>```\n    patients = data[\"patient_id\"].unique()\n    from sklearn.model_selection import KFold</p>\n\n<pre><code>kf = KFold(n_splits=n_splits, random_state=1234, shuffle=True)\n\nfor i, (train_index, valid_index) in enumerate(kf.split(patients)):\n    if (i == fold):\n        patients_train, patients_valid = patients[train_index], patients[valid_index]\n\ndata_train = data.loc[data['patient_id'].isin(patients_train)]    \ndata_valid = data.loc[data['patient_id'].isin(patients_valid)] \n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "this is split I use:\n\n```\n    patients = data[\"patient_id\"].unique()\n    from sklearn.model_selection import KFold\n    \n    kf = KFold(n_splits=n_splits, random_state=1234, shuffle=True)\n    \n    for i, (train_index, valid_index) in enumerate(kf.split(patients)):\n        if (i == fold):\n            patients_train, patients_valid = patients[train_index], patients[valid_index]\n    \n    data_train = data.loc[data['patient_id'].isin(patients_train)]    \n    data_valid = data.loc[data['patient_id'].isin(patients_valid)] \n\n```",
      "votes": null
    },
    {
      "id": "928838",
      "postDate": "07/14/2020 08:55:54",
      "content": "<p><a href=\"/aroraaman\">@aroraaman</a> do you happen to have a working implementation of label smoothing  loss with you?</p>",
      "rawMarkdown": "aroraaman do you happen to have a working implementation of label smoothing  loss with you?",
      "votes": null
    },
    {
      "id": "929173",
      "postDate": "07/14/2020 14:00:28",
      "content": "<p>What's the purpose of <code>if (i == fold):</code> ? What is the variable <code>fold</code>?</p>",
      "rawMarkdown": "What's the purpose of `if (i == fold):` ? What is the variable `fold`?",
      "votes": null
    },
    {
      "id": "929194",
      "postDate": "07/14/2020 14:14:53",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> When I train model I do it for specific fold, so after 5 runs I have 5 models</p>\n\n<p>I submitted results for single model, maybe if I blend 5 models the CV-LB gap will be smaller, will need to check it because I wasn't able to find any leak yet</p>",
      "rawMarkdown": "cdeotte When I train model I do it for specific fold, so after 5 runs I have 5 models\n\nI submitted results for single model, maybe if I blend 5 models the CV-LB gap will be smaller, will need to check it because I wasn't able to find any leak yet",
      "votes": null
    },
    {
      "id": "929207",
      "postDate": "07/14/2020 14:26:16",
      "content": "<p>ah ok, that's smart. I also run experiments on the same fold. Before i go to bed, i put my code on an infinite loop training a list of different models on the same fold and logging the validation AUC.</p>\n\n<p>I don't see a leak in your validation setup. This comp is weird. I'm struggling to match my local validation score with LB too. (My LB is higher than my local val). Today i plan to do more train test EDA. Perhaps i'll run some adversarial and try to figure out what's going on.</p>",
      "rawMarkdown": "ah ok, that's smart. I also run experiments on the same fold. Before i go to bed, i put my code on an infinite loop training a list of different models on the same fold and logging the validation AUC.\n\nI don't see a leak in your validation setup. This comp is weird. I'm struggling to match my local validation score with LB too. (My LB is higher than my local val). Today i plan to do more train test EDA. Perhaps i'll run some adversarial and try to figure out what's going on.",
      "votes": null
    },
    {
      "id": "929236",
      "postDate": "07/14/2020 14:48:57",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Do you train on GPU? Because if you are training on your computer I assume you can't use TPU? I am doing experiments on very low resolution right now. I did research on TPU as you have seen on the forum, but I am not using it to train my models now.</p>",
      "rawMarkdown": "cdeotte Do you train on GPU? Because if you are training on your computer I assume you can't use TPU? I am doing experiments on very low resolution right now. I did research on TPU as you have seen on the forum, but I am not using it to train my models now.",
      "votes": null
    },
    {
      "id": "929247",
      "postDate": "07/14/2020 14:55:34",
      "content": "<p>Yes i train on GPU. If you use multiple GPUs with mixed precision, you can run experiments as fast or faster than TPU.</p>\n\n<p>I use my <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">public notebook</a> and add </p>\n\n<pre><code>DEVICE = \"GPU\"\nstrategy = tf.distribute.MirroredStrategy()\ntf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})\n</code></pre>\n\n<p>And I run one fold over and over.</p>",
      "rawMarkdown": "Yes i train on GPU. If you use multiple GPUs with mixed precision, you can run experiments as fast or faster than TPU.\n\nI use my [public notebook][1] and add \n\n    DEVICE = \"GPU\"\n    strategy = tf.distribute.MirroredStrategy()\n    tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})\n\nAnd I run one fold over and over.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
      "votes": null
    },
    {
      "id": "929267",
      "postDate": "07/14/2020 15:14:24",
      "content": "<p>not sure what mixed precision is, is it like half float in pytorch?\nI have only one GPU, not sure should I buy more because it would be used only for Kaggle</p>",
      "rawMarkdown": "not sure what mixed precision is, is it like half float in pytorch?\nI have only one GPU, not sure should I buy more because it would be used only for Kaggle",
      "votes": null
    },
    {
      "id": "929286",
      "postDate": "07/14/2020 15:31:52",
      "content": "<p>Yes half float. Mixed uses some <code>float32</code> and some <code>float16</code>. It accelerates computation and allows more data to fit in GPU RAM. </p>",
      "rawMarkdown": "Yes half float. Mixed uses some `float32` and some `float16`. It accelerates computation and allows more data to fit in GPU RAM.",
      "votes": null
    },
    {
      "id": "929299",
      "postDate": "07/14/2020 15:39:12",
      "content": "<p>I tried that in previous competition but failed, will need to check it again here.</p>",
      "rawMarkdown": "I tried that in previous competition but failed, will need to check it again here.",
      "votes": null
    },
    {
      "id": "934432",
      "postDate": "07/18/2020 12:35:04",
      "content": "<p>sigmoid help  you to scale output to 0-1. Your target should be sigmoid(Linear()). Your target now not use sigmoid so it can be another.</p>",
      "rawMarkdown": "sigmoid help  you to scale output to 0-1. Your target should be sigmoid(Linear()). Your target now not use sigmoid so it can be another.",
      "votes": null
    },
    {
      "id": "947435",
      "postDate": "07/27/2020 09:10:47",
      "content": "<p>Do you use threshold for this competition e.g. a threshold=0.5 you smooth a value of 0.7 to 1 and submit it like that in your cv?</p>",
      "rawMarkdown": "Do you use threshold for this competition e.g. a threshold=0.5 you smooth a value of 0.7 to 1 and submit it like that in your cv?",
      "votes": null
    },
    {
      "id": "959799",
      "postDate": "08/05/2020 22:43:26",
      "content": "<p><a href=\"/aroraaman\">@aroraaman</a> are you using label smoothing? Also I think you mean't .05 not .5 but correct me if I am wrong.  </p>",
      "rawMarkdown": "aroraaman are you using label smoothing? Also I think you mean't .05 not .5 but correct me if I am wrong.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 928291,
      "author_name": "aroraaman",
      "author_url": "",
      "post_date": "07/13/2020 21:03:02",
      "content": "<p>Processing with Sigmoid seems to work pretty well for me. \nThis is a great question! I’m also interested to hear about other options. </p>",
      "votes": null,
      "replies": [
        {
          "id": 928296,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/13/2020 21:11:34",
          "content": "<p>can you show beginning of your submission.csv? </p>\n\n<p>I think for ROC only order matters (not absolute value), that's why <code>roc_auc_score</code> works even for negative numbers but I am not sure what's the implementation in the competition</p>\n\n<p>I am still in the research phase when I check how different things work and now I want to compare some simple predictions so at this point I need to prepare submission code</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928338,
          "author_name": "aroraaman",
          "author_url": "",
          "post_date": "07/13/2020 22:25:46",
          "content": "<p>The values are just between 0-1 like 0.021, 0.043, 0.53... and so on. </p>\n\n<p>Also, would be great if you add some credits for the loss function code you have borrowed. :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928341,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/13/2020 22:32:01",
          "content": "<p>:)</p>\n\n<p>```\nimport torch.nn as nn\nimport torch\nimport torch.nn.functional as F</p>\n\n<p>\"\"\"\nby Aman Arora\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/162035\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/162035</a>\n\"\"\"\nclass WeightedFocalLoss(nn.Module):\n    def <strong>init</strong>(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).<strong>init</strong>()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma</p>\n\n<pre><code>def forward(self, inputs, targets):\n    BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n    targets = targets.type(torch.long)\n    at = self.alpha.gather(0, targets.data.view(-1))\n    pt = torch.exp(-BCE_loss)\n    F_loss = at*(1-pt)**self.gamma * BCE_loss\n    return F_loss.mean()\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928345,
          "author_name": "aroraaman",
          "author_url": "",
          "post_date": "07/13/2020 22:36:03",
          "content": "<p>Thank you, now I feel guilty for asking - haha!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928348,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/13/2020 22:45:51",
          "content": "<p>if I win you will be in the published code, of course assuming focal loss will survive :)\nI switched from BCE to focal loss and now I research other areas, will back to losses later</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928352,
          "author_name": "aroraaman",
          "author_url": "",
          "post_date": "07/13/2020 22:57:39",
          "content": "<p>Thank you - appreciate it :) </p>\n\n<p>Just to mention, some have reported BCE with label smoothing 0.5 works better than FL. I wasn't able to find the exact post where this was discussed so might be worth trying that too :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928373,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/13/2020 23:52:45",
          "content": "<p>Yes I have seen these discussions but I will focus on it later, there is a lot to explore in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928838,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "07/14/2020 08:55:54",
          "content": "<p><a href=\"/aroraaman\">@aroraaman</a> do you happen to have a working implementation of label smoothing  loss with you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959799,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/05/2020 22:43:26",
          "content": "<p><a href=\"/aroraaman\">@aroraaman</a> are you using label smoothing? Also I think you mean't .05 not .5 but correct me if I am wrong.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 928366,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "07/13/2020 23:37:22",
      "content": "<p>I think all people use sigmoid to scale output 0 to 1. That works for me to get 0.951</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 928431,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/14/2020 01:39:21",
      "content": "<p>It doesn't matter for AUC. As long as the order (ranking) of all the preds stays the same, the AUC will stay the same. (And i'm pretty sure you can submit the negative values to Kaggle). Just do the same thing to all your models before you ensemble them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 928529,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 04:16:10",
          "content": "<p>confirmed, submission with or without sigmoid gives same score, so even negative numbers are ok</p>\n\n<p>however there must be some bug in my code, I am doing some simple model test and single fold score was 0.906 and LB score is 0.858</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928594,
          "author_name": "aroraaman",
          "author_url": "",
          "post_date": "07/14/2020 05:03:34",
          "content": "<p>There a lot of duplicates in the dataset and also best to create folds split by patient_id too. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928615,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 05:22:56",
          "content": "<p>this is split I use:</p>\n\n<p>```\n    patients = data[\"patient_id\"].unique()\n    from sklearn.model_selection import KFold</p>\n\n<pre><code>kf = KFold(n_splits=n_splits, random_state=1234, shuffle=True)\n\nfor i, (train_index, valid_index) in enumerate(kf.split(patients)):\n    if (i == fold):\n        patients_train, patients_valid = patients[train_index], patients[valid_index]\n\ndata_train = data.loc[data['patient_id'].isin(patients_train)]    \ndata_valid = data.loc[data['patient_id'].isin(patients_valid)] \n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929173,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/14/2020 14:00:28",
          "content": "<p>What's the purpose of <code>if (i == fold):</code> ? What is the variable <code>fold</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929194,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 14:14:53",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> When I train model I do it for specific fold, so after 5 runs I have 5 models</p>\n\n<p>I submitted results for single model, maybe if I blend 5 models the CV-LB gap will be smaller, will need to check it because I wasn't able to find any leak yet</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929207,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/14/2020 14:26:16",
          "content": "<p>ah ok, that's smart. I also run experiments on the same fold. Before i go to bed, i put my code on an infinite loop training a list of different models on the same fold and logging the validation AUC.</p>\n\n<p>I don't see a leak in your validation setup. This comp is weird. I'm struggling to match my local validation score with LB too. (My LB is higher than my local val). Today i plan to do more train test EDA. Perhaps i'll run some adversarial and try to figure out what's going on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929236,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 14:48:57",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Do you train on GPU? Because if you are training on your computer I assume you can't use TPU? I am doing experiments on very low resolution right now. I did research on TPU as you have seen on the forum, but I am not using it to train my models now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929247,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/14/2020 14:55:34",
          "content": "<p>Yes i train on GPU. If you use multiple GPUs with mixed precision, you can run experiments as fast or faster than TPU.</p>\n\n<p>I use my <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">public notebook</a> and add </p>\n\n<pre><code>DEVICE = \"GPU\"\nstrategy = tf.distribute.MirroredStrategy()\ntf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})\n</code></pre>\n\n<p>And I run one fold over and over.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929267,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 15:14:24",
          "content": "<p>not sure what mixed precision is, is it like half float in pytorch?\nI have only one GPU, not sure should I buy more because it would be used only for Kaggle</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929286,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/14/2020 15:31:52",
          "content": "<p>Yes half float. Mixed uses some <code>float32</code> and some <code>float16</code>. It accelerates computation and allows more data to fit in GPU RAM. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929299,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 15:39:12",
          "content": "<p>I tried that in previous competition but failed, will need to check it again here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 928565,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "07/14/2020 04:49:52",
      "content": "<p>It shouldn't make a difference. I'm submitting raw logits. If you plot a histogram of your out-of-fold logits and test prediction logits, you can easily see how the two distributions differ (at least in prediction)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 934432,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "07/18/2020 12:35:04",
      "content": "<p>sigmoid help  you to scale output to 0-1. Your target should be sigmoid(Linear()). Your target now not use sigmoid so it can be another.</p>",
      "votes": null,
      "replies": [
        {
          "id": 947435,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "07/27/2020 09:10:47",
          "content": "<p>Do you use threshold for this competition e.g. a threshold=0.5 you smooth a value of 0.7 to 1 and submit it like that in your cv?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "928269": "Last layer in my model is nn.Linear.\nSo there is no sigmoid or anything else after Linear, the values can be negative.\n\nThen I use following loss function:\n```\nclass WeightedFocalLoss(nn.Module):\n    def __init__(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).__init__()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma\n\n    def forward(self, inputs, targets):\n        BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        targets = targets.type(torch.long)\n        at = self.alpha.gather(0, targets.data.view(-1))\n        pt = torch.exp(-BCE_loss)\n        F_loss = at*(1-pt)**self.gamma * BCE_loss\n        return F_loss.mean()\n```\n\nand then I calculate score with `roc_auc_score`\n\nQuestion: how should I prepare submission?\nMy file looks like that:\n```\nimage_name,target\nISIC_0052060,0.012927517\nISIC_0052349,0.018632364\nISIC_0058510,0.029537585\nISIC_0073313,-0.00839683\nISIC_0073502,-0.02732876\n```\n\nAccording to competition description \"For each image_name in the test set, you must predict the probability (target) that the sample is malignant.\"\n\nSo probability should not be negative.\n\nDo you:\n- scale the prediction to fit inside 0.0 and 1.0?\n- process with sigmoid or something else?\n- use different network architecture so last layer is not just Linear?",
    "928291": "Processing with Sigmoid seems to work pretty well for me. \nThis is a great question! I’m also interested to hear about other options.",
    "928296": "can you show beginning of your submission.csv? \n\nI think for ROC only order matters (not absolute value), that's why `roc_auc_score` works even for negative numbers but I am not sure what's the implementation in the competition\n\nI am still in the research phase when I check how different things work and now I want to compare some simple predictions so at this point I need to prepare submission code",
    "928338": "The values are just between 0-1 like 0.021, 0.043, 0.53... and so on. \n\nAlso, would be great if you add some credits for the loss function code you have borrowed. :)",
    "928341": ":)\n\n```\nimport torch.nn as nn\nimport torch\nimport torch.nn.functional as F\n\n\"\"\"\nby Aman Arora\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/162035\n\"\"\"\nclass WeightedFocalLoss(nn.Module):\n    def __init__(self, device, alpha=.25, gamma=2):\n        super(WeightedFocalLoss, self).__init__()\n        self.alpha = torch.tensor([alpha, 1-alpha])\n        if (device):\n            self.alpha = self.alpha.to(device)\n        self.gamma = gamma\n\n    def forward(self, inputs, targets):\n        BCE_loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none')\n        targets = targets.type(torch.long)\n        at = self.alpha.gather(0, targets.data.view(-1))\n        pt = torch.exp(-BCE_loss)\n        F_loss = at*(1-pt)**self.gamma * BCE_loss\n        return F_loss.mean()\n```",
    "928345": "Thank you, now I feel guilty for asking - haha!",
    "928348": "if I win you will be in the published code, of course assuming focal loss will survive :)\nI switched from BCE to focal loss and now I research other areas, will back to losses later",
    "928352": "Thank you - appreciate it :) \n\nJust to mention, some have reported BCE with label smoothing 0.5 works better than FL. I wasn't able to find the exact post where this was discussed so might be worth trying that too :)",
    "928366": "I think all people use sigmoid to scale output 0 to 1. That works for me to get 0.951",
    "928373": "Yes I have seen these discussions but I will focus on it later, there is a lot to explore in this competition.",
    "928431": "It doesn't matter for AUC. As long as the order (ranking) of all the preds stays the same, the AUC will stay the same. (And i'm pretty sure you can submit the negative values to Kaggle). Just do the same thing to all your models before you ensemble them.",
    "928529": "confirmed, submission with or without sigmoid gives same score, so even negative numbers are ok\n\nhowever there must be some bug in my code, I am doing some simple model test and single fold score was 0.906 and LB score is 0.858",
    "928565": "It shouldn't make a difference. I'm submitting raw logits. If you plot a histogram of your out-of-fold logits and test prediction logits, you can easily see how the two distributions differ (at least in prediction)",
    "928594": "There a lot of duplicates in the dataset and also best to create folds split by patient_id too.",
    "928615": "this is split I use:\n\n```\n    patients = data[\"patient_id\"].unique()\n    from sklearn.model_selection import KFold\n    \n    kf = KFold(n_splits=n_splits, random_state=1234, shuffle=True)\n    \n    for i, (train_index, valid_index) in enumerate(kf.split(patients)):\n        if (i == fold):\n            patients_train, patients_valid = patients[train_index], patients[valid_index]\n    \n    data_train = data.loc[data['patient_id'].isin(patients_train)]    \n    data_valid = data.loc[data['patient_id'].isin(patients_valid)] \n\n```",
    "928838": "aroraaman do you happen to have a working implementation of label smoothing  loss with you?",
    "929173": "What's the purpose of `if (i == fold):` ? What is the variable `fold`?",
    "929194": "cdeotte When I train model I do it for specific fold, so after 5 runs I have 5 models\n\nI submitted results for single model, maybe if I blend 5 models the CV-LB gap will be smaller, will need to check it because I wasn't able to find any leak yet",
    "929207": "ah ok, that's smart. I also run experiments on the same fold. Before i go to bed, i put my code on an infinite loop training a list of different models on the same fold and logging the validation AUC.\n\nI don't see a leak in your validation setup. This comp is weird. I'm struggling to match my local validation score with LB too. (My LB is higher than my local val). Today i plan to do more train test EDA. Perhaps i'll run some adversarial and try to figure out what's going on.",
    "929236": "cdeotte Do you train on GPU? Because if you are training on your computer I assume you can't use TPU? I am doing experiments on very low resolution right now. I did research on TPU as you have seen on the forum, but I am not using it to train my models now.",
    "929247": "Yes i train on GPU. If you use multiple GPUs with mixed precision, you can run experiments as fast or faster than TPU.\n\nI use my [public notebook][1] and add \n\n    DEVICE = \"GPU\"\n    strategy = tf.distribute.MirroredStrategy()\n    tf.config.optimizer.set_experimental_options({\"auto_mixed_precision\": True})\n\nAnd I run one fold over and over.\n\n[1]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords",
    "929267": "not sure what mixed precision is, is it like half float in pytorch?\nI have only one GPU, not sure should I buy more because it would be used only for Kaggle",
    "929286": "Yes half float. Mixed uses some `float32` and some `float16`. It accelerates computation and allows more data to fit in GPU RAM.",
    "929299": "I tried that in previous competition but failed, will need to check it again here.",
    "934432": "sigmoid help  you to scale output to 0-1. Your target should be sigmoid(Linear()). Your target now not use sigmoid so it can be another.",
    "947435": "Do you use threshold for this competition e.g. a threshold=0.5 you smooth a value of 0.7 to 1 and submit it like that in your cv?",
    "959799": "aroraaman are you using label smoothing? Also I think you mean't .05 not .5 but correct me if I am wrong."
  },
  "source": "meta"
}