{
  "id": 164357,
  "title": "Values >1 after sigmoid ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164357",
  "author_name": "",
  "post_date": "2020-07-06T00:37:38.611130400Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/samuelvedrik/pytorch-melanoma-efficientnet?rvi=1\">https://www.kaggle.com/samuelvedrik/pytorch-melanoma-efficientnet?rvi=1</a></p>\n\n<p>I noticed in my submissions that some values are greater than 1. (The first one I saw was in fact 2.9). I thought that the sigmoid function converges to 1, so how is this possible? Did I make a coding error? My kernel is as above. Thank you!</p>",
  "messages": [
    {
      "id": "916715",
      "postDate": "07/06/2020 00:37:38",
      "content": "<p><a href=\"https://www.kaggle.com/samuelvedrik/pytorch-melanoma-efficientnet?rvi=1\">https://www.kaggle.com/samuelvedrik/pytorch-melanoma-efficientnet?rvi=1</a></p>\n\n<p>I noticed in my submissions that some values are greater than 1. (The first one I saw was in fact 2.9). I thought that the sigmoid function converges to 1, so how is this possible? Did I make a coding error? My kernel is as above. Thank you!</p>",
      "rawMarkdown": "https://www.kaggle.com/samuelvedrik/pytorch-melanoma-efficientnet?rvi=1\n\n\nI noticed in my submissions that some values are greater than 1. (The first one I saw was in fact 2.9). I thought that the sigmoid function converges to 1, so how is this possible? Did I make a coding error? My kernel is as above. Thank you!",
      "votes": null
    },
    {
      "id": "916857",
      "postDate": "07/06/2020 04:48:07",
      "content": "<p>are you sure you aren't looking at something with a negative exponent? <br>\nExample:</p>\n\n<p>ISIC_0371470,2.4975307e-05</p>\n\n<p>which is actually a very very small number..........in decimal:\n0.000024975307</p>\n\n<p>Also, you may want to read about the perils of combining batchnorm with dropout......and you should not impute missing data on the full dataset then split as you are leaking data if you do it that way.  Same thing if you normalize on the full set........you need to work on the splits in isolation.</p>\n\n<p>Why do you have freeze and unfreeze?  Were you going to try two methods of transfer learning, i.e. training the full model and also trying to just freeze all but the final classifier and train just the classifier?   If you are trying both how has that gone for you, did you find one worked better?</p>",
      "rawMarkdown": "are you sure you aren't looking at something with a negative exponent?  \nExample:\n\nISIC_0371470,2.4975307e-05\n\nwhich is actually a very very small number..........in decimal:\n0.000024975307\n\nAlso, you may want to read about the perils of combining batchnorm with dropout......and you should not impute missing data on the full dataset then split as you are leaking data if you do it that way.  Same thing if you normalize on the full set........you need to work on the splits in isolation.\n\nWhy do you have freeze and unfreeze?  Were you going to try two methods of transfer learning, i.e. training the full model and also trying to just freeze all but the final classifier and train just the classifier?   If you are trying both how has that gone for you, did you find one worked better?",
      "votes": null
    },
    {
      "id": "916862",
      "postDate": "07/06/2020 04:51:32",
      "content": "<p>Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.  The scoring is based on ROC AUC which is invariant to 0 or 1.  As long as you are separating the classes well with probablities i it should not matter.........it could hurt your score even but if you think it helps definitely report back.</p>",
      "rawMarkdown": "Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.  The scoring is based on ROC AUC which is invariant to 0 or 1.  As long as you are separating the classes well with probablities i it should not matter.........it could hurt your score even but if you think it helps definitely report back.",
      "votes": null
    },
    {
      "id": "916882",
      "postDate": "07/06/2020 05:03:33",
      "content": "<p>Ah I see, that might be it! My bad for not reading it properly.</p>\n\n<p>Yes, I first froze to just train the final classifier and then unfroze in hopes to make it more specialized. I lowered the learning rate during this stage to not alter the weights too much. It worked pretty well, I got an increase of 0.03 ROC AUC compared to just training the final classifier. It might just be my own confirmation bias tho. </p>\n\n<p>Also, thank you for the tip about the imputation! I certainly forgot that it might cause a data leak. I'll definitely try to fix that, thank you!</p>",
      "rawMarkdown": "Ah I see, that might be it! My bad for not reading it properly.\n\nYes, I first froze to just train the final classifier and then unfroze in hopes to make it more specialized. I lowered the learning rate during this stage to not alter the weights too much. It worked pretty well, I got an increase of 0.03 ROC AUC compared to just training the final classifier. It might just be my own confirmation bias tho. \n\nAlso, thank you for the tip about the imputation! I certainly forgot that it might cause a data leak. I'll definitely try to fix that, thank you!",
      "votes": null
    },
    {
      "id": "916938",
      "postDate": "07/06/2020 06:16:31",
      "content": "<blockquote>\n  <p>Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.</p>\n</blockquote>\n\n<p>It is true that you will get the same LB if you submit logits or probabilities. However when you use back propagation to train your CNN, you want that back propagation to go backwards through a sigmoid so that you receive the benefit of the sigmoid's derivative to more forcefully separate classes.</p>",
      "rawMarkdown": "&gt; Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.\n\nIt is true that you will get the same LB if you submit logits or probabilities. However when you use back propagation to train your CNN, you want that back propagation to go backwards through a sigmoid so that you receive the benefit of the sigmoid's derivative to more forcefully separate classes.",
      "votes": null
    },
    {
      "id": "918338",
      "postDate": "07/07/2020 07:34:25",
      "content": "<p>In my case, I am using BCEWithLogitsLoss, and it already does a Sigmoid, so I wasn't doing a Sigmoid on my NN at the end.  I have a CNN for the images, and a FC for the Meta data, and then I combine them in a Linear.  Do you think one should use a Sigmoid but then what about BCE?</p>\n\n<p>Now, I am experimenting with other loss criterion, and for those I am putting a Sigmoid at the end of mine Linear network that does the combining.  What are your thoughts?</p>",
      "rawMarkdown": "In my case, I am using BCEWithLogitsLoss, and it already does a Sigmoid, so I wasn't doing a Sigmoid on my NN at the end.  I have a CNN for the images, and a FC for the Meta data, and then I combine them in a Linear.  Do you think one should use a Sigmoid but then what about BCE?\n\nNow, I am experimenting with other loss criterion, and for those I am putting a Sigmoid at the end of mine Linear network that does the combining.  What are your thoughts?",
      "votes": null
    },
    {
      "id": "1185262",
      "postDate": "02/04/2021 04:43:45",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> super old post, but hopefully you see this.  You wrote above that the sigmoid helps in backpropagation.  So I take it you mean if you have the sigmoid applied in the forward function of your model, so say being applied to the final output.  I didn’t think that activation functions like a sigmoid on the final output were involved in influencing back propagation.  Is it so?</p>",
      "rawMarkdown": "cdeotte super old post, but hopefully you see this.  You wrote above that the sigmoid helps in backpropagation.  So I take it you mean if you have the sigmoid applied in the forward function of your model, so say being applied to the final output.  I didn’t think that activation functions like a sigmoid on the final output were involved in influencing back propagation.  Is it so?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 916857,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "07/06/2020 04:48:07",
      "content": "<p>are you sure you aren't looking at something with a negative exponent? <br>\nExample:</p>\n\n<p>ISIC_0371470,2.4975307e-05</p>\n\n<p>which is actually a very very small number..........in decimal:\n0.000024975307</p>\n\n<p>Also, you may want to read about the perils of combining batchnorm with dropout......and you should not impute missing data on the full dataset then split as you are leaking data if you do it that way.  Same thing if you normalize on the full set........you need to work on the splits in isolation.</p>\n\n<p>Why do you have freeze and unfreeze?  Were you going to try two methods of transfer learning, i.e. training the full model and also trying to just freeze all but the final classifier and train just the classifier?   If you are trying both how has that gone for you, did you find one worked better?</p>",
      "votes": null,
      "replies": [
        {
          "id": 916862,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/06/2020 04:51:32",
          "content": "<p>Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.  The scoring is based on ROC AUC which is invariant to 0 or 1.  As long as you are separating the classes well with probablities i it should not matter.........it could hurt your score even but if you think it helps definitely report back.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916882,
          "author_name": "samuelvedrik",
          "author_url": "",
          "post_date": "07/06/2020 05:03:33",
          "content": "<p>Ah I see, that might be it! My bad for not reading it properly.</p>\n\n<p>Yes, I first froze to just train the final classifier and then unfroze in hopes to make it more specialized. I lowered the learning rate during this stage to not alter the weights too much. It worked pretty well, I got an increase of 0.03 ROC AUC compared to just training the final classifier. It might just be my own confirmation bias tho. </p>\n\n<p>Also, thank you for the tip about the imputation! I certainly forgot that it might cause a data leak. I'll definitely try to fix that, thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916938,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/06/2020 06:16:31",
          "content": "<blockquote>\n  <p>Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.</p>\n</blockquote>\n\n<p>It is true that you will get the same LB if you submit logits or probabilities. However when you use back propagation to train your CNN, you want that back propagation to go backwards through a sigmoid so that you receive the benefit of the sigmoid's derivative to more forcefully separate classes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918338,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/07/2020 07:34:25",
          "content": "<p>In my case, I am using BCEWithLogitsLoss, and it already does a Sigmoid, so I wasn't doing a Sigmoid on my NN at the end.  I have a CNN for the images, and a FC for the Meta data, and then I combine them in a Linear.  Do you think one should use a Sigmoid but then what about BCE?</p>\n\n<p>Now, I am experimenting with other loss criterion, and for those I am putting a Sigmoid at the end of mine Linear network that does the combining.  What are your thoughts?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1185262,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "02/04/2021 04:43:45",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> super old post, but hopefully you see this.  You wrote above that the sigmoid helps in backpropagation.  So I take it you mean if you have the sigmoid applied in the forward function of your model, so say being applied to the final output.  I didn’t think that activation functions like a sigmoid on the final output were involved in influencing back propagation.  Is it so?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "916715": "https://www.kaggle.com/samuelvedrik/pytorch-melanoma-efficientnet?rvi=1\n\n\nI noticed in my submissions that some values are greater than 1. (The first one I saw was in fact 2.9). I thought that the sigmoid function converges to 1, so how is this possible? Did I make a coding error? My kernel is as above. Thank you!",
    "916857": "are you sure you aren't looking at something with a negative exponent?  \nExample:\n\nISIC_0371470,2.4975307e-05\n\nwhich is actually a very very small number..........in decimal:\n0.000024975307\n\nAlso, you may want to read about the perils of combining batchnorm with dropout......and you should not impute missing data on the full dataset then split as you are leaking data if you do it that way.  Same thing if you normalize on the full set........you need to work on the splits in isolation.\n\nWhy do you have freeze and unfreeze?  Were you going to try two methods of transfer learning, i.e. training the full model and also trying to just freeze all but the final classifier and train just the classifier?   If you are trying both how has that gone for you, did you find one worked better?",
    "916862": "Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.  The scoring is based on ROC AUC which is invariant to 0 or 1.  As long as you are separating the classes well with probablities i it should not matter.........it could hurt your score even but if you think it helps definitely report back.",
    "916882": "Ah I see, that might be it! My bad for not reading it properly.\n\nYes, I first froze to just train the final classifier and then unfroze in hopes to make it more specialized. I lowered the learning rate during this stage to not alter the weights too much. It worked pretty well, I got an increase of 0.03 ROC AUC compared to just training the final classifier. It might just be my own confirmation bias tho. \n\nAlso, thank you for the tip about the imputation! I certainly forgot that it might cause a data leak. I'll definitely try to fix that, thank you!",
    "916938": "&gt; Also, someone correct me if I am wrong, but I don't think there is a need for sigmoid activation on your test predictions.\n\nIt is true that you will get the same LB if you submit logits or probabilities. However when you use back propagation to train your CNN, you want that back propagation to go backwards through a sigmoid so that you receive the benefit of the sigmoid's derivative to more forcefully separate classes.",
    "918338": "In my case, I am using BCEWithLogitsLoss, and it already does a Sigmoid, so I wasn't doing a Sigmoid on my NN at the end.  I have a CNN for the images, and a FC for the Meta data, and then I combine them in a Linear.  Do you think one should use a Sigmoid but then what about BCE?\n\nNow, I am experimenting with other loss criterion, and for those I am putting a Sigmoid at the end of mine Linear network that does the combining.  What are your thoughts?",
    "1185262": "cdeotte super old post, but hopefully you see this.  You wrote above that the sigmoid helps in backpropagation.  So I take it you mean if you have the sigmoid applied in the forward function of your model, so say being applied to the final output.  I didn’t think that activation functions like a sigmoid on the final output were involved in influencing back propagation.  Is it so?"
  },
  "source": "meta"
}