{
  "id": 71484,
  "title": "What can I do for the  predictions whose probability towards all classes equals to zero?Dose it means that I have choosed a bad threshold?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/71484",
  "author_name": "",
  "post_date": "2018-11-14T03:05:24.420743Z",
  "votes": null,
  "comment_count": 10,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "420722",
      "postDate": "11/14/2018 03:05:24",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "420723",
      "postDate": "11/14/2018 03:07:19",
      "content": "<p>I'm a fresh hand to kaggle,this is my first time to begin a competition in kaggle. Thanks for your help.</p>",
      "rawMarkdown": "I'm a fresh hand to kaggle,this is my first time to begin a competition in kaggle. Thanks for your help.",
      "votes": null
    },
    {
      "id": "420725",
      "postDate": "11/14/2018 03:11:42",
      "content": "<p>Its a bit hard to answer without knowing what you actually did. Can be problem with the training, problem with the prediction, problem with the pre/post processing, problem with assembling the submission... </p>",
      "rawMarkdown": "Its a bit hard to answer without knowing what you actually did. Can be problem with the training, problem with the prediction, problem with the pre/post processing, problem with assembling the submission...",
      "votes": null
    },
    {
      "id": "420814",
      "postDate": "11/14/2018 06:50:51",
      "content": "<p>I build the resnet50 to begin this competition, and change the output of the net to be a tensor with dimension [batch,num_classes]，each output responding to a specified class represent the probability it belongs to that class. loss function uses focal loss. But when I begin a prediction(test), I found that for many samples, their probability towards all classes equals to 0! This means I cannot get a prediction for these samples. Do I choose a wrong loss function or the strategy I took is wrong? thanks for you help.</p>",
      "rawMarkdown": "I build the resnet50 to begin this competition, and change the output of the net to be a tensor with dimension [batch,num_classes]，each output responding to a specified class represent the probability it belongs to that class. loss function uses focal loss. But when I begin a prediction(test), I found that for many samples, their probability towards all classes equals to 0! This means I cannot get a prediction for these samples. Do I choose a wrong loss function or the strategy I took is wrong? thanks for you help.",
      "votes": null
    },
    {
      "id": "420816",
      "postDate": "11/14/2018 06:52:21",
      "content": "<p>I set the threshold to 0.5.</p>",
      "rawMarkdown": "I set the threshold to 0.5.",
      "votes": null
    },
    {
      "id": "420857",
      "postDate": "11/14/2018 08:34:07",
      "content": "<p>It can still be many things. One possibility would be not training enough or training too much. Did you use the rgb out of the rgby? Did you use pretrained model (imagenet)? What was the loss and val loss during the training? I suggest you revisit all the steps and check. </p>",
      "rawMarkdown": "It can still be many things. One possibility would be not training enough or training too much. Did you use the rgb out of the rgby? Did you use pretrained model (imagenet)? What was the loss and val loss during the training? I suggest you revisit all the steps and check.",
      "votes": null
    },
    {
      "id": "420934",
      "postDate": "11/14/2018 10:56:27",
      "content": "<p>Thanks for your guide, I will make changes according to your suggestiions. </p>",
      "rawMarkdown": "Thanks for your guide, I will make changes according to your suggestiions.",
      "votes": null
    },
    {
      "id": "421114",
      "postDate": "11/14/2018 15:49:04",
      "content": "<p>If the target metric is assumed to be real \"macro\" f1 score it can be favorable for your model to not give a prediction for every sample.</p>\n\n<p>However, the actual evaluation that they use might not be macro f1 score since a lot of people here have noted that the LB score does not match within any margin of error of what it should be. So if you wanna experiment with finding out if it helps this unknown \"real\" metric to predict every sample, one thing you could try is to just assign the class with the highest probability in cases of all zeros.</p>",
      "rawMarkdown": "If the target metric is assumed to be real \"macro\" f1 score it can be favorable for your model to not give a prediction for every sample.\n\nHowever, the actual evaluation that they use might not be macro f1 score since a lot of people here have noted that the LB score does not match within any margin of error of what it should be. So if you wanna experiment with finding out if it helps this unknown \"real\" metric to predict every sample, one thing you could try is to just assign the class with the highest probability in cases of all zeros.",
      "votes": null
    },
    {
      "id": "421187",
      "postDate": "11/14/2018 17:31:15",
      "content": "<p>If you're using regular loss functions that don't consider relations between classes or that don't consider any kind of unbalance,  such as binary crossentropy, or maybe mean squared error, the models will certainly (at first) go to all zeros, then a fine tuning might raise some ones. This happens because there are way more zeros in the targets than ones.</p>\n\n<p>The zero rate is very probably above 90%, because you've got usually 3 or less classes/ones among 28 classes.</p>\n\n<p>This would be less drastic with loss functions that consider that only one class is correct, such as categorical crossentropy (but this is not ideal because we want more than one correct class).   </p>\n\n<p>Ideally, you should taylor your own loss, perhaps creating a better balance, or using some sort of F1 loss. </p>\n\n<p>I'm using F1 loss and trying some other options (not saying this is the ideal loss, but seems pretty reasonable - needs big batches).   </p>\n\n<pre><code>#for Keras, but easily adaptable to other libraries   \ndef competitionLoss(true,pred):\n    groundPositives = K.sum(true, axis=0) + K.epsilon()\n    correctPositives = K.sum(true * pred, axis=0) + K.epsilon()\n    predictedPositives = K.sum(pred, axis=0) + K.epsilon()\n\n    precision = correctPositives / predictedPositives\n    recall = correctPositives / groundPositives\n\n    m = (2 * precision * recall) / (precision + recall)\n\n    return 1-K.mean(m)\n</code></pre>",
      "rawMarkdown": "If you're using regular loss functions that don't consider relations between classes or that don't consider any kind of unbalance,  such as binary crossentropy, or maybe mean squared error, the models will certainly (at first) go to all zeros, then a fine tuning might raise some ones. This happens because there are way more zeros in the targets than ones.\n\nThe zero rate is very probably above 90%, because you've got usually 3 or less classes/ones among 28 classes.\n\nThis would be less drastic with loss functions that consider that only one class is correct, such as categorical crossentropy (but this is not ideal because we want more than one correct class).   \n\nIdeally, you should taylor your own loss, perhaps creating a better balance, or using some sort of F1 loss. \n\nI'm using F1 loss and trying some other options (not saying this is the ideal loss, but seems pretty reasonable - needs big batches).   \n\n    #for Keras, but easily adaptable to other libraries   \n    def competitionLoss(true,pred):\n        groundPositives = K.sum(true, axis=0) + K.epsilon()\n        correctPositives = K.sum(true * pred, axis=0) + K.epsilon()\n        predictedPositives = K.sum(pred, axis=0) + K.epsilon()\n\n        precision = correctPositives / predictedPositives\n        recall = correctPositives / groundPositives\n\n        m = (2 * precision * recall) / (precision + recall)\n\n        return 1-K.mean(m)",
      "votes": null
    },
    {
      "id": "423790",
      "postDate": "11/19/2018 03:30:44",
      "content": "<p>Thanks for you code and great help. In the training, I use the popular Focal loss as the loss function. I have failed to try to use the inititial F1 score, but it has no gradient to propogation.  I guess you might use a different function with some changes and I have no experience on this, could you give me some tips to solve this problem? Thank you!</p>",
      "rawMarkdown": "Thanks for you code and great help. In the training, I use the popular Focal loss as the loss function. I have failed to try to use the inititial F1 score, but it has no gradient to propogation.  I guess you might use a different function with some changes and I have no experience on this, could you give me some tips to solve this problem? Thank you!",
      "votes": null
    },
    {
      "id": "423792",
      "postDate": "11/19/2018 03:35:54",
      "content": "<p>I used the Focal loss as our team's loss function, at the begining, we struggle from get a \"proper\" prediction, since many output are all zeros. But this is mostly caused by ineffcient training. After we made some changes, things became better. Another problem is that many samples to be predicted output the same, does this means that we still have a ineffcient training or there are some other problems? </p>",
      "rawMarkdown": "I used the Focal loss as our team's loss function, at the begining, we struggle from get a \"proper\" prediction, since many output are all zeros. But this is mostly caused by ineffcient training. After we made some changes, things became better. Another problem is that many samples to be predicted output the same, does this means that we still have a ineffcient training or there are some other problems?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 420723,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2018 03:07:19",
      "content": "<p>I'm a fresh hand to kaggle,this is my first time to begin a competition in kaggle. Thanks for your help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420725,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "11/14/2018 03:11:42",
      "content": "<p>Its a bit hard to answer without knowing what you actually did. Can be problem with the training, problem with the prediction, problem with the pre/post processing, problem with assembling the submission... </p>",
      "votes": null,
      "replies": [
        {
          "id": 420814,
          "author_name": "",
          "author_url": "",
          "post_date": "11/14/2018 06:50:51",
          "content": "<p>I build the resnet50 to begin this competition, and change the output of the net to be a tensor with dimension [batch,num_classes]，each output responding to a specified class represent the probability it belongs to that class. loss function uses focal loss. But when I begin a prediction(test), I found that for many samples, their probability towards all classes equals to 0! This means I cannot get a prediction for these samples. Do I choose a wrong loss function or the strategy I took is wrong? thanks for you help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420816,
          "author_name": "",
          "author_url": "",
          "post_date": "11/14/2018 06:52:21",
          "content": "<p>I set the threshold to 0.5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420857,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "11/14/2018 08:34:07",
          "content": "<p>It can still be many things. One possibility would be not training enough or training too much. Did you use the rgb out of the rgby? Did you use pretrained model (imagenet)? What was the loss and val loss during the training? I suggest you revisit all the steps and check. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420934,
          "author_name": "",
          "author_url": "",
          "post_date": "11/14/2018 10:56:27",
          "content": "<p>Thanks for your guide, I will make changes according to your suggestiions. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 421114,
      "author_name": "chrissikek",
      "author_url": "",
      "post_date": "11/14/2018 15:49:04",
      "content": "<p>If the target metric is assumed to be real \"macro\" f1 score it can be favorable for your model to not give a prediction for every sample.</p>\n\n<p>However, the actual evaluation that they use might not be macro f1 score since a lot of people here have noted that the LB score does not match within any margin of error of what it should be. So if you wanna experiment with finding out if it helps this unknown \"real\" metric to predict every sample, one thing you could try is to just assign the class with the highest probability in cases of all zeros.</p>",
      "votes": null,
      "replies": [
        {
          "id": 423792,
          "author_name": "",
          "author_url": "",
          "post_date": "11/19/2018 03:35:54",
          "content": "<p>I used the Focal loss as our team's loss function, at the begining, we struggle from get a \"proper\" prediction, since many output are all zeros. But this is mostly caused by ineffcient training. After we made some changes, things became better. Another problem is that many samples to be predicted output the same, does this means that we still have a ineffcient training or there are some other problems? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 421187,
      "author_name": "danmoller",
      "author_url": "",
      "post_date": "11/14/2018 17:31:15",
      "content": "<p>If you're using regular loss functions that don't consider relations between classes or that don't consider any kind of unbalance,  such as binary crossentropy, or maybe mean squared error, the models will certainly (at first) go to all zeros, then a fine tuning might raise some ones. This happens because there are way more zeros in the targets than ones.</p>\n\n<p>The zero rate is very probably above 90%, because you've got usually 3 or less classes/ones among 28 classes.</p>\n\n<p>This would be less drastic with loss functions that consider that only one class is correct, such as categorical crossentropy (but this is not ideal because we want more than one correct class).   </p>\n\n<p>Ideally, you should taylor your own loss, perhaps creating a better balance, or using some sort of F1 loss. </p>\n\n<p>I'm using F1 loss and trying some other options (not saying this is the ideal loss, but seems pretty reasonable - needs big batches).   </p>\n\n<pre><code>#for Keras, but easily adaptable to other libraries   \ndef competitionLoss(true,pred):\n    groundPositives = K.sum(true, axis=0) + K.epsilon()\n    correctPositives = K.sum(true * pred, axis=0) + K.epsilon()\n    predictedPositives = K.sum(pred, axis=0) + K.epsilon()\n\n    precision = correctPositives / predictedPositives\n    recall = correctPositives / groundPositives\n\n    m = (2 * precision * recall) / (precision + recall)\n\n    return 1-K.mean(m)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 423790,
          "author_name": "",
          "author_url": "",
          "post_date": "11/19/2018 03:30:44",
          "content": "<p>Thanks for you code and great help. In the training, I use the popular Focal loss as the loss function. I have failed to try to use the inititial F1 score, but it has no gradient to propogation.  I guess you might use a different function with some changes and I have no experience on this, could you give me some tips to solve this problem? Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "420722": "",
    "420723": "I'm a fresh hand to kaggle,this is my first time to begin a competition in kaggle. Thanks for your help.",
    "420725": "Its a bit hard to answer without knowing what you actually did. Can be problem with the training, problem with the prediction, problem with the pre/post processing, problem with assembling the submission...",
    "420814": "I build the resnet50 to begin this competition, and change the output of the net to be a tensor with dimension [batch,num_classes]，each output responding to a specified class represent the probability it belongs to that class. loss function uses focal loss. But when I begin a prediction(test), I found that for many samples, their probability towards all classes equals to 0! This means I cannot get a prediction for these samples. Do I choose a wrong loss function or the strategy I took is wrong? thanks for you help.",
    "420816": "I set the threshold to 0.5.",
    "420857": "It can still be many things. One possibility would be not training enough or training too much. Did you use the rgb out of the rgby? Did you use pretrained model (imagenet)? What was the loss and val loss during the training? I suggest you revisit all the steps and check.",
    "420934": "Thanks for your guide, I will make changes according to your suggestiions.",
    "421114": "If the target metric is assumed to be real \"macro\" f1 score it can be favorable for your model to not give a prediction for every sample.\n\nHowever, the actual evaluation that they use might not be macro f1 score since a lot of people here have noted that the LB score does not match within any margin of error of what it should be. So if you wanna experiment with finding out if it helps this unknown \"real\" metric to predict every sample, one thing you could try is to just assign the class with the highest probability in cases of all zeros.",
    "421187": "If you're using regular loss functions that don't consider relations between classes or that don't consider any kind of unbalance,  such as binary crossentropy, or maybe mean squared error, the models will certainly (at first) go to all zeros, then a fine tuning might raise some ones. This happens because there are way more zeros in the targets than ones.\n\nThe zero rate is very probably above 90%, because you've got usually 3 or less classes/ones among 28 classes.\n\nThis would be less drastic with loss functions that consider that only one class is correct, such as categorical crossentropy (but this is not ideal because we want more than one correct class).   \n\nIdeally, you should taylor your own loss, perhaps creating a better balance, or using some sort of F1 loss. \n\nI'm using F1 loss and trying some other options (not saying this is the ideal loss, but seems pretty reasonable - needs big batches).   \n\n    #for Keras, but easily adaptable to other libraries   \n    def competitionLoss(true,pred):\n        groundPositives = K.sum(true, axis=0) + K.epsilon()\n        correctPositives = K.sum(true * pred, axis=0) + K.epsilon()\n        predictedPositives = K.sum(pred, axis=0) + K.epsilon()\n\n        precision = correctPositives / predictedPositives\n        recall = correctPositives / groundPositives\n\n        m = (2 * precision * recall) / (precision + recall)\n\n        return 1-K.mean(m)",
    "423790": "Thanks for you code and great help. In the training, I use the popular Focal loss as the loss function. I have failed to try to use the inititial F1 score, but it has no gradient to propogation.  I guess you might use a different function with some changes and I have no experience on this, could you give me some tips to solve this problem? Thank you!",
    "423792": "I used the Focal loss as our team's loss function, at the begining, we struggle from get a \"proper\" prediction, since many output are all zeros. But this is mostly caused by ineffcient training. After we made some changes, things became better. Another problem is that many samples to be predicted output the same, does this means that we still have a ineffcient training or there are some other problems?"
  },
  "source": "meta"
}