{
  "id": 221022,
  "title": "Question About Multi-Label Classification",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/221022",
  "author_name": "Darien Schettler",
  "post_date": "2021-02-20T16:17:00.314000",
  "votes": 10,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi there.</p>\n<p>I think this is a pretty basic question but I wanted to ask here just in case anyone else had the same inquiry (I'm not afraid to ask the simple questions 😅). I did some Googling and have my own speculations but I always appreciate the wisdom of other Kagglers… so here we go.</p>\n<hr>\n<p><br></p>\n<p><strong>SETUP:</strong> </p>\n<p>I'm training a multi-label cell classifier.<br>\nWe have 19 classes (18 organelle structures, and a negative label). </p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong>QUESTION:</strong></p>\n<p>Do I frame this as a 19 class multi-label classification problem or an 18 class classification problem where a lack of confidence in the 18 organelle classes indicates the presence of the negative class?</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong>FOLLOW-UP 1:</strong></p>\n<p>Let's say I decide to frame it as an 18 class multi-label classification problem. Therefore, during training, if I have <strong><em>n</em></strong> negative training samples. </p>\n<p>Knowing this, should I pass the ground-truth label as a vector of length 18 full of 0s for all <strong><em>n</em></strong> negative training samples?</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong>FOLLOW-UP 2:</strong></p>\n<p>During the training of a multi-label classification model with more than a few classes, it's obvious that most of the time the majority of the outputs will be 0.</p>\n<p>i.e. Let's look at 5 outputs for a 5 class model (assuming no negative classes which would only exacerbate the problem).</p>\n<pre><code>model_output_1 = [0.05, 0.23, 0.10, 0.80, 0.10]\nground_truth_1  = [0.00, 0.00, 0.00, 1.00, 0.00,]\n\nmodel_output_2 = [0.95, 0.02, 0.04, 0.02, 0.10]\nground_truth_2  = [1.00, 0.00, 0.00, 0.00, 0.00]\n\nmodel_output_3 = [0.02, 0.66, 0.09, 0.08, 0.13]\nground_truth_3  = [0.00, 1.00, 0.00, 0.00, 0.00,]\n\nmodel_output_4 = [0.15, 0.03, 0.08, 0.38, 0.55]\nground_truth_4  = [0.00, 0.00, 0.00, 0.00, 1.00,]\n\nmodel_output_5 = [0.22, 0.04, 0.98, 0.32, 0.12]\nground_truth_5  = [0.00, 0.00, 1.00, 0.00, 0.00,]\n</code></pre>\n<p>Would it not be true in this case that the model will just eventually learn to predict all 0s all the time? Without a class weighting function that is?</p>\n<p>And if that's true, is it normal practice to pass a class-wise binary, class-weighting?</p>\n<p><br></p>\n<hr>\n<p><strong>Thanks in advance!!</strong></p>",
  "messages": [
    {
      "id": 1211873,
      "postDate": "2021-02-20T16:17:00.313Z",
      "content": "<p>Hi there.</p>\n<p>I think this is a pretty basic question but I wanted to ask here just in case anyone else had the same inquiry (I'm not afraid to ask the simple questions 😅). I did some Googling and have my own speculations but I always appreciate the wisdom of other Kagglers… so here we go.</p>\n<hr>\n<p><br></p>\n<p><strong>SETUP:</strong> </p>\n<p>I'm training a multi-label cell classifier.<br>\nWe have 19 classes (18 organelle structures, and a negative label). </p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong>QUESTION:</strong></p>\n<p>Do I frame this as a 19 class multi-label classification problem or an 18 class classification problem where a lack of confidence in the 18 organelle classes indicates the presence of the negative class?</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong>FOLLOW-UP 1:</strong></p>\n<p>Let's say I decide to frame it as an 18 class multi-label classification problem. Therefore, during training, if I have <strong><em>n</em></strong> negative training samples. </p>\n<p>Knowing this, should I pass the ground-truth label as a vector of length 18 full of 0s for all <strong><em>n</em></strong> negative training samples?</p>\n<p><br></p>\n<hr>\n<p><br></p>\n<p><strong>FOLLOW-UP 2:</strong></p>\n<p>During the training of a multi-label classification model with more than a few classes, it's obvious that most of the time the majority of the outputs will be 0.</p>\n<p>i.e. Let's look at 5 outputs for a 5 class model (assuming no negative classes which would only exacerbate the problem).</p>\n<pre><code>model_output_1 = [0.05, 0.23, 0.10, 0.80, 0.10]\nground_truth_1  = [0.00, 0.00, 0.00, 1.00, 0.00,]\n\nmodel_output_2 = [0.95, 0.02, 0.04, 0.02, 0.10]\nground_truth_2  = [1.00, 0.00, 0.00, 0.00, 0.00]\n\nmodel_output_3 = [0.02, 0.66, 0.09, 0.08, 0.13]\nground_truth_3  = [0.00, 1.00, 0.00, 0.00, 0.00,]\n\nmodel_output_4 = [0.15, 0.03, 0.08, 0.38, 0.55]\nground_truth_4  = [0.00, 0.00, 0.00, 0.00, 1.00,]\n\nmodel_output_5 = [0.22, 0.04, 0.98, 0.32, 0.12]\nground_truth_5  = [0.00, 0.00, 1.00, 0.00, 0.00,]\n</code></pre>\n<p>Would it not be true in this case that the model will just eventually learn to predict all 0s all the time? Without a class weighting function that is?</p>\n<p>And if that's true, is it normal practice to pass a class-wise binary, class-weighting?</p>\n<p><br></p>\n<hr>\n<p><strong>Thanks in advance!!</strong></p>",
      "rawMarkdown": "Hi there.\n\nI think this is a pretty basic question but I wanted to ask here just in case anyone else had the same inquiry (I'm not afraid to ask the simple questions 😅). I did some Googling and have my own speculations but I always appreciate the wisdom of other Kagglers... so here we go.\n\n---\n\n<br>\n\n**SETUP:** \n\nI'm training a multi-label cell classifier.\nWe have 19 classes (18 organelle structures, and a negative label). \n\n<br>\n\n---\n\n<br>\n\n**QUESTION:**\n\nDo I frame this as a 19 class multi-label classification problem or an 18 class classification problem where a lack of confidence in the 18 organelle classes indicates the presence of the negative class?\n\n<br>\n\n---\n\n<br>\n\n**FOLLOW-UP 1:**\n\nLet's say I decide to frame it as an 18 class multi-label classification problem. Therefore, during training, if I have ***n*** negative training samples. \n\nKnowing this, should I pass the ground-truth label as a vector of length 18 full of 0s for all ***n*** negative training samples?\n\n<br>\n\n---\n\n<br>\n\n**FOLLOW-UP 2:**\n\nDuring the training of a multi-label classification model with more than a few classes, it's obvious that most of the time the majority of the outputs will be 0.\n\ni.e. Let's look at 5 outputs for a 5 class model (assuming no negative classes which would only exacerbate the problem).\n\n```python\n\nmodel_output_1 = [0.05, 0.23, 0.10, 0.80, 0.10]\nground_truth_1  = [0.00, 0.00, 0.00, 1.00, 0.00,]\n\nmodel_output_2 = [0.95, 0.02, 0.04, 0.02, 0.10]\nground_truth_2  = [1.00, 0.00, 0.00, 0.00, 0.00]\n\nmodel_output_3 = [0.02, 0.66, 0.09, 0.08, 0.13]\nground_truth_3  = [0.00, 1.00, 0.00, 0.00, 0.00,]\n\nmodel_output_4 = [0.15, 0.03, 0.08, 0.38, 0.55]\nground_truth_4  = [0.00, 0.00, 0.00, 0.00, 1.00,]\n\nmodel_output_5 = [0.22, 0.04, 0.98, 0.32, 0.12]\nground_truth_5  = [0.00, 0.00, 1.00, 0.00, 0.00,]\n\n```\n\nWould it not be true in this case that the model will just eventually learn to predict all 0s all the time? Without a class weighting function that is?\n\nAnd if that's true, is it normal practice to pass a class-wise binary, class-weighting?\n\n<br>\n\n---\n\n**Thanks in advance!!**",
      "votes": 9
    },
    {
      "id": 1212250,
      "postDate": "2021-02-21T03:42:23.807Z",
      "content": "<p>IMO there are 19 classes.  But we are hampered by a lack of very many ground truth images with only that class.  In another <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748\" target=\"_blank\">post</a> in this competition someone examined the \"negative\" images and as I recall felt even as a amateur they would have classified about half of them with one of our base 18.</p>\n<p>We do however have an huge number (once again IMO) of the negative cells in many images.  I am referring to the cells falling on the edge of the main image itself.  I would expect that my models should end up calling many of these negative.</p>\n<p>Comments from the hosts suggest that border cells are not likely to be annotated and in the hand correction of segments I assume this means no segment will be present in the private test set.</p>\n<p>There would seem to be several options for handling the edge of image cells - I currently plan to code towards them being classified as negative.  But I am open for other suggestions on how to handle partial cells from edges.</p>",
      "rawMarkdown": "IMO there are 19 classes.  But we are hampered by a lack of very many ground truth images with only that class.  In another [post](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748) in this competition someone examined the \"negative\" images and as I recall felt even as a amateur they would have classified about half of them with one of our base 18.\n\nWe do however have an huge number (once again IMO) of the negative cells in many images.  I am referring to the cells falling on the edge of the main image itself.  I would expect that my models should end up calling many of these negative.\n\nComments from the hosts suggest that border cells are not likely to be annotated and in the hand correction of segments I assume this means no segment will be present in the private test set.\n\nThere would seem to be several options for handling the edge of image cells - I currently plan to code towards them being classified as negative.  But I am open for other suggestions on how to handle partial cells from edges.",
      "votes": 3
    },
    {
      "id": 1212023,
      "postDate": "2021-02-20T19:38:03.740Z",
      "content": "<p>Good questions! My initial intuitions below:</p>\n<ul>\n<li>framing this as 19 vs. 18 class problem: both are feasible, I've started with 19, but am planning to experiment with 18 classes as well and compare the results</li>\n<li>for the negative examples, if we frame this as 18 class problem, then the target vector will be all zeros</li>\n<li>in multilabel problems, the loss function (e.g. BCE with logits loss) should evaluate each class independently (e.g. sigmoid applied for each class prediction) - so the model should learn decision boundary for each class. If it learns to classify the positive examples correctly for each class, it will be rewarded by lowering the loss, so that is an incentive not to predict all zeros all the time. Having said that, with few positive examples for rare classes, the model may have hard time learning the decision boundary, so adding class weighting, oversampling etc. may be a good idea :) </li>\n</ul>",
      "rawMarkdown": "Good questions! My initial intuitions below:\n- framing this as 19 vs. 18 class problem: both are feasible, I've started with 19, but am planning to experiment with 18 classes as well and compare the results\n- for the negative examples, if we frame this as 18 class problem, then the target vector will be all zeros\n- in multilabel problems, the loss function (e.g. BCE with logits loss) should evaluate each class independently (e.g. sigmoid applied for each class prediction) - so the model should learn decision boundary for each class. If it learns to classify the positive examples correctly for each class, it will be rewarded by lowering the loss, so that is an incentive not to predict all zeros all the time. Having said that, with few positive examples for rare classes, the model may have hard time learning the decision boundary, so adding class weighting, oversampling etc. may be a good idea :) ",
      "votes": 3,
      "replies": [
        {
          "id": 1212070,
          "postDate": "2021-02-20T21:00:39.587Z",
          "content": "<p>Thanks for the comment! </p>\n<hr>\n<ul>\n<li><p>Agreed it can be framed in either way. I have, up until now, been treating it as a 19 class problem. Now that I've created a new dataset with a plethora of <strong>Negative</strong> labels I will be pivoting to an 18 class problem (either, as you indicated, passing a vector of zeros as the label for the <strong>Negative</strong> images, or, as Ayush indicated, having a multi-stage pipeline where you first perform Binary Classification and then if the class is positive, perform multi-label classification using the 18 classes.</p></li>\n<li><p>Agreed</p></li>\n<li><p>I think I understand. I guess I'm just struggling with the implementation side of things. Normally, for a multi-class classification problem (not multi-label), if I have a class imbalance, I will pass a dictionary of class weights to my fit function. If I had a model with 18 outputs, where each was a sigmoid activation function, I would normally pass 18 class weighting dictionaries (<strong><code>{0:0.75, 1:0.25}</code></strong>) indicating the percentage of positives v. negatives for every class. However, when I build my network now, I have a single output layer of length 18. I can't pass individual class weight dictionaries… I think. I guess I just have to get more comfortable with multi-label class weighting in Tensorflow. Or, alternatively, I can create a custom loss function with the class weighting baked in.</p></li>\n</ul>\n<hr>\n<p>I'm currently training the two models to be used in the multi-stage approach… which should be finished in 6-8 hours. Following that, I will update my inference notebook and submit and we'll get some feedback on this new technique!</p>",
          "rawMarkdown": "Thanks for the comment! \n\n---\n\n- Agreed it can be framed in either way. I have, up until now, been treating it as a 19 class problem. Now that I've created a new dataset with a plethora of **Negative** labels I will be pivoting to an 18 class problem (either, as you indicated, passing a vector of zeros as the label for the **Negative** images, or, as Ayush indicated, having a multi-stage pipeline where you first perform Binary Classification and then if the class is positive, perform multi-label classification using the 18 classes.\n\n- Agreed\n\n- I think I understand. I guess I'm just struggling with the implementation side of things. Normally, for a multi-class classification problem (not multi-label), if I have a class imbalance, I will pass a dictionary of class weights to my fit function. If I had a model with 18 outputs, where each was a sigmoid activation function, I would normally pass 18 class weighting dictionaries (**`{0:0.75, 1:0.25}`**) indicating the percentage of positives v. negatives for every class. However, when I build my network now, I have a single output layer of length 18. I can't pass individual class weight dictionaries... I think. I guess I just have to get more comfortable with multi-label class weighting in Tensorflow. Or, alternatively, I can create a custom loss function with the class weighting baked in.\n\n---\n\nI'm currently training the two models to be used in the multi-stage approach... which should be finished in 6-8 hours. Following that, I will update my inference notebook and submit and we'll get some feedback on this new technique!\n",
          "votes": 1
        },
        {
          "id": 1212274,
          "postDate": "2021-02-21T04:13:14.593Z",
          "content": "<blockquote>\n  <p>{0:0.75, 1:0.25}</p>\n</blockquote>\n<p>Is there no way to pass a single dictionary having the weights? Something like {0:0.5, 1:0.7, 2:1,……,18:0.8}</p>",
          "rawMarkdown": "> {0:0.75, 1:0.25}\n\nIs there no way to pass a single dictionary having the weights? Something like {0:0.5, 1:0.7, 2:1,......,18:0.8}"
        },
        {
          "id": 1212384,
          "postDate": "2021-02-21T07:26:35.987Z",
          "content": "<p>I haven't done this myself yet. I started by creating a more balanced sample of train and use it to train my models. Planning to explore a balanced sampler as well, e.g. like this one: <a href=\"https://github.com/issamemari/pytorch-multilabel-balanced-sampler\" target=\"_blank\">https://github.com/issamemari/pytorch-multilabel-balanced-sampler</a></p>",
          "rawMarkdown": "I haven't done this myself yet. I started by creating a more balanced sample of train and use it to train my models. Planning to explore a balanced sampler as well, e.g. like this one: https://github.com/issamemari/pytorch-multilabel-balanced-sampler",
          "votes": 1
        }
      ]
    },
    {
      "id": 1212502,
      "postDate": "2021-02-21T09:29:51.773Z",
      "content": "<p>Hi. What's your CV strategy?<br>\nI think multilabel stratified kfold is good, but also we may have to consider multilabel stratified GROUP kfold because it is said that there are 17 different cell types.</p>",
      "rawMarkdown": "Hi. What's your CV strategy?\nI think multilabel stratified kfold is good, but also we may have to consider multilabel stratified GROUP kfold because it is said that there are 17 different cell types.",
      "votes": 1,
      "replies": [
        {
          "id": 1212566,
          "postDate": "2021-02-21T10:53:33.847Z",
          "content": "<p>For the classification models, I'm using MultilabelStratifiedKFold:  <br>\n<a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">https://github.com/trent-b/iterative-stratification</a>. For group kfold do you mean stratifying by cell line? We don't have that available in training though, do we? </p>",
          "rawMarkdown": "For the classification models, I'm using MultilabelStratifiedKFold:  \nhttps://github.com/trent-b/iterative-stratification. For group kfold do you mean stratifying by cell line? We don't have that available in training though, do we? ",
          "votes": 2
        },
        {
          "id": 1212573,
          "postDate": "2021-02-21T10:59:45.803Z",
          "content": "<p>Thanks for the reply.<br>\nSo, clustering is one choice.<br>\nIf there is no tendency that the same type of cells is easy to predict, we don't have to care about it :)</p>",
          "rawMarkdown": "Thanks for the reply.\nSo, clustering is one choice.\nIf there is no tendency that the same type of cells is easy to predict, we don't have to care about it :)",
          "votes": 2
        },
        {
          "id": 1212642,
          "postDate": "2021-02-21T12:33:28.143Z",
          "content": "<p>I have not implemented a CV strategy yet. I plan to use a strategy similar to the one Darek noted. That being said I will most likely implement something using <strong><code>tf.data</code></strong>.</p>\n<p>Clustering on cell-type and stratifying would be an excellent idea. Definitely, something I hadn't thought of! Thanks for suggesting it and I look forward to seeing it implemented. (I may give it a shot if I have the bandwidth 😄).</p>",
          "rawMarkdown": "I have not implemented a CV strategy yet. I plan to use a strategy similar to the one Darek noted. That being said I will most likely implement something using **`tf.data`**.\n\nClustering on cell-type and stratifying would be an excellent idea. Definitely, something I hadn't thought of! Thanks for suggesting it and I look forward to seeing it implemented. (I may give it a shot if I have the bandwidth 😄).",
          "votes": 2
        },
        {
          "id": 1234239,
          "postDate": "2021-03-11T04:23:09.877Z",
          "content": "<p>How do we distinguish cell lines?</p>",
          "rawMarkdown": "How do we distinguish cell lines?",
          "votes": 1
        },
        {
          "id": 1235121,
          "postDate": "2021-03-11T20:57:25.333Z",
          "content": "<p>Actually several days ago I made a simple cellline classifier using public data, but it didn't work well for me😂</p>",
          "rawMarkdown": "Actually several days ago I made a simple cellline classifier using public data, but it didn't work well for me😂",
          "votes": 2
        }
      ]
    },
    {
      "id": 1211887,
      "postDate": "2021-02-20T16:27:52.420Z",
      "content": "<p>If there were a sufficient number of negative examples I would have tried to do binary classification to classify negative from positive images. Given positive images do multi-label classification. But that's not the case, unfortunately. </p>\n<p>Formulating this as a 19 multi-label classification problem should be a good starting point in my opinion. The catch as you know is to associate a negative label to a cell. And most probably many cells would fall under the negative label but that's subject to further investigation.</p>\n<p>Hoping others have a better answer. :)</p>",
      "rawMarkdown": "If there were a sufficient number of negative examples I would have tried to do binary classification to classify negative from positive images. Given positive images do multi-label classification. But that's not the case, unfortunately. \n\nFormulating this as a 19 multi-label classification problem should be a good starting point in my opinion. The catch as you know is to associate a negative label to a cell. And most probably many cells would fall under the negative label but that's subject to further investigation.\n\nHoping others have a better answer. :)",
      "votes": 1,
      "replies": [
        {
          "id": 1211949,
          "postDate": "2021-02-20T17:31:26.430Z",
          "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> !</p>\n<p>I have created a cell-level dataset following the application of some heuristics to identify negative cells. There is now a large number of negative class cells. ~3 times more than the next highest class and thousands of times more than the lower classes.</p>\n<p>The idea of prefacing my multilabel with a negative/positive classifier is a good idea. I think I may try and implement that.</p>\n<p>I'd also like to see other options for solutions!</p>",
          "rawMarkdown": "Thanks for the reply @ayuraj !\n\nI have created a cell-level dataset following the application of some heuristics to identify negative cells. There is now a large number of negative class cells. ~3 times more than the next highest class and thousands of times more than the lower classes.\n\nThe idea of prefacing my multilabel with a negative/positive classifier is a good idea. I think I may try and implement that.\n\nI'd also like to see other options for solutions!",
          "votes": 1
        },
        {
          "id": 1212381,
          "postDate": "2021-02-21T07:22:30.377Z",
          "content": "<p>Looking forward to see your results! I've been playing with some heuristics to identify negative classes via probing the leaderboard, but unsuccessful so far…</p>",
          "rawMarkdown": "Looking forward to see your results! I've been playing with some heuristics to identify negative classes via probing the leaderboard, but unsuccessful so far..."
        },
        {
          "id": 1212473,
          "postDate": "2021-02-21T08:49:05.120Z",
          "content": "<p>If you want more negative images you can use the <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg#Downloading-HPA-public-data\" target=\"_blank\">HPA public images</a> where negative images are labeled \"No staining\".</p>",
          "rawMarkdown": "If you want more negative images you can use the [HPA public images](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg#Downloading-HPA-public-data) where negative images are labeled \"No staining\".",
          "votes": 4
        },
        {
          "id": 1212526,
          "postDate": "2021-02-21T09:54:58.930Z",
          "content": "<p>That would be helpful. Thank you for the info.</p>",
          "rawMarkdown": "That would be helpful. Thank you for the info."
        }
      ]
    },
    {
      "id": 1211985,
      "postDate": "2021-02-20T18:17:14.190Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1212063,
          "postDate": "2021-02-20T20:51:24.180Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/weka511\" target=\"_blank\">@weka511</a>, thanks for the comment. I have seen Darek's excellent notebook. He has given me quite a lot of inspiration in this competition.</p>\n<p>I agree with your points across the board and enjoy the Kuiper Belt analogy!</p>",
          "rawMarkdown": "Hi @weka511, thanks for the comment. I have seen Darek's excellent notebook. He has given me quite a lot of inspiration in this competition.\n\nI agree with your points across the board and enjoy the Kuiper Belt analogy!",
          "votes": 1
        },
        {
          "id": 1212168,
          "postDate": "2021-02-21T00:26:59.073Z",
          "content": "<p><a href=\"https://www.kaggle.com/weka511\" target=\"_blank\">Simon</a></p>\n<p>There are more than 18 - the list on the HPA site shows 35.  They have some imaginative names :)  It could really be an unbalanced problem if we needed to indentify all 35.  I agree with your guess that they are Kuiper Belt members of a cell.</p>\n<blockquote>\n  <p>Protein localization data is derived from antibody-based profiling, using immunofluorescence (ICC-IF) and confocal microscopy, and classified into 35 different organelles and fine subcellular structures.&gt; </p>\n</blockquote>",
          "rawMarkdown": "[Simon](https://www.kaggle.com/weka511)\n\n There are more than 18 - the list on the HPA site shows 35.  They have some imaginative names :)  It could really be an unbalanced problem if we needed to indentify all 35.  I agree with your guess that they are Kuiper Belt members of a cell.\n\n> Protein localization data is derived from antibody-based profiling, using immunofluorescence (ICC-IF) and confocal microscopy, and classified into 35 different organelles and fine subcellular structures.> "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1212250,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-21T03:42:23.807000",
      "content": "<p>IMO there are 19 classes.  But we are hampered by a lack of very many ground truth images with only that class.  In another <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748\" target=\"_blank\">post</a> in this competition someone examined the \"negative\" images and as I recall felt even as a amateur they would have classified about half of them with one of our base 18.</p>\n<p>We do however have an huge number (once again IMO) of the negative cells in many images.  I am referring to the cells falling on the edge of the main image itself.  I would expect that my models should end up calling many of these negative.</p>\n<p>Comments from the hosts suggest that border cells are not likely to be annotated and in the hand correction of segments I assume this means no segment will be present in the private test set.</p>\n<p>There would seem to be several options for handling the edge of image cells - I currently plan to code towards them being classified as negative.  But I am open for other suggestions on how to handle partial cells from edges.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1212023,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-02-20T19:38:03.740000",
      "content": "<p>Good questions! My initial intuitions below:</p>\n<ul>\n<li>framing this as 19 vs. 18 class problem: both are feasible, I've started with 19, but am planning to experiment with 18 classes as well and compare the results</li>\n<li>for the negative examples, if we frame this as 18 class problem, then the target vector will be all zeros</li>\n<li>in multilabel problems, the loss function (e.g. BCE with logits loss) should evaluate each class independently (e.g. sigmoid applied for each class prediction) - so the model should learn decision boundary for each class. If it learns to classify the positive examples correctly for each class, it will be rewarded by lowering the loss, so that is an incentive not to predict all zeros all the time. Having said that, with few positive examples for rare classes, the model may have hard time learning the decision boundary, so adding class weighting, oversampling etc. may be a good idea :) </li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 1212070,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-20T21:00:39.587000",
          "content": "<p>Thanks for the comment! </p>\n<hr>\n<ul>\n<li><p>Agreed it can be framed in either way. I have, up until now, been treating it as a 19 class problem. Now that I've created a new dataset with a plethora of <strong>Negative</strong> labels I will be pivoting to an 18 class problem (either, as you indicated, passing a vector of zeros as the label for the <strong>Negative</strong> images, or, as Ayush indicated, having a multi-stage pipeline where you first perform Binary Classification and then if the class is positive, perform multi-label classification using the 18 classes.</p></li>\n<li><p>Agreed</p></li>\n<li><p>I think I understand. I guess I'm just struggling with the implementation side of things. Normally, for a multi-class classification problem (not multi-label), if I have a class imbalance, I will pass a dictionary of class weights to my fit function. If I had a model with 18 outputs, where each was a sigmoid activation function, I would normally pass 18 class weighting dictionaries (<strong><code>{0:0.75, 1:0.25}</code></strong>) indicating the percentage of positives v. negatives for every class. However, when I build my network now, I have a single output layer of length 18. I can't pass individual class weight dictionaries… I think. I guess I just have to get more comfortable with multi-label class weighting in Tensorflow. Or, alternatively, I can create a custom loss function with the class weighting baked in.</p></li>\n</ul>\n<hr>\n<p>I'm currently training the two models to be used in the multi-stage approach… which should be finished in 6-8 hours. Following that, I will update my inference notebook and submit and we'll get some feedback on this new technique!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212274,
          "author_name": "Arka Saha",
          "author_url": "",
          "post_date": "2021-02-21T04:13:14.593000",
          "content": "<blockquote>\n  <p>{0:0.75, 1:0.25}</p>\n</blockquote>\n<p>Is there no way to pass a single dictionary having the weights? Something like {0:0.5, 1:0.7, 2:1,……,18:0.8}</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212384,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-02-21T07:26:35.987000",
          "content": "<p>I haven't done this myself yet. I started by creating a more balanced sample of train and use it to train my models. Planning to explore a balanced sampler as well, e.g. like this one: <a href=\"https://github.com/issamemari/pytorch-multilabel-balanced-sampler\" target=\"_blank\">https://github.com/issamemari/pytorch-multilabel-balanced-sampler</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1212502,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-02-21T09:29:51.773000",
      "content": "<p>Hi. What's your CV strategy?<br>\nI think multilabel stratified kfold is good, but also we may have to consider multilabel stratified GROUP kfold because it is said that there are 17 different cell types.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1212566,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-02-21T10:53:33.847000",
          "content": "<p>For the classification models, I'm using MultilabelStratifiedKFold:  <br>\n<a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">https://github.com/trent-b/iterative-stratification</a>. For group kfold do you mean stratifying by cell line? We don't have that available in training though, do we? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1212573,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-02-21T10:59:45.803000",
          "content": "<p>Thanks for the reply.<br>\nSo, clustering is one choice.<br>\nIf there is no tendency that the same type of cells is easy to predict, we don't have to care about it :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1212642,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-21T12:33:28.143000",
          "content": "<p>I have not implemented a CV strategy yet. I plan to use a strategy similar to the one Darek noted. That being said I will most likely implement something using <strong><code>tf.data</code></strong>.</p>\n<p>Clustering on cell-type and stratifying would be an excellent idea. Definitely, something I hadn't thought of! Thanks for suggesting it and I look forward to seeing it implemented. (I may give it a shot if I have the bandwidth 😄).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1234239,
          "author_name": "Correlation",
          "author_url": "",
          "post_date": "2021-03-11T04:23:09.877000",
          "content": "<p>How do we distinguish cell lines?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1235121,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-03-11T20:57:25.333000",
          "content": "<p>Actually several days ago I made a simple cellline classifier using public data, but it didn't work well for me😂</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1211887,
      "author_name": "Ayush Thakur",
      "author_url": "",
      "post_date": "2021-02-20T16:27:52.420000",
      "content": "<p>If there were a sufficient number of negative examples I would have tried to do binary classification to classify negative from positive images. Given positive images do multi-label classification. But that's not the case, unfortunately. </p>\n<p>Formulating this as a 19 multi-label classification problem should be a good starting point in my opinion. The catch as you know is to associate a negative label to a cell. And most probably many cells would fall under the negative label but that's subject to further investigation.</p>\n<p>Hoping others have a better answer. :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1211949,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-20T17:31:26.430000",
          "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> !</p>\n<p>I have created a cell-level dataset following the application of some heuristics to identify negative cells. There is now a large number of negative class cells. ~3 times more than the next highest class and thousands of times more than the lower classes.</p>\n<p>The idea of prefacing my multilabel with a negative/positive classifier is a good idea. I think I may try and implement that.</p>\n<p>I'd also like to see other options for solutions!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212381,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-02-21T07:22:30.377000",
          "content": "<p>Looking forward to see your results! I've been playing with some heuristics to identify negative classes via probing the leaderboard, but unsuccessful so far…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212473,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-02-21T08:49:05.120000",
          "content": "<p>If you want more negative images you can use the <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg#Downloading-HPA-public-data\" target=\"_blank\">HPA public images</a> where negative images are labeled \"No staining\".</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1212526,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-02-21T09:54:58.930000",
          "content": "<p>That would be helpful. Thank you for the info.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1211985,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-20T18:17:14.190000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1212063,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-20T20:51:24.180000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/weka511\" target=\"_blank\">@weka511</a>, thanks for the comment. I have seen Darek's excellent notebook. He has given me quite a lot of inspiration in this competition.</p>\n<p>I agree with your points across the board and enjoy the Kuiper Belt analogy!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212168,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-02-21T00:26:59.073000",
          "content": "<p><a href=\"https://www.kaggle.com/weka511\" target=\"_blank\">Simon</a></p>\n<p>There are more than 18 - the list on the HPA site shows 35.  They have some imaginative names :)  It could really be an unbalanced problem if we needed to indentify all 35.  I agree with your guess that they are Kuiper Belt members of a cell.</p>\n<blockquote>\n  <p>Protein localization data is derived from antibody-based profiling, using immunofluorescence (ICC-IF) and confocal microscopy, and classified into 35 different organelles and fine subcellular structures.&gt; </p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1211873": "Hi there.\n\nI think this is a pretty basic question but I wanted to ask here just in case anyone else had the same inquiry (I'm not afraid to ask the simple questions 😅). I did some Googling and have my own speculations but I always appreciate the wisdom of other Kagglers... so here we go.\n\n---\n\n<br>\n\n**SETUP:** \n\nI'm training a multi-label cell classifier.\nWe have 19 classes (18 organelle structures, and a negative label). \n\n<br>\n\n---\n\n<br>\n\n**QUESTION:**\n\nDo I frame this as a 19 class multi-label classification problem or an 18 class classification problem where a lack of confidence in the 18 organelle classes indicates the presence of the negative class?\n\n<br>\n\n---\n\n<br>\n\n**FOLLOW-UP 1:**\n\nLet's say I decide to frame it as an 18 class multi-label classification problem. Therefore, during training, if I have ***n*** negative training samples. \n\nKnowing this, should I pass the ground-truth label as a vector of length 18 full of 0s for all ***n*** negative training samples?\n\n<br>\n\n---\n\n<br>\n\n**FOLLOW-UP 2:**\n\nDuring the training of a multi-label classification model with more than a few classes, it's obvious that most of the time the majority of the outputs will be 0.\n\ni.e. Let's look at 5 outputs for a 5 class model (assuming no negative classes which would only exacerbate the problem).\n\n```python\n\nmodel_output_1 = [0.05, 0.23, 0.10, 0.80, 0.10]\nground_truth_1  = [0.00, 0.00, 0.00, 1.00, 0.00,]\n\nmodel_output_2 = [0.95, 0.02, 0.04, 0.02, 0.10]\nground_truth_2  = [1.00, 0.00, 0.00, 0.00, 0.00]\n\nmodel_output_3 = [0.02, 0.66, 0.09, 0.08, 0.13]\nground_truth_3  = [0.00, 1.00, 0.00, 0.00, 0.00,]\n\nmodel_output_4 = [0.15, 0.03, 0.08, 0.38, 0.55]\nground_truth_4  = [0.00, 0.00, 0.00, 0.00, 1.00,]\n\nmodel_output_5 = [0.22, 0.04, 0.98, 0.32, 0.12]\nground_truth_5  = [0.00, 0.00, 1.00, 0.00, 0.00,]\n\n```\n\nWould it not be true in this case that the model will just eventually learn to predict all 0s all the time? Without a class weighting function that is?\n\nAnd if that's true, is it normal practice to pass a class-wise binary, class-weighting?\n\n<br>\n\n---\n\n**Thanks in advance!!**",
    "1212250": "IMO there are 19 classes.  But we are hampered by a lack of very many ground truth images with only that class.  In another [post](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748) in this competition someone examined the \"negative\" images and as I recall felt even as a amateur they would have classified about half of them with one of our base 18.\n\nWe do however have an huge number (once again IMO) of the negative cells in many images.  I am referring to the cells falling on the edge of the main image itself.  I would expect that my models should end up calling many of these negative.\n\nComments from the hosts suggest that border cells are not likely to be annotated and in the hand correction of segments I assume this means no segment will be present in the private test set.\n\nThere would seem to be several options for handling the edge of image cells - I currently plan to code towards them being classified as negative.  But I am open for other suggestions on how to handle partial cells from edges.",
    "1212023": "Good questions! My initial intuitions below:\n- framing this as 19 vs. 18 class problem: both are feasible, I've started with 19, but am planning to experiment with 18 classes as well and compare the results\n- for the negative examples, if we frame this as 18 class problem, then the target vector will be all zeros\n- in multilabel problems, the loss function (e.g. BCE with logits loss) should evaluate each class independently (e.g. sigmoid applied for each class prediction) - so the model should learn decision boundary for each class. If it learns to classify the positive examples correctly for each class, it will be rewarded by lowering the loss, so that is an incentive not to predict all zeros all the time. Having said that, with few positive examples for rare classes, the model may have hard time learning the decision boundary, so adding class weighting, oversampling etc. may be a good idea :) ",
    "1212502": "Hi. What's your CV strategy?\nI think multilabel stratified kfold is good, but also we may have to consider multilabel stratified GROUP kfold because it is said that there are 17 different cell types.",
    "1211887": "If there were a sufficient number of negative examples I would have tried to do binary classification to classify negative from positive images. Given positive images do multi-label classification. But that's not the case, unfortunately. \n\nFormulating this as a 19 multi-label classification problem should be a good starting point in my opinion. The catch as you know is to associate a negative label to a cell. And most probably many cells would fall under the negative label but that's subject to further investigation.\n\nHoping others have a better answer. :)",
    "1211985": ""
  }
}