{
  "id": 76303,
  "title": "the hidden tricks",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/76303",
  "author_name": "hengck23",
  "post_date": "2018-12-31T20:56:50.930000",
  "votes": 18,
  "comment_count": 6,
  "views": 0,
  "content": "<p>i will list more as i discover more tricks ... here is the first ones:</p>\n\n<ol>\n<li><p>image size (1024, 2048) as addition input information . e.g. labels of test image-2048 are more likely to come from train image-2048 than train image-1024?</p></li>\n<li><p>rare classes : you should love them! first you are going to rank better if you can handle them  because most kagglers tend to score poorly on them. It is difficult to apply \"pure machine learning\" on them due to small data size. Use domain knowledge (e.g. co-occurrence of class labels). Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc (since they are rare, there are only to be few images and and annotations don't take long time)</p></li>\n<li><p>size of objects/blobs/cells</p></li>\n<li><p>balanced sampling : balance between classes, between  single/multiple labels, between external/kaggle data</p></li>\n<li><p>there are a few misaligned images (the red,green,blue,yellow channels are offset). if these are from rare classes, you will gain good score. </p></li>\n<li><p>multi-label images : most kagglers will predict this with more mistakes. This is because multi-label images are less than single-label images in training.  Use domain knowledge (e.g. co-occurrence of class labels) to improve. you can use fancy methods like rnn/lstm or stacking of networks (aka. bestfitting trick)</p></li>\n<li><p>thresholding for ensemble: for each model, construct precision/recall or threshold/F1 curve for a class. Average the curves for different models in the ensemble. determine threshold based on the averaged curve. \n...</p></li>\n</ol>",
  "messages": [
    {
      "id": 448349,
      "postDate": "2018-12-31T20:56:50.930Z",
      "content": "<p>i will list more as i discover more tricks ... here is the first ones:</p>\n\n<ol>\n<li><p>image size (1024, 2048) as addition input information . e.g. labels of test image-2048 are more likely to come from train image-2048 than train image-1024?</p></li>\n<li><p>rare classes : you should love them! first you are going to rank better if you can handle them  because most kagglers tend to score poorly on them. It is difficult to apply \"pure machine learning\" on them due to small data size. Use domain knowledge (e.g. co-occurrence of class labels). Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc (since they are rare, there are only to be few images and and annotations don't take long time)</p></li>\n<li><p>size of objects/blobs/cells</p></li>\n<li><p>balanced sampling : balance between classes, between  single/multiple labels, between external/kaggle data</p></li>\n<li><p>there are a few misaligned images (the red,green,blue,yellow channels are offset). if these are from rare classes, you will gain good score. </p></li>\n<li><p>multi-label images : most kagglers will predict this with more mistakes. This is because multi-label images are less than single-label images in training.  Use domain knowledge (e.g. co-occurrence of class labels) to improve. you can use fancy methods like rnn/lstm or stacking of networks (aka. bestfitting trick)</p></li>\n<li><p>thresholding for ensemble: for each model, construct precision/recall or threshold/F1 curve for a class. Average the curves for different models in the ensemble. determine threshold based on the averaged curve. \n...</p></li>\n</ol>",
      "rawMarkdown": "i will list more as i discover more tricks ... here is the first ones:\n\n1. image size (1024, 2048) as addition input information . e.g. labels of test image-2048 are more likely to come from train image-2048 than train image-1024?\n\n2. rare classes : you should love them! first you are going to rank better if you can handle them  because most kagglers tend to score poorly on them. It is difficult to apply \"pure machine learning\" on them due to small data size. Use domain knowledge (e.g. co-occurrence of class labels). Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc (since they are rare, there are only to be few images and and annotations don't take long time)\n\n3. size of objects/blobs/cells\n\n4. balanced sampling : balance between classes, between  single/multiple labels, between external/kaggle data\n\n5. there are a few misaligned images (the red,green,blue,yellow channels are offset). if these are from rare classes, you will gain good score. \n\n6. multi-label images : most kagglers will predict this with more mistakes. This is because multi-label images are less than single-label images in training.  Use domain knowledge (e.g. co-occurrence of class labels) to improve. you can use fancy methods like rnn/lstm or stacking of networks (aka. bestfitting trick)\n\n7. thresholding for ensemble: for each model, construct precision/recall or threshold/F1 curve for a class. Average the curves for different models in the ensemble. determine threshold based on the averaged curve. \n ...",
      "votes": 17
    },
    {
      "id": 449473,
      "postDate": "2019-01-03T07:41:26.323Z",
      "content": "<p>Regarding the \"Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc\" in the point number 2, how would you incorporate this into a model? Since only the rare labels annotated would we need a different model for the rare annotated labels, or is there a way to train the rare annotated and non-annotated labels? One way that I could think of is to use one head for classification and one for object detection -&gt; classification, then during training, the loss of the object detection head could be set to 0 for the classes that aren't annotated, am I going in the right direction? Also, what is the bestfitting trick that you are referring to?</p>",
      "rawMarkdown": "Regarding the \"Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc\" in the point number 2, how would you incorporate this into a model? Since only the rare labels annotated would we need a different model for the rare annotated labels, or is there a way to train the rare annotated and non-annotated labels? One way that I could think of is to use one head for classification and one for object detection -&gt; classification, then during training, the loss of the object detection head could be set to 0 for the classes that aren't annotated, am I going in the right direction? Also, what is the bestfitting trick that you are referring to?"
    },
    {
      "id": 448711,
      "postDate": "2019-01-01T22:48:42.520Z",
      "content": "<p>Hi Heng. Thank for your post.\nThough I have not succeded in the competition yet let me make  comments on some points.\n3. It is stated, that protein could be from different part of body cells, so the same protein can be in the cells of different size. Won't tacking into account the size of cell cause overfittion to the availible cell types? \n6. There might not be enough data for stacking networks, especially this will affect rare classes. Moreover, multilabel classification CNN model trained on its own account for correlation between the classes. Correct me if I am wrong. Maybe stacking will help in case of lots of data avalible, but I can not see feasible imporvement. I made experiments on pretrained features and tried to train lightGBM on  them. Though I got imporvement on my val part ( I devided set into 1. Train data for  CNN 2. val  data for CNN. 2.1. Train 2nd level model 2.2. Validate 2nd level model. , so I train 2nd level model on part of val data), I saw decrease in public LB, so I think , that I overfitted on  val data. Did you have any progress using 2nd level model or stacking? </p>",
      "rawMarkdown": "Hi Heng. Thank for your post.\nThough I have not succeded in the competition yet let me make  comments on some points.\n3. It is stated, that protein could be from different part of body cells, so the same protein can be in the cells of different size. Won't tacking into account the size of cell cause overfittion to the availible cell types? \n6. There might not be enough data for stacking networks, especially this will affect rare classes. Moreover, multilabel classification CNN model trained on its own account for correlation between the classes. Correct me if I am wrong. Maybe stacking will help in case of lots of data avalible, but I can not see feasible imporvement. I made experiments on pretrained features and tried to train lightGBM on  them. Though I got imporvement on my val part ( I devided set into 1. Train data for  CNN 2. val  data for CNN. 2.1. Train 2nd level model 2.2. Validate 2nd level model. , so I train 2nd level model on part of val data), I saw decrease in public LB, so I think , that I overfitted on  val data. Did you have any progress using 2nd level model or stacking? "
    },
    {
      "id": 448367,
      "postDate": "2018-12-31T22:16:58.080Z",
      "content": "<p>What batch size did you need to train the larger image sizes effectively? Anything that needed to change from 512 model to make larger image sizes work?</p>",
      "rawMarkdown": "What batch size did you need to train the larger image sizes effectively? Anything that needed to change from 512 model to make larger image sizes work?",
      "replies": [
        {
          "id": 448431,
          "postDate": "2019-01-01T06:18:50.470Z",
          "content": "<p>Train on crops of larger image first. Then freezes the bottom layers and finetune on full image size.</p>\n\n<p>Alternatively, if you can get rid of batch norm ( eg using group norm etc) you can use gradient accumulation during training. Small batch size is not a problem then</p>",
          "rawMarkdown": "Train on crops of larger image first. Then freezes the bottom layers and finetune on full image size.\n\n\nAlternatively, if you can get rid of batch norm ( eg using group norm etc) you can use gradient accumulation during training. Small batch size is not a problem then",
          "votes": 3
        },
        {
          "id": 448785,
          "postDate": "2019-01-02T05:00:55.080Z",
          "content": "<p>Thanks for the ideas. I was trying gradient accumulation yesterday but was unaware of the problems batch norm would give me. I'll look into group norm. </p>",
          "rawMarkdown": "Thanks for the ideas. I was trying gradient accumulation yesterday but was unaware of the problems batch norm would give me. I'll look into group norm. "
        }
      ]
    },
    {
      "id": 448761,
      "postDate": "2019-01-02T03:15:30.977Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 449473,
      "author_name": "Kevin Lu",
      "author_url": "",
      "post_date": "2019-01-03T07:41:26.323000",
      "content": "<p>Regarding the \"Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc\" in the point number 2, how would you incorporate this into a model? Since only the rare labels annotated would we need a different model for the rare annotated labels, or is there a way to train the rare annotated and non-annotated labels? One way that I could think of is to use one head for classification and one for object detection -&gt; classification, then during training, the loss of the object detection head could be set to 0 for the classes that aren't annotated, am I going in the right direction? Also, what is the bestfitting trick that you are referring to?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448711,
      "author_name": "Vaagn Minasian",
      "author_url": "",
      "post_date": "2019-01-01T22:48:42.520000",
      "content": "<p>Hi Heng. Thank for your post.\nThough I have not succeded in the competition yet let me make  comments on some points.\n3. It is stated, that protein could be from different part of body cells, so the same protein can be in the cells of different size. Won't tacking into account the size of cell cause overfittion to the availible cell types? \n6. There might not be enough data for stacking networks, especially this will affect rare classes. Moreover, multilabel classification CNN model trained on its own account for correlation between the classes. Correct me if I am wrong. Maybe stacking will help in case of lots of data avalible, but I can not see feasible imporvement. I made experiments on pretrained features and tried to train lightGBM on  them. Though I got imporvement on my val part ( I devided set into 1. Train data for  CNN 2. val  data for CNN. 2.1. Train 2nd level model 2.2. Validate 2nd level model. , so I train 2nd level model on part of val data), I saw decrease in public LB, so I think , that I overfitted on  val data. Did you have any progress using 2nd level model or stacking? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448367,
      "author_name": "David Wagner",
      "author_url": "",
      "post_date": "2018-12-31T22:16:58.080000",
      "content": "<p>What batch size did you need to train the larger image sizes effectively? Anything that needed to change from 512 model to make larger image sizes work?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 448431,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-01-01T06:18:50.470000",
          "content": "<p>Train on crops of larger image first. Then freezes the bottom layers and finetune on full image size.</p>\n\n<p>Alternatively, if you can get rid of batch norm ( eg using group norm etc) you can use gradient accumulation during training. Small batch size is not a problem then</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 448785,
          "author_name": "David Wagner",
          "author_url": "",
          "post_date": "2019-01-02T05:00:55.080000",
          "content": "<p>Thanks for the ideas. I was trying gradient accumulation yesterday but was unaware of the problems batch norm would give me. I'll look into group norm. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 448761,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-02T03:15:30.977000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "448349": "i will list more as i discover more tricks ... here is the first ones:\n\n1. image size (1024, 2048) as addition input information . e.g. labels of test image-2048 are more likely to come from train image-2048 than train image-1024?\n\n2. rare classes : you should love them! first you are going to rank better if you can handle them  because most kagglers tend to score poorly on them. It is difficult to apply \"pure machine learning\" on them due to small data size. Use domain knowledge (e.g. co-occurrence of class labels). Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc (since they are rare, there are only to be few images and and annotations don't take long time)\n\n3. size of objects/blobs/cells\n\n4. balanced sampling : balance between classes, between  single/multiple labels, between external/kaggle data\n\n5. there are a few misaligned images (the red,green,blue,yellow channels are offset). if these are from rare classes, you will gain good score. \n\n6. multi-label images : most kagglers will predict this with more mistakes. This is because multi-label images are less than single-label images in training.  Use domain knowledge (e.g. co-occurrence of class labels) to improve. you can use fancy methods like rnn/lstm or stacking of networks (aka. bestfitting trick)\n\n7. thresholding for ensemble: for each model, construct precision/recall or threshold/F1 curve for a class. Average the curves for different models in the ensemble. determine threshold based on the averaged curve. \n ...",
    "449473": "Regarding the \"Provide more supervisory signals, e.g. hand segment the train data and treat them as object detection or segmentation etc\" in the point number 2, how would you incorporate this into a model? Since only the rare labels annotated would we need a different model for the rare annotated labels, or is there a way to train the rare annotated and non-annotated labels? One way that I could think of is to use one head for classification and one for object detection -&gt; classification, then during training, the loss of the object detection head could be set to 0 for the classes that aren't annotated, am I going in the right direction? Also, what is the bestfitting trick that you are referring to?",
    "448711": "Hi Heng. Thank for your post.\nThough I have not succeded in the competition yet let me make  comments on some points.\n3. It is stated, that protein could be from different part of body cells, so the same protein can be in the cells of different size. Won't tacking into account the size of cell cause overfittion to the availible cell types? \n6. There might not be enough data for stacking networks, especially this will affect rare classes. Moreover, multilabel classification CNN model trained on its own account for correlation between the classes. Correct me if I am wrong. Maybe stacking will help in case of lots of data avalible, but I can not see feasible imporvement. I made experiments on pretrained features and tried to train lightGBM on  them. Though I got imporvement on my val part ( I devided set into 1. Train data for  CNN 2. val  data for CNN. 2.1. Train 2nd level model 2.2. Validate 2nd level model. , so I train 2nd level model on part of val data), I saw decrease in public LB, so I think , that I overfitted on  val data. Did you have any progress using 2nd level model or stacking? ",
    "448367": "What batch size did you need to train the larger image sizes effectively? Anything that needed to change from 512 model to make larger image sizes work?",
    "448761": ""
  }
}