{
  "id": 238385,
  "title": "18th place solution private 0.515 public 0.514",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238385",
  "author_name": "yuvaramsingh",
  "post_date": "2021-05-12T04:27:21.095000",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to kaggle and the organizers for putting together a competition on weakly supervised learning approach. learned a lot from this experience.</p>\n<h2>solution overview</h2>\n<p>since this is a weakly supervised learning task and HPA cell segment model was able to give instance mask, i chose to use Multiple instance learning method here. this made sense as the image level labels are not precise and through discussion it was clear that there could be cells that are negative class but not mentioned on the image level label. </p>\n<h2>stage 0 model</h2>\n<p>Here i used Effb4 and resnest50 as the CNN backbone. the training setup consist of extracting individual cells from an image and selecting 16 cells (max, else resampling to match 16 count) at random to be used as Bag of images. the target is same as image level label. here i use 16 cells just to make sure i have at least some instance of cells corresponding to the label. (in stage 0 we cant be sure on cell labels) </p>\n<p>CNN Feature Extractor : Effb4, resnest 50<br>\ncell used per image : 16<br>\nPooling layer (per cell): GAP<br>\nAttention pooling layer (Per image, 16 cells)<br>\nimage level label<br>\nFocal loss<br>\nTime : roughly 1hr per epoch </p>\n<p>The magic of this approach is the attention pooling layer which weights the 16 cell feature embedding by providing importance score based on the contribution of individual cells to the image level label. this enabled me to train a model that can find the appropriate cell level label by looking at the image level label. </p>\n<h2>stage 1 model</h2>\n<p>after training stage 0 model, now we have a model that is capable of providing cell level protein score that can be submitted to the competition. along with this we also have a well trained attention layer that can serve as a module to score individual cells. this can be used as a pseudo label generator. i used this approach to find the missing cell labels in a image and append this to the image level label. this helped with finding images which consist of Negative classes but not labeled correctly. This also helped with reducing the cell count per image as now we are more sure about the class of cells present on the image.</p>\n<p>CNN Feature Extractor : Effb4, resnest 50<br>\ncell used per image : 8<br>\nPooling layer (per cell): Transformer based attention pooling<br>\nAttention pooling layer (Per image, 8 cells)<br>\nimage level label + pseudo label<br>\nFocal loss<br>\nTime : roughly 20min per epoch </p>\n<p>we can clearly see how the training time dropped because of reducing the no of cells.  </p>\n<h2>final inference</h2>\n<p>My final inference consist of  5 model where 3 stage 0 and 2 stage 1 models were used. i did use 4 TTA to make my model prediction robust</p>\n<p>i am happy that my public and private score for this submission did not change a lot Private 0.515 public 0.514. i did suffer from accuracy drop due to GPU hardware change. i trained my model on volta architecture GPU and Kaggle uses P100 (pascal). i got a scored of 0.524 on public dataset by generating the submission file from volta GPU. this was a surprise for me. i did try to calculate the mean of difference between the GPU submission files and used it to correct my final submission score. it did help me a bit but i lost a lot of score just because of Hardware difference. </p>\n<h2>final thoughts</h2>\n<p>overall i got a good experience with working on Weakly supervised learning problem. Thanks again</p>",
  "messages": [
    {
      "id": 1303436,
      "postDate": "2021-05-12T04:27:21.097Z",
      "content": "<p>Thanks to kaggle and the organizers for putting together a competition on weakly supervised learning approach. learned a lot from this experience.</p>\n<h2>solution overview</h2>\n<p>since this is a weakly supervised learning task and HPA cell segment model was able to give instance mask, i chose to use Multiple instance learning method here. this made sense as the image level labels are not precise and through discussion it was clear that there could be cells that are negative class but not mentioned on the image level label. </p>\n<h2>stage 0 model</h2>\n<p>Here i used Effb4 and resnest50 as the CNN backbone. the training setup consist of extracting individual cells from an image and selecting 16 cells (max, else resampling to match 16 count) at random to be used as Bag of images. the target is same as image level label. here i use 16 cells just to make sure i have at least some instance of cells corresponding to the label. (in stage 0 we cant be sure on cell labels) </p>\n<p>CNN Feature Extractor : Effb4, resnest 50<br>\ncell used per image : 16<br>\nPooling layer (per cell): GAP<br>\nAttention pooling layer (Per image, 16 cells)<br>\nimage level label<br>\nFocal loss<br>\nTime : roughly 1hr per epoch </p>\n<p>The magic of this approach is the attention pooling layer which weights the 16 cell feature embedding by providing importance score based on the contribution of individual cells to the image level label. this enabled me to train a model that can find the appropriate cell level label by looking at the image level label. </p>\n<h2>stage 1 model</h2>\n<p>after training stage 0 model, now we have a model that is capable of providing cell level protein score that can be submitted to the competition. along with this we also have a well trained attention layer that can serve as a module to score individual cells. this can be used as a pseudo label generator. i used this approach to find the missing cell labels in a image and append this to the image level label. this helped with finding images which consist of Negative classes but not labeled correctly. This also helped with reducing the cell count per image as now we are more sure about the class of cells present on the image.</p>\n<p>CNN Feature Extractor : Effb4, resnest 50<br>\ncell used per image : 8<br>\nPooling layer (per cell): Transformer based attention pooling<br>\nAttention pooling layer (Per image, 8 cells)<br>\nimage level label + pseudo label<br>\nFocal loss<br>\nTime : roughly 20min per epoch </p>\n<p>we can clearly see how the training time dropped because of reducing the no of cells.  </p>\n<h2>final inference</h2>\n<p>My final inference consist of  5 model where 3 stage 0 and 2 stage 1 models were used. i did use 4 TTA to make my model prediction robust</p>\n<p>i am happy that my public and private score for this submission did not change a lot Private 0.515 public 0.514. i did suffer from accuracy drop due to GPU hardware change. i trained my model on volta architecture GPU and Kaggle uses P100 (pascal). i got a scored of 0.524 on public dataset by generating the submission file from volta GPU. this was a surprise for me. i did try to calculate the mean of difference between the GPU submission files and used it to correct my final submission score. it did help me a bit but i lost a lot of score just because of Hardware difference. </p>\n<h2>final thoughts</h2>\n<p>overall i got a good experience with working on Weakly supervised learning problem. Thanks again</p>",
      "rawMarkdown": "Thanks to kaggle and the organizers for putting together a competition on weakly supervised learning approach. learned a lot from this experience.\n\n## solution overview\nsince this is a weakly supervised learning task and HPA cell segment model was able to give instance mask, i chose to use Multiple instance learning method here. this made sense as the image level labels are not precise and through discussion it was clear that there could be cells that are negative class but not mentioned on the image level label. \n\n## stage 0 model\nHere i used Effb4 and resnest50 as the CNN backbone. the training setup consist of extracting individual cells from an image and selecting 16 cells (max, else resampling to match 16 count) at random to be used as Bag of images. the target is same as image level label. here i use 16 cells just to make sure i have at least some instance of cells corresponding to the label. (in stage 0 we cant be sure on cell labels) \n\nCNN Feature Extractor : Effb4, resnest 50\ncell used per image : 16\nPooling layer (per cell): GAP\nAttention pooling layer (Per image, 16 cells)\nimage level label\nFocal loss\nTime : roughly 1hr per epoch \n\nThe magic of this approach is the attention pooling layer which weights the 16 cell feature embedding by providing importance score based on the contribution of individual cells to the image level label. this enabled me to train a model that can find the appropriate cell level label by looking at the image level label. \n\n## stage 1 model\nafter training stage 0 model, now we have a model that is capable of providing cell level protein score that can be submitted to the competition. along with this we also have a well trained attention layer that can serve as a module to score individual cells. this can be used as a pseudo label generator. i used this approach to find the missing cell labels in a image and append this to the image level label. this helped with finding images which consist of Negative classes but not labeled correctly. This also helped with reducing the cell count per image as now we are more sure about the class of cells present on the image.\n\nCNN Feature Extractor : Effb4, resnest 50\ncell used per image : 8\nPooling layer (per cell): Transformer based attention pooling\nAttention pooling layer (Per image, 8 cells)\nimage level label + pseudo label\nFocal loss\nTime : roughly 20min per epoch \n\nwe can clearly see how the training time dropped because of reducing the no of cells.  \n\n## final inference \nMy final inference consist of  5 model where 3 stage 0 and 2 stage 1 models were used. i did use 4 TTA to make my model prediction robust\n\ni am happy that my public and private score for this submission did not change a lot Private 0.515 public 0.514. i did suffer from accuracy drop due to GPU hardware change. i trained my model on volta architecture GPU and Kaggle uses P100 (pascal). i got a scored of 0.524 on public dataset by generating the submission file from volta GPU. this was a surprise for me. i did try to calculate the mean of difference between the GPU submission files and used it to correct my final submission score. it did help me a bit but i lost a lot of score just because of Hardware difference. \n\n## final thoughts\noverall i got a good experience with working on Weakly supervised learning problem. Thanks again\n",
      "votes": 17
    },
    {
      "id": 1305038,
      "postDate": "2021-05-13T04:31:12.957Z",
      "content": "<p><a href=\"https://www.kaggle.com/yuvaramsingh\" target=\"_blank\">@yuvaramsingh</a> Congratulations on Solo Medal and thanks for sharing the approach </p>",
      "rawMarkdown": "@yuvaramsingh Congratulations on Solo Medal and thanks for sharing the approach ",
      "votes": 1
    },
    {
      "id": 1304149,
      "postDate": "2021-05-12T12:58:06.767Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1305038,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:31:12.957000",
      "content": "<p><a href=\"https://www.kaggle.com/yuvaramsingh\" target=\"_blank\">@yuvaramsingh</a> Congratulations on Solo Medal and thanks for sharing the approach </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304149,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-12T12:58:06.767000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1303436": "Thanks to kaggle and the organizers for putting together a competition on weakly supervised learning approach. learned a lot from this experience.\n\n## solution overview\nsince this is a weakly supervised learning task and HPA cell segment model was able to give instance mask, i chose to use Multiple instance learning method here. this made sense as the image level labels are not precise and through discussion it was clear that there could be cells that are negative class but not mentioned on the image level label. \n\n## stage 0 model\nHere i used Effb4 and resnest50 as the CNN backbone. the training setup consist of extracting individual cells from an image and selecting 16 cells (max, else resampling to match 16 count) at random to be used as Bag of images. the target is same as image level label. here i use 16 cells just to make sure i have at least some instance of cells corresponding to the label. (in stage 0 we cant be sure on cell labels) \n\nCNN Feature Extractor : Effb4, resnest 50\ncell used per image : 16\nPooling layer (per cell): GAP\nAttention pooling layer (Per image, 16 cells)\nimage level label\nFocal loss\nTime : roughly 1hr per epoch \n\nThe magic of this approach is the attention pooling layer which weights the 16 cell feature embedding by providing importance score based on the contribution of individual cells to the image level label. this enabled me to train a model that can find the appropriate cell level label by looking at the image level label. \n\n## stage 1 model\nafter training stage 0 model, now we have a model that is capable of providing cell level protein score that can be submitted to the competition. along with this we also have a well trained attention layer that can serve as a module to score individual cells. this can be used as a pseudo label generator. i used this approach to find the missing cell labels in a image and append this to the image level label. this helped with finding images which consist of Negative classes but not labeled correctly. This also helped with reducing the cell count per image as now we are more sure about the class of cells present on the image.\n\nCNN Feature Extractor : Effb4, resnest 50\ncell used per image : 8\nPooling layer (per cell): Transformer based attention pooling\nAttention pooling layer (Per image, 8 cells)\nimage level label + pseudo label\nFocal loss\nTime : roughly 20min per epoch \n\nwe can clearly see how the training time dropped because of reducing the no of cells.  \n\n## final inference \nMy final inference consist of  5 model where 3 stage 0 and 2 stage 1 models were used. i did use 4 TTA to make my model prediction robust\n\ni am happy that my public and private score for this submission did not change a lot Private 0.515 public 0.514. i did suffer from accuracy drop due to GPU hardware change. i trained my model on volta architecture GPU and Kaggle uses P100 (pascal). i got a scored of 0.524 on public dataset by generating the submission file from volta GPU. this was a surprise for me. i did try to calculate the mean of difference between the GPU submission files and used it to correct my final submission score. it did help me a bit but i lost a lot of score just because of Hardware difference. \n\n## final thoughts\noverall i got a good experience with working on Weakly supervised learning problem. Thanks again\n",
    "1305038": "@yuvaramsingh Congratulations on Solo Medal and thanks for sharing the approach ",
    "1304149": "Thanks for sharing!"
  }
}