{
  "id": 238678,
  "title": "9th Place Solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238678",
  "author_name": "Qishen Ha",
  "post_date": "2021-05-13T03:44:05.307000",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>We had a tough competition, congratulations to all the kagglers who persevered to the end and many thanks to the organizers and my wonderful teammates <a href=\"https://www.kaggle.com/daishu\" target=\"_blank\">@daishu</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a>  <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> </p>\n<h1>Summary</h1>\n<p>We have two kinds of pipeline in this competition, they are:</p>\n<ul>\n<li>Pipeline1: Train on full image, test on single cells (mask out other cells)</li>\n<li>Pipeline2: Train on cropped cells, test on cropped cells also</li>\n</ul>\n<p>Finally, we trained 20~ folds of model with these two pipelines and ensembled them by taking the mean value.</p>\n<p><strong><em>Notice that we do not use image-level prediction.</em></strong></p>\n<h1>Methods</h1>\n<h3>Pipeline1</h3>\n<p>We used two methods to train-test this pipeline.</p>\n<p>The first one is to train with 512 images, and the test input is also 512. We loop n times for each image (n is the number of cells in the image), leaving only one cell in each time and masking out the other cells to get single cell predictions.</p>\n<p>The second one is trained with 768 random crop 512, and then tested almost the same way as the first one, but not only mask out the other cells, but we also put the position of the cells left in the center of the image.</p>\n<h3>Pipeline2</h3>\n<p>We pre-crop all the cells of each image and save them locally. Then during training, for each image we randomly select 16 cells. We then set bs=32, so for each batch we have 32x16=512 cells in total.</p>\n<p>We resize each cell to 128x128, so the returned data shape from the dataloader is <code>(32, 16, 4, 128, 128)</code> . Next we reshape it into <code>(512, 4, 128, 128)</code> and then use a very common CNN to forward it, the output shape is <code>(512, 19)</code></p>\n<p>In the prediction phase, we will directly take this output and use it as the predicted value for each cell. ↑</p>\n<p>But during the training process, we rereshape this <code>(512, 19)</code> prediction back into <code>(32, 16, 19)</code> . Then the loss is calculated for each cell with image-level GT label.</p>\n<h1>Acknowledge</h1>\n<p>Special Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU!</p>\n<p>Since I got the Z8G4 workstation and the ZBook last December, this is the third gold medal I've won from Kaggle ;)</p>",
  "messages": [
    {
      "id": 1304985,
      "postDate": "2021-05-13T03:44:05.307Z",
      "content": "<p>Hi,</p>\n<p>We had a tough competition, congratulations to all the kagglers who persevered to the end and many thanks to the organizers and my wonderful teammates <a href=\"https://www.kaggle.com/daishu\" target=\"_blank\">@daishu</a> <a href=\"https://www.kaggle.com/garybios\" target=\"_blank\">@garybios</a>  <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> </p>\n<h1>Summary</h1>\n<p>We have two kinds of pipeline in this competition, they are:</p>\n<ul>\n<li>Pipeline1: Train on full image, test on single cells (mask out other cells)</li>\n<li>Pipeline2: Train on cropped cells, test on cropped cells also</li>\n</ul>\n<p>Finally, we trained 20~ folds of model with these two pipelines and ensembled them by taking the mean value.</p>\n<p><strong><em>Notice that we do not use image-level prediction.</em></strong></p>\n<h1>Methods</h1>\n<h3>Pipeline1</h3>\n<p>We used two methods to train-test this pipeline.</p>\n<p>The first one is to train with 512 images, and the test input is also 512. We loop n times for each image (n is the number of cells in the image), leaving only one cell in each time and masking out the other cells to get single cell predictions.</p>\n<p>The second one is trained with 768 random crop 512, and then tested almost the same way as the first one, but not only mask out the other cells, but we also put the position of the cells left in the center of the image.</p>\n<h3>Pipeline2</h3>\n<p>We pre-crop all the cells of each image and save them locally. Then during training, for each image we randomly select 16 cells. We then set bs=32, so for each batch we have 32x16=512 cells in total.</p>\n<p>We resize each cell to 128x128, so the returned data shape from the dataloader is <code>(32, 16, 4, 128, 128)</code> . Next we reshape it into <code>(512, 4, 128, 128)</code> and then use a very common CNN to forward it, the output shape is <code>(512, 19)</code></p>\n<p>In the prediction phase, we will directly take this output and use it as the predicted value for each cell. ↑</p>\n<p>But during the training process, we rereshape this <code>(512, 19)</code> prediction back into <code>(32, 16, 19)</code> . Then the loss is calculated for each cell with image-level GT label.</p>\n<h1>Acknowledge</h1>\n<p>Special Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU!</p>\n<p>Since I got the Z8G4 workstation and the ZBook last December, this is the third gold medal I've won from Kaggle ;)</p>",
      "rawMarkdown": "Hi,\n\nWe had a tough competition, congratulations to all the kagglers who persevered to the end and many thanks to the organizers and my wonderful teammates @daishu @garybios  @boliu0 \n\n# Summary\n\nWe have two kinds of pipeline in this competition, they are:\n\n* Pipeline1: Train on full image, test on single cells (mask out other cells)\n* Pipeline2: Train on cropped cells, test on cropped cells also\n\nFinally, we trained 20~ folds of model with these two pipelines and ensembled them by taking the mean value.\n\n***Notice that we do not use image-level prediction.***\n\n# Methods\n\n### Pipeline1\n\nWe used two methods to train-test this pipeline.\n\nThe first one is to train with 512 images, and the test input is also 512. We loop n times for each image (n is the number of cells in the image), leaving only one cell in each time and masking out the other cells to get single cell predictions.\n\nThe second one is trained with 768 random crop 512, and then tested almost the same way as the first one, but not only mask out the other cells, but we also put the position of the cells left in the center of the image.\n\n### Pipeline2\n\nWe pre-crop all the cells of each image and save them locally. Then during training, for each image we randomly select 16 cells. We then set bs=32, so for each batch we have 32x16=512 cells in total.\n\nWe resize each cell to 128x128, so the returned data shape from the dataloader is `(32, 16, 4, 128, 128)` . Next we reshape it into `(512, 4, 128, 128)` and then use a very common CNN to forward it, the output shape is `(512, 19)`\n\nIn the prediction phase, we will directly take this output and use it as the predicted value for each cell. ↑\n\nBut during the training process, we rereshape this `(512, 19)` prediction back into `(32, 16, 19)` . Then the loss is calculated for each cell with image-level GT label.\n\n\n# Acknowledge\n\nSpecial Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU!\n\nSince I got the Z8G4 workstation and the ZBook last December, this is the third gold medal I've won from Kaggle ;)",
      "votes": 24
    },
    {
      "id": 1304995,
      "postDate": "2021-05-13T03:55:04.240Z",
      "content": "<p>Congratulations all the winners and thanks my wonderful teammates. <br>\nThis is one of the most difficult competition I have ever participated in Kaggle.<br>\nAfter reading other top solutions, we realized that the solutions were very similar, but we might have missed many details, which led to our failure to optimize them well.</p>\n<p>And Special Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU. I've been doing all experiments with Z8Workstation in this competition.</p>",
      "rawMarkdown": "Congratulations all the winners and thanks my wonderful teammates. \nThis is one of the most difficult competition I have ever participated in Kaggle.\nAfter reading other top solutions, we realized that the solutions were very similar, but we might have missed many details, which led to our failure to optimize them well.\n\nAnd Special Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU. I've been doing all experiments with Z8Workstation in this competition.\n",
      "votes": 6
    },
    {
      "id": 1311349,
      "postDate": "2021-05-17T10:44:55.880Z",
      "content": "<p>Congratulations and Thanks for sharing the approach</p>",
      "rawMarkdown": "Congratulations and Thanks for sharing the approach",
      "votes": 1
    },
    {
      "id": 1308078,
      "postDate": "2021-05-15T00:46:51.197Z",
      "content": "<p>Congratulations Qishen and team!</p>",
      "rawMarkdown": "Congratulations Qishen and team!",
      "votes": 1
    },
    {
      "id": 1305073,
      "postDate": "2021-05-13T04:52:14.913Z",
      "content": "<p>Congrats on 9th place <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and team! </p>",
      "rawMarkdown": "Congrats on 9th place @haqishen and team! ",
      "votes": 1
    },
    {
      "id": 1305020,
      "postDate": "2021-05-13T04:24:23.300Z",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Congratulations on Gold and Thanks for sharing the approach</p>",
      "rawMarkdown": "@haqishen Congratulations on Gold and Thanks for sharing the approach",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1304995,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2021-05-13T03:55:04.240000",
      "content": "<p>Congratulations all the winners and thanks my wonderful teammates. <br>\nThis is one of the most difficult competition I have ever participated in Kaggle.<br>\nAfter reading other top solutions, we realized that the solutions were very similar, but we might have missed many details, which led to our failure to optimize them well.</p>\n<p>And Special Thanks to Z by HP &amp; NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU. I've been doing all experiments with Z8Workstation in this competition.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1311349,
      "author_name": "Sani Kamal",
      "author_url": "",
      "post_date": "2021-05-17T10:44:55.880000",
      "content": "<p>Congratulations and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1308078,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-05-15T00:46:51.197000",
      "content": "<p>Congratulations Qishen and team!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1305073,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-05-13T04:52:14.913000",
      "content": "<p>Congrats on 9th place <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> and team! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1305020,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:24:23.300000",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Congratulations on Gold and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1304985": "Hi,\n\nWe had a tough competition, congratulations to all the kagglers who persevered to the end and many thanks to the organizers and my wonderful teammates @daishu @garybios  @boliu0 \n\n# Summary\n\nWe have two kinds of pipeline in this competition, they are:\n\n* Pipeline1: Train on full image, test on single cells (mask out other cells)\n* Pipeline2: Train on cropped cells, test on cropped cells also\n\nFinally, we trained 20~ folds of model with these two pipelines and ensembled them by taking the mean value.\n\n***Notice that we do not use image-level prediction.***\n\n# Methods\n\n### Pipeline1\n\nWe used two methods to train-test this pipeline.\n\nThe first one is to train with 512 images, and the test input is also 512. We loop n times for each image (n is the number of cells in the image), leaving only one cell in each time and masking out the other cells to get single cell predictions.\n\nThe second one is trained with 768 random crop 512, and then tested almost the same way as the first one, but not only mask out the other cells, but we also put the position of the cells left in the center of the image.\n\n### Pipeline2\n\nWe pre-crop all the cells of each image and save them locally. Then during training, for each image we randomly select 16 cells. We then set bs=32, so for each batch we have 32x16=512 cells in total.\n\nWe resize each cell to 128x128, so the returned data shape from the dataloader is `(32, 16, 4, 128, 128)` . Next we reshape it into `(512, 4, 128, 128)` and then use a very common CNN to forward it, the output shape is `(512, 19)`\n\nIn the prediction phase, we will directly take this output and use it as the predicted value for each cell. ↑\n\nBut during the training process, we rereshape this `(512, 19)` prediction back into `(32, 16, 19)` . Then the loss is calculated for each cell with image-level GT label.\n\n\n# Acknowledge\n\nSpecial Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU!\n\nSince I got the Z8G4 workstation and the ZBook last December, this is the third gold medal I've won from Kaggle ;)",
    "1304995": "Congratulations all the winners and thanks my wonderful teammates. \nThis is one of the most difficult competition I have ever participated in Kaggle.\nAfter reading other top solutions, we realized that the solutions were very similar, but we might have missed many details, which led to our failure to optimize them well.\n\nAnd Special Thanks to Z by HP & NVIDIA for sponsoring me a Z8G4 Workstation with dual RTX6000 GPU and a ZBook with RTX5000 GPU. I've been doing all experiments with Z8Workstation in this competition.\n",
    "1311349": "Congratulations and Thanks for sharing the approach",
    "1308078": "Congratulations Qishen and team!",
    "1305073": "Congrats on 9th place @haqishen and team! ",
    "1305020": "@haqishen Congratulations on Gold and Thanks for sharing the approach"
  }
}