{
  "id": 239861,
  "title": "HPA 36th Place Solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/239861",
  "author_name": "tacorice",
  "post_date": "2021-05-17T21:40:38.486000",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all we would like to thank the host for this really interesting competition, and congrats to all the winners.</p>\n<h1>Overview</h1>\n<p>These are the issues and solutions we had.</p>\n<p><strong>Issue</strong><br>\n① Learning with weak supervised learning<br>\n② The nucleus protruding from the image<br>\n③ The problem of imbalance in the number of labels</p>\n<p><strong>Solution</strong><br>\nFor ③, the problem of imbalance in the number of labels could be reinforced by using external data (downsampling was performed so that each label would have about 10,000 labels). This resulted in an LB score of +0.01. In the past competitions, when the number of labels was unbalanced, using focal loss for loss was a good example of learning, but this time it didn't.</p>\n<p>As a measure against weak supervised learning in ①, we used the difference in pixel values. If there were cells with different labels in the image,we assume that the staining intensities are different and that when cropped into a single cell, the average pixel values ​​in the image will vary. Therefore, we contributed to improving the accuracy of the model by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images. In addition, by setting a threshold value, we were able to eliminate a certain number of problems ② in which the nuclei protrude from the image. (Unfortunately, we couldn't compare the thresholds because we didn't have enough time.) As a result, the LB score was +0.03~0.04.</p>\n<h1>Models</h1>\n<p>The models are divided into two types, a single cell model and an image level model.</p>\n<p><strong>Single cell model</strong><br>\nModel：ResNet50+EfficientNet B4<br>\nImage size：128x128<br>\nLoss：BCEWithLogitsLoss<br>\nAugmentation：<br>\nFlip<br>\nTTA (n=3)<br>\nDataset:<br>\nTrain：179.2k<br>\nValidation：44.7k</p>\n<p><strong>Image level model</strong><br>\nModel：SEResNeXt50 32×4d+EfficientNet B7<br>\nImage size：640x640<br>\nLoss：<br>\nBCEWithLogitsLoss (EfficientNet B7)<br>\nFocalLoss (SEResNeXt50 32×4d)<br>\nAugmentation：HorizontalFlip(p=0.5)<br>\nDataset:Use only green channel<br>\nTrain: 17.4k<br>\nValidation: 4.4k</p>\n<p>That's how I was able to win the silver medal.<br>\nAnd with this medal, I was able to be promoted to Kaggle Master! !<br>\nVery glad！</p>\n<p>A year before I joined kaggle, I didn't understand Python and machine learning at all, but I feel that by aiming for medals, I've gradually become able to do what I couldn't do. In addition, I think that I was able to train a lot mentally by experiencing a lot of shake-ups and shakedowns XD</p>\n<p>Kaggle is the best data science learning platform for me.<br>\nThanks to kaggle and all kagglers.</p>\n<p>I will continue to take on challenges</p>",
  "messages": [
    {
      "id": 1312163,
      "postDate": "2021-05-17T21:40:38.487Z",
      "content": "<p>First of all we would like to thank the host for this really interesting competition, and congrats to all the winners.</p>\n<h1>Overview</h1>\n<p>These are the issues and solutions we had.</p>\n<p><strong>Issue</strong><br>\n① Learning with weak supervised learning<br>\n② The nucleus protruding from the image<br>\n③ The problem of imbalance in the number of labels</p>\n<p><strong>Solution</strong><br>\nFor ③, the problem of imbalance in the number of labels could be reinforced by using external data (downsampling was performed so that each label would have about 10,000 labels). This resulted in an LB score of +0.01. In the past competitions, when the number of labels was unbalanced, using focal loss for loss was a good example of learning, but this time it didn't.</p>\n<p>As a measure against weak supervised learning in ①, we used the difference in pixel values. If there were cells with different labels in the image,we assume that the staining intensities are different and that when cropped into a single cell, the average pixel values ​​in the image will vary. Therefore, we contributed to improving the accuracy of the model by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images. In addition, by setting a threshold value, we were able to eliminate a certain number of problems ② in which the nuclei protrude from the image. (Unfortunately, we couldn't compare the thresholds because we didn't have enough time.) As a result, the LB score was +0.03~0.04.</p>\n<h1>Models</h1>\n<p>The models are divided into two types, a single cell model and an image level model.</p>\n<p><strong>Single cell model</strong><br>\nModel：ResNet50+EfficientNet B4<br>\nImage size：128x128<br>\nLoss：BCEWithLogitsLoss<br>\nAugmentation：<br>\nFlip<br>\nTTA (n=3)<br>\nDataset:<br>\nTrain：179.2k<br>\nValidation：44.7k</p>\n<p><strong>Image level model</strong><br>\nModel：SEResNeXt50 32×4d+EfficientNet B7<br>\nImage size：640x640<br>\nLoss：<br>\nBCEWithLogitsLoss (EfficientNet B7)<br>\nFocalLoss (SEResNeXt50 32×4d)<br>\nAugmentation：HorizontalFlip(p=0.5)<br>\nDataset:Use only green channel<br>\nTrain: 17.4k<br>\nValidation: 4.4k</p>\n<p>That's how I was able to win the silver medal.<br>\nAnd with this medal, I was able to be promoted to Kaggle Master! !<br>\nVery glad！</p>\n<p>A year before I joined kaggle, I didn't understand Python and machine learning at all, but I feel that by aiming for medals, I've gradually become able to do what I couldn't do. In addition, I think that I was able to train a lot mentally by experiencing a lot of shake-ups and shakedowns XD</p>\n<p>Kaggle is the best data science learning platform for me.<br>\nThanks to kaggle and all kagglers.</p>\n<p>I will continue to take on challenges</p>",
      "rawMarkdown": "First of all we would like to thank the host for this really interesting competition, and congrats to all the winners.\n\n\n# Overview\n\nThese are the issues and solutions we had.\n\n**Issue**\n① Learning with weak supervised learning\n② The nucleus protruding from the image\n③ The problem of imbalance in the number of labels\n\n**Solution**\nFor ③, the problem of imbalance in the number of labels could be reinforced by using external data (downsampling was performed so that each label would have about 10,000 labels). This resulted in an LB score of +0.01. In the past competitions, when the number of labels was unbalanced, using focal loss for loss was a good example of learning, but this time it didn't.\n\nAs a measure against weak supervised learning in ①, we used the difference in pixel values. If there were cells with different labels in the image,we assume that the staining intensities are different and that when cropped into a single cell, the average pixel values ​​in the image will vary. Therefore, we contributed to improving the accuracy of the model by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images. In addition, by setting a threshold value, we were able to eliminate a certain number of problems ② in which the nuclei protrude from the image. (Unfortunately, we couldn't compare the thresholds because we didn't have enough time.) As a result, the LB score was +0.03~0.04.\n\n\n# Models\n\nThe models are divided into two types, a single cell model and an image level model.\n\n**Single cell model**\nModel：ResNet50+EfficientNet B4\nImage size：128x128\nLoss：BCEWithLogitsLoss\nAugmentation：\nFlip\nTTA (n=3)\nDataset:\nTrain：179.2k\nValidation：44.7k\n\n**Image level model**\nModel：SEResNeXt50 32×4d+EfficientNet B7\nImage size：640x640\nLoss：\nBCEWithLogitsLoss (EfficientNet B7)\nFocalLoss (SEResNeXt50 32×4d)\nAugmentation：HorizontalFlip(p=0.5)\nDataset:Use only green channel\nTrain: 17.4k\nValidation: 4.4k\n\nThat's how I was able to win the silver medal.\nAnd with this medal, I was able to be promoted to Kaggle Master! !\nVery glad！\n\nA year before I joined kaggle, I didn't understand Python and machine learning at all, but I feel that by aiming for medals, I've gradually become able to do what I couldn't do. In addition, I think that I was able to train a lot mentally by experiencing a lot of shake-ups and shakedowns XD\n\nKaggle is the best data science learning platform for me.\nThanks to kaggle and all kagglers.\n\nI will continue to take on challenges",
      "votes": 10
    },
    {
      "id": 1312555,
      "postDate": "2021-05-18T05:40:14.810Z",
      "content": "<p>Congratulations for becoming a competitions master. Thank you for the write-up. Could you elaborate a bit more on:</p>\n<blockquote>\n  <p>by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images.</p>\n</blockquote>\n<p>I dont understand yet.</p>",
      "rawMarkdown": "Congratulations for becoming a competitions master. Thank you for the write-up. Could you elaborate a bit more on:\n\n> by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images.\n\nI dont understand yet.",
      "replies": [
        {
          "id": 1312792,
          "postDate": "2021-05-18T08:51:35.467Z",
          "content": "<p>I appreciate your celebration message.</p>\n<p>I would like to explain the content you pointed out below.</p>\n<p><em>U</em>: Mean value of all single cell images<br>\n<em>μ_i</em>: Mean value of <em>i</em>-th single cell image<br>\n<em>thr</em>: Coefficient for threshold (0.2 here)<br>\n<em>thr</em>× <em>U</em>: Threshold</p>\n<p>If <em>μ_i</em> &lt;<em>thr</em> ×<em>U</em>, the <em>i</em>-th single cell image is removed from the dataset.</p>\n<p>However, I have not been able to verify whether <em>thr</em> = 0.2 is better, <em>thr</em> = 0.3 is better, or something else is.</p>",
          "rawMarkdown": "I appreciate your celebration message.\n\nI would like to explain the content you pointed out below.\n\n*U*: Mean value of all single cell images\n*μ_i*: Mean value of *i*-th single cell image\n*thr*: Coefficient for threshold (0.2 here)\n*thr*× *U*: Threshold\n\nIf *μ_i* <*thr* ×*U*, the *i*-th single cell image is removed from the dataset.\n\nHowever, I have not been able to verify whether *thr* = 0.2 is better, *thr* = 0.3 is better, or something else is.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1312555,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2021-05-18T05:40:14.810000",
      "content": "<p>Congratulations for becoming a competitions master. Thank you for the write-up. Could you elaborate a bit more on:</p>\n<blockquote>\n  <p>by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images.</p>\n</blockquote>\n<p>I dont understand yet.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1312792,
          "author_name": "tacorice",
          "author_url": "",
          "post_date": "2021-05-18T08:51:35.467000",
          "content": "<p>I appreciate your celebration message.</p>\n<p>I would like to explain the content you pointed out below.</p>\n<p><em>U</em>: Mean value of all single cell images<br>\n<em>μ_i</em>: Mean value of <em>i</em>-th single cell image<br>\n<em>thr</em>: Coefficient for threshold (0.2 here)<br>\n<em>thr</em>× <em>U</em>: Threshold</p>\n<p>If <em>μ_i</em> &lt;<em>thr</em> ×<em>U</em>, the <em>i</em>-th single cell image is removed from the dataset.</p>\n<p>However, I have not been able to verify whether <em>thr</em> = 0.2 is better, <em>thr</em> = 0.3 is better, or something else is.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1312163": "First of all we would like to thank the host for this really interesting competition, and congrats to all the winners.\n\n\n# Overview\n\nThese are the issues and solutions we had.\n\n**Issue**\n① Learning with weak supervised learning\n② The nucleus protruding from the image\n③ The problem of imbalance in the number of labels\n\n**Solution**\nFor ③, the problem of imbalance in the number of labels could be reinforced by using external data (downsampling was performed so that each label would have about 10,000 labels). This resulted in an LB score of +0.01. In the past competitions, when the number of labels was unbalanced, using focal loss for loss was a good example of learning, but this time it didn't.\n\nAs a measure against weak supervised learning in ①, we used the difference in pixel values. If there were cells with different labels in the image,we assume that the staining intensities are different and that when cropped into a single cell, the average pixel values ​​in the image will vary. Therefore, we contributed to improving the accuracy of the model by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images. In addition, by setting a threshold value, we were able to eliminate a certain number of problems ② in which the nuclei protrude from the image. (Unfortunately, we couldn't compare the thresholds because we didn't have enough time.) As a result, the LB score was +0.03~0.04.\n\n\n# Models\n\nThe models are divided into two types, a single cell model and an image level model.\n\n**Single cell model**\nModel：ResNet50+EfficientNet B4\nImage size：128x128\nLoss：BCEWithLogitsLoss\nAugmentation：\nFlip\nTTA (n=3)\nDataset:\nTrain：179.2k\nValidation：44.7k\n\n**Image level model**\nModel：SEResNeXt50 32×4d+EfficientNet B7\nImage size：640x640\nLoss：\nBCEWithLogitsLoss (EfficientNet B7)\nFocalLoss (SEResNeXt50 32×4d)\nAugmentation：HorizontalFlip(p=0.5)\nDataset:Use only green channel\nTrain: 17.4k\nValidation: 4.4k\n\nThat's how I was able to win the silver medal.\nAnd with this medal, I was able to be promoted to Kaggle Master! !\nVery glad！\n\nA year before I joined kaggle, I didn't understand Python and machine learning at all, but I feel that by aiming for medals, I've gradually become able to do what I couldn't do. In addition, I think that I was able to train a lot mentally by experiencing a lot of shake-ups and shakedowns XD\n\nKaggle is the best data science learning platform for me.\nThanks to kaggle and all kagglers.\n\nI will continue to take on challenges",
    "1312555": "Congratulations for becoming a competitions master. Thank you for the write-up. Could you elaborate a bit more on:\n\n> by using a data set in which 20% of the average value is removed as the threshold value for the average value of the images.\n\nI dont understand yet."
  }
}