{
  "id": 238120,
  "title": "44th place solution",
  "url": "/competitions/hubmap-kidney-segmentation/writeups/fabiendaniel-44th-place-solution",
  "author_name": "",
  "post_date": "2021-05-11T16:25:11.993Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I give here a brief overview of my solution mostly to share the methodology I used to tackle that competition: I used chunks of codes that track the evolution of the different experiments through git, which turns out to be quite useful in order to monitor the evolution along the various attempts. </p>\n<p>Since I had not previous experience on computer vision, I entered HubMap using the code shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (I take this opportunity to thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for all the material and knowledge he shares along its kaggle's journey).</p>\n<p>The code used for training  and inference is available on <a href=\"https://github.com/FabienDaniel/HubMap/\" target=\"_blank\">git</a>. <br>\nThe inference code on Kaggle is <a href=\"https://www.kaggle.com/fabiendaniel/submit-hubmap\" target=\"_blank\">here</a>. (utility scripts are used inside this script.) </p>\n<p>My solution consists in a two stage modeling:</p>\n<ul>\n<li>a first stage with 256 x 256 tiles, with a reduction by a factor 4 in resolution: this network is trained with emphasis on false negatives, using images that cover the entire .tiff images.</li>\n<li>the output of the previous network is used to locate the potential glomeruli on the image. A second network is then trained using images that are mostly located around known glomeruli. The train images have sizes of 515 x 512 with a scale factor of 2. </li>\n</ul>\n<p>For each of these two networks, I used Unet models with Efficientnet b0 backbones. I quickly checked other backbones but didn't see any noticeable impact on the results. It quickly appear that selecting the dataset to train the model was the key for model performance.</p>\n<h1>CV scheme</h1>\n<p>As mentioned many times, finding a correlation between local CV and LB was difficult because of the <code>d48</code> image. During the course of this competition, I tried various CV schemes:</p>\n<ol>\n<li>Kfold on whole tiles</li>\n<li>a random sampling of images with a random 20% of each tile used as validation</li>\n<li>a whole part of each tile kept as validation.</li>\n</ol>\n<p>A posteriori, the two latest schemes give good results with scores well aligned with the private LB. During the competition, it was not clear if any of these scheme would be reliable but I auto-convinced myself that the 3rd one should be reliable and I used it to select the details of the models.</p>\n<h1>Pseudo-labelisation</h1>\n<p>In the last days, I started to include the test tiles in order to train some models. For these images, the mask were generated using the best model I had so far. I subsequently checked \"visually\" that the glomeruli selected that way and the corresponding masks were reliable with respect to what I expected.</p>\n<p>[TO COMPLETE]</p>",
  "messages": [
    {
      "id": "1301875",
      "postDate": "05/11/2021 09:28:06",
      "content": "<p>I give here a brief overview of my solution mostly to share the methodology I used to tackle that competition: I used chunks of codes that track the evolution of the different experiments through git, which turns out to be quite useful in order to monitor the evolution along the various attempts. </p>\n<p>Since I had not previous experience on computer vision, I entered HubMap using the code shared by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (I take this opportunity to thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for all the material and knowledge he shares along its kaggle's journey).</p>\n<p>The code used for training  and inference is available on <a href=\"https://github.com/FabienDaniel/HubMap/\" target=\"_blank\">git</a>. <br>\nThe inference code on Kaggle is <a href=\"https://www.kaggle.com/fabiendaniel/submit-hubmap\" target=\"_blank\">here</a>. (utility scripts are used inside this script.) </p>\n<p>My solution consists in a two stage modeling:</p>\n<ul>\n<li>a first stage with 256 x 256 tiles, with a reduction by a factor 4 in resolution: this network is trained with emphasis on false negatives, using images that cover the entire .tiff images.</li>\n<li>the output of the previous network is used to locate the potential glomeruli on the image. A second network is then trained using images that are mostly located around known glomeruli. The train images have sizes of 515 x 512 with a scale factor of 2. </li>\n</ul>\n<p>For each of these two networks, I used Unet models with Efficientnet b0 backbones. I quickly checked other backbones but didn't see any noticeable impact on the results. It quickly appear that selecting the dataset to train the model was the key for model performance.</p>\n<h1>CV scheme</h1>\n<p>As mentioned many times, finding a correlation between local CV and LB was difficult because of the <code>d48</code> image. During the course of this competition, I tried various CV schemes:</p>\n<ol>\n<li>Kfold on whole tiles</li>\n<li>a random sampling of images with a random 20% of each tile used as validation</li>\n<li>a whole part of each tile kept as validation.</li>\n</ol>\n<p>A posteriori, the two latest schemes give good results with scores well aligned with the private LB. During the competition, it was not clear if any of these scheme would be reliable but I auto-convinced myself that the 3rd one should be reliable and I used it to select the details of the models.</p>\n<h1>Pseudo-labelisation</h1>\n<p>In the last days, I started to include the test tiles in order to train some models. For these images, the mask were generated using the best model I had so far. I subsequently checked \"visually\" that the glomeruli selected that way and the corresponding masks were reliable with respect to what I expected.</p>\n<p>[TO COMPLETE]</p>",
      "rawMarkdown": "I give here a brief overview of my solution mostly to share the methodology I used to tackle that competition: I used chunks of codes that track the evolution of the different experiments through git, which turns out to be quite useful in order to monitor the evolution along the various attempts. \n\nSince I had not previous experience on computer vision, I entered HubMap using the code shared by @hengck23 (I take this opportunity to thanks @hengck23 for all the material and knowledge he shares along its kaggle's journey).\n\nThe code used for training  and inference is available on [git](https://github.com/FabienDaniel/HubMap/). \nThe inference code on Kaggle is [here](https://www.kaggle.com/fabiendaniel/submit-hubmap). (utility scripts are used inside this script.) \n\n\n\nMy solution consists in a two stage modeling:\n- a first stage with 256 x 256 tiles, with a reduction by a factor 4 in resolution: this network is trained with emphasis on false negatives, using images that cover the entire .tiff images.\n- the output of the previous network is used to locate the potential glomeruli on the image. A second network is then trained using images that are mostly located around known glomeruli. The train images have sizes of 515 x 512 with a scale factor of 2. \n\nFor each of these two networks, I used Unet models with Efficientnet b0 backbones. I quickly checked other backbones but didn't see any noticeable impact on the results. It quickly appear that selecting the dataset to train the model was the key for model performance.\n\n\n# CV scheme\n\nAs mentioned many times, finding a correlation between local CV and LB was difficult because of the `d48` image. During the course of this competition, I tried various CV schemes:\n1. Kfold on whole tiles\n2. a random sampling of images with a random 20% of each tile used as validation\n3. a whole part of each tile kept as validation.\n\nA posteriori, the two latest schemes give good results with scores well aligned with the private LB. During the competition, it was not clear if any of these scheme would be reliable but I auto-convinced myself that the 3rd one should be reliable and I used it to select the details of the models.\n\n\n# Pseudo-labelisation\n\nIn the last days, I started to include the test tiles in order to train some models. For these images, the mask were generated using the best model I had so far. I subsequently checked \"visually\" that the glomeruli selected that way and the corresponding masks were reliable with respect to what I expected.\n\n[TO COMPLETE]",
      "votes": null
    },
    {
      "id": "1301915",
      "postDate": "05/11/2021 10:02:14",
      "content": "<p>big up for <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 👍</p>\n<p>I also followed his forum posts early in the competition and it helped me a lot (my best model, not chosen for submission is 0.948)</p>",
      "rawMarkdown": "big up for @hengck23 👍\n\nI also followed his forum posts early in the competition and it helped me a lot (my best model, not chosen for submission is 0.948)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1301915,
      "author_name": "rosuluc",
      "author_url": "",
      "post_date": "05/11/2021 10:02:14",
      "content": "<p>big up for <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 👍</p>\n<p>I also followed his forum posts early in the competition and it helped me a lot (my best model, not chosen for submission is 0.948)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1301875": "I give here a brief overview of my solution mostly to share the methodology I used to tackle that competition: I used chunks of codes that track the evolution of the different experiments through git, which turns out to be quite useful in order to monitor the evolution along the various attempts. \n\nSince I had not previous experience on computer vision, I entered HubMap using the code shared by @hengck23 (I take this opportunity to thanks @hengck23 for all the material and knowledge he shares along its kaggle's journey).\n\nThe code used for training  and inference is available on [git](https://github.com/FabienDaniel/HubMap/). \nThe inference code on Kaggle is [here](https://www.kaggle.com/fabiendaniel/submit-hubmap). (utility scripts are used inside this script.) \n\n\n\nMy solution consists in a two stage modeling:\n- a first stage with 256 x 256 tiles, with a reduction by a factor 4 in resolution: this network is trained with emphasis on false negatives, using images that cover the entire .tiff images.\n- the output of the previous network is used to locate the potential glomeruli on the image. A second network is then trained using images that are mostly located around known glomeruli. The train images have sizes of 515 x 512 with a scale factor of 2. \n\nFor each of these two networks, I used Unet models with Efficientnet b0 backbones. I quickly checked other backbones but didn't see any noticeable impact on the results. It quickly appear that selecting the dataset to train the model was the key for model performance.\n\n\n# CV scheme\n\nAs mentioned many times, finding a correlation between local CV and LB was difficult because of the `d48` image. During the course of this competition, I tried various CV schemes:\n1. Kfold on whole tiles\n2. a random sampling of images with a random 20% of each tile used as validation\n3. a whole part of each tile kept as validation.\n\nA posteriori, the two latest schemes give good results with scores well aligned with the private LB. During the competition, it was not clear if any of these scheme would be reliable but I auto-convinced myself that the 3rd one should be reliable and I used it to select the details of the models.\n\n\n# Pseudo-labelisation\n\nIn the last days, I started to include the test tiles in order to train some models. For these images, the mask were generated using the best model I had so far. I subsequently checked \"visually\" that the glomeruli selected that way and the corresponding masks were reliable with respect to what I expected.\n\n[TO COMPLETE]",
    "1301915": "big up for @hengck23 👍\n\nI also followed his forum posts early in the competition and it helped me a lot (my best model, not chosen for submission is 0.948)"
  },
  "source": "meta"
}