{
  "id": 343291,
  "title": "Too much cv vs lb gap..(CV 78 vs LB 60), also this will become my log book of this competition",
  "url": "/competitions/hubmap-organ-segmentation/discussion/343291",
  "author_name": "kaggler",
  "post_date": "2022-08-10T17:47:28.545000",
  "votes": 18,
  "comment_count": 38,
  "views": 0,
  "content": "<p>My Efficientnet7 + DeeplabV3Plus can score around 0.7837 with <strong>tiling</strong>, 100epoch in Local CV(single fold) but when i submit it, I just get only 0.60<br>\n2022/8/12 <br>\none fold tiling cv score : <strong>0.7837</strong><br>\n<strong>(HUBMAP LB 39 , HPA LB 21)</strong>  <br>\n<strong>spleen val dice score : 0.7772</strong><br>\n<strong>lung val dice score : 0.2246</strong><br>\n<strong>kidney val dice score : 0.9474</strong><br>\n<strong>largeintestine val dice score : 0.9093</strong>  <br>\n<strong>prostate val dice score : 0.8610</strong></p>\n<p>2022/8/14  <br>\nI had investigated the Hubmap scores of each organ<br>\n<strong>spleen hubmap score  : 0.11</strong>  <br>\n<strong>lung hubmap score : 0.03</strong>  <br>\n<strong>large-intestine score : 0.05</strong>  <br>\n<strong>kidney hubmap score : 0.11</strong>  <br>\n<strong>prostate hubmap score : 0.06</strong>  </p>\n<p>Local Image prediction  <br>\n<img src=\"https://i.imgur.com/DDdnYwT.png\" alt=\"\"><br>\nPrediction on the test image<br>\n<img src=\"https://i.imgur.com/7EPLSG8.png\" alt=\"\"></p>\n<p>lung and prostate scores seem to bad compared to other competitors<br>\nlung and prostate hugely failed to be generalized to HUBMAP</p>\n<p>I have been working day and night for this comp.. but i am very depressed by the fact that i can't cross even 0.6</p>\n<p>2022/8/15  <br>\nI just added more augmentations than before, no external data<br>\n<strong>(Hubmap LB 51 &lt;- super increased , HPA LB 21 &lt;- cv Up, but not reflected in LB)</strong><br>\nas <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> said in his topic, not much domain miss match between Hubmap and HPA?   <br>\ntommorrow(2022/8/16, i will investigate scores of each organ)<br>\none fold tiling cv score : <strong>0.7885</strong><br>\n<strong>spleen val dice score : 0.8092</strong><br>\n<strong>lung val dice score : 0.2186</strong><br>\n<strong>kidney val dice score : 0.9471</strong><br>\n<strong>largeintestine val dice score : 0.9073</strong>  <br>\n<strong>prostate val dice score : 0.8677</strong></p>\n<p>Local Image prediction  <br>\n<img src=\"https://i.imgur.com/sszC68I.png\" alt=\"\"><br>\nPrediction on the test image <br>\n<img src=\"https://i.imgur.com/9UnkJJF.png\" alt=\"\"></p>\n<p>2022/8/16<br>\nMore Epochs improved the cv score but not LB<br>\n<strong>(Hubmap 51  , HPA 21, LB73 )</strong><br>\none fold tiling cv score : <strong>0.7939</strong><br>\n<strong>spleen val dice score : 0.8249</strong><br>\n<strong>lung val dice score : 0.2455</strong><br>\n<strong>kidney val dice score : 0.9549</strong><br>\n<strong>largeintestine val dice score : 0.9048</strong>  <br>\n<strong>prostate val dice score : 0.8541</strong>  <br>\nwhy on earth prostate local score decreased?  </p>\n<p><strong>Hubmap prostate LB : 0.13</strong>  <br>\n<strong>Spleen LB : 0.12</strong>  <br>\n<strong>Lung LB : 0.07</strong>  <br>\n<strong>Kidney LB : 0.11</strong>  <br>\n<strong>LargeIntestine LB : 0.05</strong>  <br>\nMy improvements are from spleen and lung, prostate!<br>\nIt seems like the spleen can't be generalized well with the method of tiling since we need to see a large area</p>\n<p>2022/8/18  <br>\nLocal Image Prediction  <br>\n<img src=\"https://i.imgur.com/pDiq3dQ.png\" alt=\"\">  <br>\nI have started to train the whole images.  CV dropped from 0.7885 -&gt; 0.7635 but the inference became more realistic.  </p>\n<p>2022/8/19  <br>\nimproved with some tuning .. <strong>CV 80, LB 75, HUBMAP LB 53</strong>  <br>\nOnly using provided competition dataset, same model, my model has improved. awesome!</p>",
  "messages": [
    {
      "id": 1893308,
      "postDate": "2022-08-10T17:47:28.547Z",
      "content": "<p>My Efficientnet7 + DeeplabV3Plus can score around 0.7837 with <strong>tiling</strong>, 100epoch in Local CV(single fold) but when i submit it, I just get only 0.60<br>\n2022/8/12 <br>\none fold tiling cv score : <strong>0.7837</strong><br>\n<strong>(HUBMAP LB 39 , HPA LB 21)</strong>  <br>\n<strong>spleen val dice score : 0.7772</strong><br>\n<strong>lung val dice score : 0.2246</strong><br>\n<strong>kidney val dice score : 0.9474</strong><br>\n<strong>largeintestine val dice score : 0.9093</strong>  <br>\n<strong>prostate val dice score : 0.8610</strong></p>\n<p>2022/8/14  <br>\nI had investigated the Hubmap scores of each organ<br>\n<strong>spleen hubmap score  : 0.11</strong>  <br>\n<strong>lung hubmap score : 0.03</strong>  <br>\n<strong>large-intestine score : 0.05</strong>  <br>\n<strong>kidney hubmap score : 0.11</strong>  <br>\n<strong>prostate hubmap score : 0.06</strong>  </p>\n<p>Local Image prediction  <br>\n<img src=\"https://i.imgur.com/DDdnYwT.png\" alt=\"\"><br>\nPrediction on the test image<br>\n<img src=\"https://i.imgur.com/7EPLSG8.png\" alt=\"\"></p>\n<p>lung and prostate scores seem to bad compared to other competitors<br>\nlung and prostate hugely failed to be generalized to HUBMAP</p>\n<p>I have been working day and night for this comp.. but i am very depressed by the fact that i can't cross even 0.6</p>\n<p>2022/8/15  <br>\nI just added more augmentations than before, no external data<br>\n<strong>(Hubmap LB 51 &lt;- super increased , HPA LB 21 &lt;- cv Up, but not reflected in LB)</strong><br>\nas <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> said in his topic, not much domain miss match between Hubmap and HPA?   <br>\ntommorrow(2022/8/16, i will investigate scores of each organ)<br>\none fold tiling cv score : <strong>0.7885</strong><br>\n<strong>spleen val dice score : 0.8092</strong><br>\n<strong>lung val dice score : 0.2186</strong><br>\n<strong>kidney val dice score : 0.9471</strong><br>\n<strong>largeintestine val dice score : 0.9073</strong>  <br>\n<strong>prostate val dice score : 0.8677</strong></p>\n<p>Local Image prediction  <br>\n<img src=\"https://i.imgur.com/sszC68I.png\" alt=\"\"><br>\nPrediction on the test image <br>\n<img src=\"https://i.imgur.com/9UnkJJF.png\" alt=\"\"></p>\n<p>2022/8/16<br>\nMore Epochs improved the cv score but not LB<br>\n<strong>(Hubmap 51  , HPA 21, LB73 )</strong><br>\none fold tiling cv score : <strong>0.7939</strong><br>\n<strong>spleen val dice score : 0.8249</strong><br>\n<strong>lung val dice score : 0.2455</strong><br>\n<strong>kidney val dice score : 0.9549</strong><br>\n<strong>largeintestine val dice score : 0.9048</strong>  <br>\n<strong>prostate val dice score : 0.8541</strong>  <br>\nwhy on earth prostate local score decreased?  </p>\n<p><strong>Hubmap prostate LB : 0.13</strong>  <br>\n<strong>Spleen LB : 0.12</strong>  <br>\n<strong>Lung LB : 0.07</strong>  <br>\n<strong>Kidney LB : 0.11</strong>  <br>\n<strong>LargeIntestine LB : 0.05</strong>  <br>\nMy improvements are from spleen and lung, prostate!<br>\nIt seems like the spleen can't be generalized well with the method of tiling since we need to see a large area</p>\n<p>2022/8/18  <br>\nLocal Image Prediction  <br>\n<img src=\"https://i.imgur.com/pDiq3dQ.png\" alt=\"\">  <br>\nI have started to train the whole images.  CV dropped from 0.7885 -&gt; 0.7635 but the inference became more realistic.  </p>\n<p>2022/8/19  <br>\nimproved with some tuning .. <strong>CV 80, LB 75, HUBMAP LB 53</strong>  <br>\nOnly using provided competition dataset, same model, my model has improved. awesome!</p>",
      "rawMarkdown": "My Efficientnet7 + DeeplabV3Plus can score around 0.7837 with **tiling**, 100epoch in Local CV(single fold) but when i submit it, I just get only 0.60\n2022/8/12 \none fold tiling cv score : **0.7837**\n**(HUBMAP LB 39 , HPA LB 21)**  \n**spleen val dice score : 0.7772**\n**lung val dice score : 0.2246**\n**kidney val dice score : 0.9474**\n**largeintestine val dice score : 0.9093**  \n**prostate val dice score : 0.8610**\n\n2022/8/14  \nI had investigated the Hubmap scores of each organ\n**spleen hubmap score  : 0.11**  \n**lung hubmap score : 0.03**  \n**large-intestine score : 0.05**  \n**kidney hubmap score : 0.11**  \n**prostate hubmap score : 0.06**  \n\nLocal Image prediction  \n![](https://i.imgur.com/DDdnYwT.png)\nPrediction on the test image\n![](https://i.imgur.com/7EPLSG8.png)\n\nlung and prostate scores seem to bad compared to other competitors\nlung and prostate hugely failed to be generalized to HUBMAP\n\nI have been working day and night for this comp.. but i am very depressed by the fact that i can't cross even 0.6\n\n2022/8/15  \nI just added more augmentations than before, no external data\n**(Hubmap LB 51 <- super increased , HPA LB 21 <- cv Up, but not reflected in LB)**\nas @hengck23 said in his topic, not much domain miss match between Hubmap and HPA?   \ntommorrow(2022/8/16, i will investigate scores of each organ)\none fold tiling cv score : **0.7885**\n**spleen val dice score : 0.8092**\n**lung val dice score : 0.2186**\n**kidney val dice score : 0.9471**\n**largeintestine val dice score : 0.9073**  \n**prostate val dice score : 0.8677**\n\nLocal Image prediction  \n![](https://i.imgur.com/sszC68I.png)\nPrediction on the test image \n![](https://i.imgur.com/9UnkJJF.png)\n\n2022/8/16\nMore Epochs improved the cv score but not LB\n**(Hubmap 51  , HPA 21, LB73 )**\none fold tiling cv score : **0.7939**\n**spleen val dice score : 0.8249**\n**lung val dice score : 0.2455**\n**kidney val dice score : 0.9549**\n**largeintestine val dice score : 0.9048**  \n**prostate val dice score : 0.8541**  \nwhy on earth prostate local score decreased?  \n  \n**Hubmap prostate LB : 0.13**  \n**Spleen LB : 0.12**  \n**Lung LB : 0.07**  \n**Kidney LB : 0.11**  \n**LargeIntestine LB : 0.05**  \nMy improvements are from spleen and lung, prostate!\nIt seems like the spleen can't be generalized well with the method of tiling since we need to see a large area\n\n\n2022/8/18  \nLocal Image Prediction  \n![](https://i.imgur.com/pDiq3dQ.png)  \nI have started to train the whole images.  CV dropped from 0.7885 -> 0.7635 but the inference became more realistic.  \n\n2022/8/19  \nimproved with some tuning .. **CV 80, LB 75, HUBMAP LB 53**  \nOnly using provided competition dataset, same model, my model has improved. awesome!\n",
      "votes": 18
    },
    {
      "id": 1899047,
      "postDate": "2022-08-15T03:21:29.307Z",
      "content": "<p>now that you have provide more information, it is easier to debug</p>\n<p>your scores</p>\n<pre><code>2022/8/14 \nI had investigated the Hubmap scores of each organ\nspleen hubmap score : 0.11 \nlung hubmap score : 0.03 \nlarge-intestine score : 0.05 \nkidney hubmap score : 0.11 \nprostate hubmap score : 0.06\n</code></pre>\n<p>my scores</p>\n<pre><code>        unet        unet\n        resnet101d  effnetb7 \n        [1]         [2]\n\n                      CV-HPA  0.7722    0.7886 \n       0.278350515    LB-HPA   0.754  0.754 \n       0.721649485    LB-Hubmap   0.721   0.748 \n                      LB-public   0.730     0.750 \n\nHubmap    all         0.520   0.540 \n    kidney        0.100     0.110 \n    prostate        0.130 \n    largeintestine  0.050   0.050 \n    spleen          0.140   0.150 \n    lung            0.080   0.090 \n\n---\nlocal CV score\n    [1] [2]\nall    0.77222     0.78857 \nkidney    0.94029     0.94566 \nprostate    0.82681     0.82665 \nlargeintestine    0.88933     0.89904 \nspleen    0.77511     0.81372 \nlung    0.17511     0.22954 \n</code></pre>\n<p>we note that kidney and large intestine are correct.<br>\n(although there are tiliing artifacts from the prediction mask you shown)</p>\n<p>your prostate score is very low. i suspect you simply just tile the prostate test images. this is wrong. the hidden hubmap prostate images at at 6 um (train hubmap is 0.4 um).<br>\nyou should rescale test (about 15 to 16x) prostate images before tiling.<br>\nyou can probe the prostate test images. i think they are about 160x160 only </p>\n<p>your local CV for spleen is low<br>\nfrom you spleen test image, you have over predicted. there are too many FTUs,<br>\nthis could be reason of your low score.<br>\nthis is because spleen FTU are larger (and less distinctive) than that of kidney and largeintestine.<br>\nthey probably require larger context (larger tile window) for accurate prediction.</p>\n<p>to prove this make a separate model for spleen using smaller resize image (while keeping your current tile window the same) for train and test</p>\n<p>for lung, i think you use the same threshold as the rest of the organ. you can lower it by learning a new threshold</p>\n<hr>\n<p>further,</p>\n<p>but your model and mine are using efficientB7.</p>\n<p>you are using tiling  and I am using whole image.<br>\nit is obvious that for the earlier FTU (kidney and large intestine) your score is better becuase you have larger resolution. (and maybe ASPP of deeplab can capture more \"complex mask\" where as unet mask are \"smooth\", i.e. without sharp turns, curvy outlines, …)</p>\n<hr>\n<p>yet further,</p>\n<ul>\n<li>kidney and large intestine are the easiest. most kagglers have differents core.</li>\n<li>spleen is difficult. this differentiate the gold medals from the rest.</li>\n<li>lung: a bit of \"by luck\" prediction. but nevertheless, the gold medal still produce better results than the rest.</li>\n</ul>\n<hr>\n<p>\"lung and prostate hugely failed to be generalized to HUBMAP\". this statement is wrong.</p>\n<p>prostate is a pre-processing bug.</p>\n<p>lung and spleen is modeling bug. (becuase your performance is already low in your local hpa domain). it should be problem of scale and context</p>\n<p>to verify you may want to waste another 5 submissions to probe your test HPA score</p>\n<hr>\n<p>maybe best of both world - local and global context?</p>\n<ul>\n<li>whole image prediction output some feature map</li>\n<li>concat this feature map to input rgb, then perform tiling</li>\n</ul>\n<p>but the annotation is already bad, so you maybe gain much after much efforts.</p>\n<hr>\n<p>actually many kagglers have problems. but they don't provide enough information in the post. hence there is little we can help.<br>\nthe first step in algorithm debug is to generate some results :)</p>",
      "rawMarkdown": "now that you have provide more information, it is easier to debug\n\nyour scores\n```\n2022/8/14 \nI had investigated the Hubmap scores of each organ\nspleen hubmap score : 0.11 \nlung hubmap score : 0.03 \nlarge-intestine score : 0.05 \nkidney hubmap score : 0.11 \nprostate hubmap score : 0.06\n\n\n```\n\nmy scores\n```\n\t\tunet        unet\n\t\tresnet101d\teffnetb7 \n\t\t[1]\t        [2]\n \t\t\n\t                  CV-HPA  0.7722 \t0.7886 \n       0.278350515\tLB-HPA   0.754 \t0.754 \n       0.721649485\tLB-Hubmap \t0.721 \t0.748 \n\t                  LB-public\t  0.730 \t0.750 \n\t\t\t\nHubmap\tall\t       \t0.520 \t0.540 \n\tkidney\t  \t  0.100 \t0.110 \n\tprostate\t\t0.130 \n\tlargeintestine\t0.050 \t0.050 \n\tspleen\t\t    0.140 \t0.150 \n\tlung\t\t\t0.080 \t0.090 \n\n---\nlocal CV score\n\t[1]\t[2]\nall\t0.77222 \t0.78857 \nkidney\t0.94029 \t0.94566 \nprostate\t0.82681 \t0.82665 \nlargeintestine\t0.88933 \t0.89904 \nspleen\t0.77511 \t0.81372 \nlung\t0.17511 \t0.22954 \n\n\n```\n\nwe note that kidney and large intestine are correct.\n(although there are tiliing artifacts from the prediction mask you shown)\n\nyour prostate score is very low. i suspect you simply just tile the prostate test images. this is wrong. the hidden hubmap prostate images at at 6 um (train hubmap is 0.4 um).\nyou should rescale test (about 15 to 16x) prostate images before tiling.\nyou can probe the prostate test images. i think they are about 160x160 only \n\n\nyour local CV for spleen is low\nfrom you spleen test image, you have over predicted. there are too many FTUs,\nthis could be reason of your low score.\nthis is because spleen FTU are larger (and less distinctive) than that of kidney and largeintestine.\nthey probably require larger context (larger tile window) for accurate prediction.\n\nto prove this make a separate model for spleen using smaller resize image (while keeping your current tile window the same) for train and test\n\nfor lung, i think you use the same threshold as the rest of the organ. you can lower it by learning a new threshold\n\n---\nfurther,\n\nbut your model and mine are using efficientB7.\n\nyou are using tiling  and I am using whole image.\nit is obvious that for the earlier FTU (kidney and large intestine) your score is better becuase you have larger resolution. (and maybe ASPP of deeplab can capture more \"complex mask\" where as unet mask are \"smooth\", i.e. without sharp turns, curvy outlines, ...)\n\n\n---\n\nyet further,\n- kidney and large intestine are the easiest. most kagglers have differents core.\n- spleen is difficult. this differentiate the gold medals from the rest.\n- lung: a bit of \"by luck\" prediction. but nevertheless, the gold medal still produce better results than the rest.\n\n\n---\n\n\"lung and prostate hugely failed to be generalized to HUBMAP\". this statement is wrong.\n\nprostate is a pre-processing bug.\n\nlung and spleen is modeling bug. (becuase your performance is already low in your local hpa domain). it should be problem of scale and context\n\nto verify you may want to waste another 5 submissions to probe your test HPA score\n\n---\nmaybe best of both world - local and global context?\n- whole image prediction output some feature map\n- concat this feature map to input rgb, then perform tiling\n\nbut the annotation is already bad, so you maybe gain much after much efforts.\n\n---\n\nactually many kagglers have problems. but they don't provide enough information in the post. hence there is little we can help.\nthe first step in algorithm debug is to generate some results :)",
      "votes": 5,
      "replies": [
        {
          "id": 1901363,
          "postDate": "2022-08-16T15:47:28.067Z",
          "content": "<p>Thanks, following your advice, much improved.. but i have to narrow the gap more!</p>",
          "rawMarkdown": "Thanks, following your advice, much improved.. but i have to narrow the gap more!"
        },
        {
          "id": 1901385,
          "postDate": "2022-08-16T16:03:40.737Z",
          "content": "<p>if you can post some results and Lb scores, I can help you to anaylse what is wrong</p>",
          "rawMarkdown": "if you can post some results and Lb scores, I can help you to anaylse what is wrong",
          "votes": 1
        }
      ]
    },
    {
      "id": 1893731,
      "postDate": "2022-08-11T03:32:53.157Z",
      "content": "<p>best LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.<br>\ni suppose local CV (HPA) is also about 0.80.</p>\n<p>(if you are getting 0.90, there is something wrong with your evaluation. note that evaluation on tiles is different from evaluation full image. you should always compute dice for the whole image after any post processing)</p>\n<p>So there is really not so much domain difference as i have initially expected.<br>\n(separate probe shows that LB-humap, LB-HPA are smiliar in public  test set)</p>\n<p>both HPA and Human have a common stain: the blue color.<br>\njust that one is blue+brown and other is blue+pink</p>",
      "rawMarkdown": "best LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.\ni suppose local CV (HPA) is also about 0.80.\n\n(if you are getting 0.90, there is something wrong with your evaluation. note that evaluation on tiles is different from evaluation full image. you should always compute dice for the whole image after any post processing)\n\nSo there is really not so much domain difference as i have initially expected.\n(separate probe shows that LB-humap, LB-HPA are smiliar in public  test set)\n\nboth HPA and Human have a common stain: the blue color.\njust that one is blue+brown and other is blue+pink\n\n\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 1894028,
          "postDate": "2022-08-11T08:25:21.377Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Do you mean only training \"HPA images\" can bring us LB 0.8 without (any domain adaptation or pixel size adjusting when training?<br>\nDo you mean only reading tiff images and reshaping the images to some sizes, after normalizing the images to imagenet mean and std or dividing it 255, after training some segmentation models, it can bring us LB 0.8? with the method of tiling or whole images</p>",
          "rawMarkdown": "@hengck23  Do you mean only training \"HPA images\" can bring us LB 0.8 without (any domain adaptation or pixel size adjusting when training?\nDo you mean only reading tiff images and reshaping the images to some sizes, after normalizing the images to imagenet mean and std or dividing it 255, after training some segmentation models, it can bring us LB 0.8? with the method of tiling or whole images\n"
        },
        {
          "id": 1899074,
          "postDate": "2022-08-15T03:55:51.570Z",
          "content": "<p>check my inference framwork here:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb-0-78-coat-with-no-decoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-78-coat-with-no-decoder</a></p>\n<p>for tiling, refer to<br>\n<a href=\"https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds\" target=\"_blank\">https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds</a></p>",
          "rawMarkdown": "check my inference framwork here:\nhttps://www.kaggle.com/code/hengck23/lb-0-78-coat-with-no-decoder\n\nfor tiling, refer to\nhttps://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds",
          "votes": 1
        },
        {
          "id": 1900279,
          "postDate": "2022-08-15T21:35:56.360Z",
          "content": "<p>Hi, do you think that training with augmented images(as shown here <a href=\"https://www.kaggle.com/code/gunesevitan/pixel-size-and-tissue-thickness-domain-adaptation/notebook\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/pixel-size-and-tissue-thickness-domain-adaptation/notebook</a>) would make any difference? I think it will give lower score on LB, but it seems to me it will have quite decent score on private set, as all of them contain HuBMAP data.</p>",
          "rawMarkdown": "Hi, do you think that training with augmented images(as shown here https://www.kaggle.com/code/gunesevitan/pixel-size-and-tissue-thickness-domain-adaptation/notebook) would make any difference? I think it will give lower score on LB, but it seems to me it will have quite decent score on private set, as all of them contain HuBMAP data."
        },
        {
          "id": 1903350,
          "postDate": "2022-08-17T10:48:02.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/nurkhanlaiyk\" target=\"_blank\">@nurkhanlaiyk</a> in my case, augmentation made differences in terms of HUBMAP but not much in HPA  </p>",
          "rawMarkdown": "@nurkhanlaiyk in my case, augmentation made differences in terms of HUBMAP but not much in HPA  "
        },
        {
          "id": 1903378,
          "postDate": "2022-08-17T11:26:32.867Z",
          "content": "<p>Did you you use tiled images? Or took it as whole?</p>",
          "rawMarkdown": "Did you you use tiled images? Or took it as whole?"
        },
        {
          "id": 1903419,
          "postDate": "2022-08-17T12:24:06.777Z",
          "content": "<p>I tried both, but only tiling worked</p>",
          "rawMarkdown": "I tried both, but only tiling worked"
        }
      ]
    },
    {
      "id": 1893386,
      "postDate": "2022-08-10T18:47:38.397Z",
      "content": "<p>I don't know what's weird. Domain adaptation is a very common problem in real world computer vision projects.</p>",
      "rawMarkdown": "I don't know what's weird. Domain adaptation is a very common problem in real world computer vision projects.",
      "votes": 4,
      "replies": [
        {
          "id": 1894025,
          "postDate": "2022-08-11T08:24:33.720Z",
          "content": "<p>I will study your domain adatation notebook, thanks</p>",
          "rawMarkdown": "I will study your domain adatation notebook, thanks"
        }
      ]
    },
    {
      "id": 1893338,
      "postDate": "2022-08-10T18:04:50.163Z",
      "content": "<p>That is because the LB consists of HuBMAP + HPA while the train set only consists of HuBMAP, you can read more about it here: <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/333704\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/333704</a></p>",
      "rawMarkdown": "That is because the LB consists of HuBMAP + HPA while the train set only consists of HuBMAP, you can read more about it here: https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/333704",
      "votes": 4,
      "replies": [
        {
          "id": 1894029,
          "postDate": "2022-08-11T08:25:37.550Z",
          "content": "<p>Thanks for the good link!</p>",
          "rawMarkdown": "Thanks for the good link!"
        }
      ]
    },
    {
      "id": 1930585,
      "postDate": "2022-09-08T04:06:56.327Z",
      "content": "<p>May I ask what kind of augmentations did you use?</p>",
      "rawMarkdown": "May I ask what kind of augmentations did you use?",
      "votes": 1,
      "replies": [
        {
          "id": 1944841,
          "postDate": "2022-09-18T16:08:30.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> sorry to late reply  <br>\nI mostly apply flip, rotate, hsv, grid distortion </p>",
          "rawMarkdown": "@leonshangguan sorry to late reply  \nI mostly apply flip, rotate, hsv, grid distortion ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1901565,
      "postDate": "2022-08-16T18:43:28.460Z",
      "content": "<p>\"Hubmap 51 , HPA 21, LB73 \"</p>\n<p>this is reasonable score.<br>\nHPA 32 can gives Hubmap 52 to 54 depends on encoder, decoder.<br>\nyou can try different combination once you fixed your tiling and whole image bug. </p>",
      "rawMarkdown": "\"Hubmap 51 , HPA 21, LB73 \"\n\nthis is reasonable score.\nHPA 32 can gives Hubmap 52 to 54 depends on encoder, decoder.\nyou can try different combination once you fixed your tiling and whole image bug. \n\n ",
      "votes": 1
    },
    {
      "id": 1901555,
      "postDate": "2022-08-16T18:31:09.280Z",
      "content": "<p>\"why on earth prostate local score decreased?\"</p>\n<p>there are only a few prostate validation images. <br>\nchanges in just one image or one FTU can causes large change in scores.</p>\n<p>the label are not consistent (label noise). lower score can mean \"better results\".<br>\nit is better that you inspect the results and find the reason for low score, e.g.</p>\n<p>high score = prediction missed a non labelled FTU.<br>\nlow score = predicted detects the  non labelled FTU.</p>",
      "rawMarkdown": "\"why on earth prostate local score decreased?\"\n\nthere are only a few prostate validation images. \nchanges in just one image or one FTU can causes large change in scores.\n\nthe label are not consistent (label noise). lower score can mean \"better results\".\nit is better that you inspect the results and find the reason for low score, e.g.\n\nhigh score = prediction missed a non labelled FTU.\nlow score = predicted detects the  non labelled FTU.",
      "votes": 1
    },
    {
      "id": 1901032,
      "postDate": "2022-08-16T12:38:27.813Z",
      "content": "<p>I am same to you when I began this game, and the lb-score is rising when I give the more epochs and mode parameters, try more SOTA models like <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I learn much  from his work</p>",
      "rawMarkdown": "I am same to you when I began this game, and the lb-score is rising when I give the more epochs and mode parameters, try more SOTA models like @hengck23, I learn much  from his work",
      "votes": 1,
      "replies": [
        {
          "id": 1901035,
          "postDate": "2022-08-16T12:39:04.650Z",
          "content": "<p>more parameters, my english is not good</p>",
          "rawMarkdown": "more parameters, my english is not good",
          "votes": 1
        },
        {
          "id": 1901370,
          "postDate": "2022-08-16T15:52:22.380Z",
          "content": "<p>it seems like <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> made models trained on full images, not by tiliing<br>\ndo you also use full images when training?</p>",
          "rawMarkdown": "it seems like @hengck23 made models trained on full images, not by tiliing\ndo you also use full images when training?"
        },
        {
          "id": 1901398,
          "postDate": "2022-08-16T16:08:12.837Z",
          "content": "<p>whether to use tiling or not as your final solution, you should perform both whole image and tiling.<br>\ne.g. say I want to make a system of 1000x1000, tile at 500x500, stride at 250.</p>\n<p>the full image image system at 1000x1000 is for me to debug the post processing of my tiling system.<br>\nthe accuracy should not differs by too much.</p>\n<p>but then there is one question, how do you train a 1000x1000 system?<br>\nuse tiling as pretrain and whole image as finetune</p>",
          "rawMarkdown": "whether to use tiling or not as your final solution, you should perform both whole image and tiling.\ne.g. say I want to make a system of 1000x1000, tile at 500x500, stride at 250.\n\nthe full image image system at 1000x1000 is for me to debug the post processing of my tiling system.\nthe accuracy should not differs by too much.\n\nbut then there is one question, how do you train a 1000x1000 system?\nuse tiling as pretrain and whole image as finetune",
          "votes": 1
        },
        {
          "id": 1901424,
          "postDate": "2022-08-16T16:32:39.453Z",
          "content": "<p>At this step, I trained full image(resized from 3000x3000 to 768x768) and tile image(patches cropped with 1024 size and resized to 768 sizes) \"at the same time\" and check cv scores for tiled patch and  full images  <br>\nI mean when I define dataset in PyTorch, sometimes we get patches, sometimes we get full images<br>\nThe resultant cv score of the tile method is around 0.78~0.79 but the cv score of the full image is around 0.4 dice score(i am debugging where my training method is wrong)  <br>\nalso, I tried to train the full image based on the pre-trained patch-based model, but I can't score more than 0.5 dice score. it's weird. Considering having seen some notebooks just using TensorFlow, efficientnetb0+unet can get 0.58 and efficientnetb7+deeplab can score 0.68. so I am debugging the process</p>",
          "rawMarkdown": "At this step, I trained full image(resized from 3000x3000 to 768x768) and tile image(patches cropped with 1024 size and resized to 768 sizes) \"at the same time\" and check cv scores for tiled patch and  full images  \nI mean when I define dataset in PyTorch, sometimes we get patches, sometimes we get full images\nThe resultant cv score of the tile method is around 0.78~0.79 but the cv score of the full image is around 0.4 dice score(i am debugging where my training method is wrong)  \nalso, I tried to train the full image based on the pre-trained patch-based model, but I can't score more than 0.5 dice score. it's weird. Considering having seen some notebooks just using TensorFlow, efficientnetb0+unet can get 0.58 and efficientnetb7+deeplab can score 0.68. so I am debugging the process"
        },
        {
          "id": 1903083,
          "postDate": "2022-08-17T05:01:27.320Z",
          "content": "<p>I think the data-distribution  for training and evaluation needs to be same, such as pixel size, or object scale .etc, which is the best state of input for model processing</p>",
          "rawMarkdown": "I think the data-distribution  for training and evaluation needs to be same, such as pixel size, or object scale .etc, which is the best state of input for model processing",
          "votes": 1
        },
        {
          "id": 1904173,
          "postDate": "2022-08-18T02:57:27.557Z",
          "content": "<p>\"data-distribution for training and evaluation needs to be same\"</p>\n<p>it is probility true if you are talking about the  loss (e.g bce).</p>\n<p>but when it comes to performance metric, it may not be true.<br>\ne.g single crop classification is better if you enlarge your image for evaluation in imagenet papers.<br>\ngraph of \" train_loss vs valid_loss  \" and  \" train_loss vs valid_dice  \" also reveal that the two grph don't coincide</p>",
          "rawMarkdown": "\"data-distribution for training and evaluation needs to be same\"\n\nit is probility true if you are talking about the  loss (e.g bce).\n\nbut when it comes to performance metric, it may not be true.\ne.g single crop classification is better if you enlarge your image for evaluation in imagenet papers.\ngraph of \" train\\_loss vs valid\\_loss  \" and  \" train\\_loss vs valid\\_dice  \" also reveal that the two grph don't coincide",
          "votes": 1
        }
      ]
    },
    {
      "id": 1893705,
      "postDate": "2022-08-11T02:37:39.100Z",
      "content": "<p>Try more augmentation, it could be better</p>",
      "rawMarkdown": "Try more augmentation, it could be better",
      "votes": 1,
      "replies": [
        {
          "id": 1894024,
          "postDate": "2022-08-11T08:24:11.757Z",
          "content": "<p>I will try, thanks!</p>",
          "rawMarkdown": "I will try, thanks!"
        },
        {
          "id": 1895879,
          "postDate": "2022-08-12T12:18:51.657Z",
          "content": "<p>Using heavy augmentations improved my LB score from ~0.50 to ~0.60, it definitely seems to help!</p>",
          "rawMarkdown": "Using heavy augmentations improved my LB score from ~0.50 to ~0.60, it definitely seems to help!",
          "votes": 2
        },
        {
          "id": 1897062,
          "postDate": "2022-08-13T12:02:58.310Z",
          "content": "<p>Thanks a lot</p>",
          "rawMarkdown": "Thanks a lot"
        },
        {
          "id": 1944068,
          "postDate": "2022-09-18T03:31:51.370Z",
          "content": "<p>Hi, I am a beginner, can I ask you what kind of heavy data augmentation you use to reduce the gap between CV and LB during training and testing, thank you very much!</p>",
          "rawMarkdown": "Hi, I am a beginner, can I ask you what kind of heavy data augmentation you use to reduce the gap between CV and LB during training and testing, thank you very much!"
        }
      ]
    },
    {
      "id": 1893809,
      "postDate": "2022-08-11T05:03:47.683Z",
      "content": "<p>I'm sorry guys for my comment that \"this competition is a joke\", it joust a little bit frustrating to put so much work in to get only 50%.<br>\nI'm quite puzzled by my results.</p>\n<pre><code>best LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.\ni suppose local CV (HPA) is also about 0.80.\n</code></pre>\n<p>I'm even more puzzled by this😄<br>\nHow long did you train your model to get such a result?</p>",
      "rawMarkdown": "I'm sorry guys for my comment that \"this competition is a joke\", it joust a little bit frustrating to put so much work in to get only 50%.\nI'm quite puzzled by my results.\n\n```\nbest LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.\ni suppose local CV (HPA) is also about 0.80.\n```\n\nI'm even more puzzled by this😄\nHow long did you train your model to get such a result?",
      "votes": 2,
      "replies": [
        {
          "id": 1893813,
          "postDate": "2022-08-11T05:14:09.563Z",
          "content": "<p>check my comment and code at <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941</a><br>\njust one fold with training about 2 hr can get you ~0.78</p>",
          "rawMarkdown": "check my comment and code at https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941\njust one fold with training about 2 hr can get you ~0.78",
          "votes": 2
        },
        {
          "id": 1894040,
          "postDate": "2022-08-11T08:34:02.300Z",
          "content": "<p>some new notebook by  <a href=\"https://www.kaggle.com/w3579628328\" target=\"_blank\">@w3579628328</a> </p>\n<p>mmsegmentation LB0.78 inference (1/5folds)<br>\n<a href=\"https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds\" target=\"_blank\">https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds</a></p>",
          "rawMarkdown": "some new notebook by  @w3579628328 \n\nmmsegmentation LB0.78 inference (1/5folds)\nhttps://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds",
          "votes": 1
        }
      ]
    },
    {
      "id": 1927956,
      "postDate": "2022-09-06T05:51:26.347Z",
      "content": "<p>How did you solve the issue in the end? I keep having CV 0.8 LB 0.6 even with augmentatios</p>",
      "rawMarkdown": "How did you solve the issue in the end? I keep having CV 0.8 LB 0.6 even with augmentatios",
      "replies": [
        {
          "id": 1944842,
          "postDate": "2022-09-18T16:09:14.743Z",
          "content": "<p>very long epoch training</p>",
          "rawMarkdown": "very long epoch training"
        }
      ]
    },
    {
      "id": 1894194,
      "postDate": "2022-08-11T10:13:43.697Z",
      "content": "<p>I meet the same condition, <a href=\"https://www.kaggle.com/befunny\" target=\"_blank\">@befunny</a> Hi,wangkui ,I know you use the mmsegmentation in the competition ,i use it published too,i resize the whole img to (1536,1536) to train and test in (1536,1536)size,it is perform good in my CV(0.75),but i just got 0.52 in  my LB,I don't know why, did i  <br>\nmiss some details ?can you help me?which really bothering me for a long time，i will really appreciate it!🙏</p>",
      "rawMarkdown": "I meet the same condition, @befunny Hi,wangkui ,I know you use the mmsegmentation in the competition ,i use it published too,i resize the whole img to (1536,1536) to train and test in (1536,1536)size,it is perform good in my CV(0.75),but i just got 0.52 in  my LB,I don't know why, did i  \nmiss some details ?can you help me?which really bothering me for a long time，i will really appreciate it!🙏",
      "replies": [
        {
          "id": 1895055,
          "postDate": "2022-08-11T23:36:15.740Z",
          "content": "<p>I haven't used mmseg yet，I think a possible problem is that the resolution is too large, you should notice the hidden test images's size. </p>\n<blockquote>\n  <p>The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.<br>\n  I just use 768 follow <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>\n</blockquote>",
          "rawMarkdown": "I haven't used mmseg yet，I think a possible problem is that the resolution is too large, you should notice the hidden test images's size. \n> The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.\nI just use 768 follow @hengck23"
        }
      ]
    },
    {
      "id": 1893359,
      "postDate": "2022-08-10T18:25:32.777Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1899047,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-08-15T03:21:29.307000",
      "content": "<p>now that you have provide more information, it is easier to debug</p>\n<p>your scores</p>\n<pre><code>2022/8/14 \nI had investigated the Hubmap scores of each organ\nspleen hubmap score : 0.11 \nlung hubmap score : 0.03 \nlarge-intestine score : 0.05 \nkidney hubmap score : 0.11 \nprostate hubmap score : 0.06\n</code></pre>\n<p>my scores</p>\n<pre><code>        unet        unet\n        resnet101d  effnetb7 \n        [1]         [2]\n\n                      CV-HPA  0.7722    0.7886 \n       0.278350515    LB-HPA   0.754  0.754 \n       0.721649485    LB-Hubmap   0.721   0.748 \n                      LB-public   0.730     0.750 \n\nHubmap    all         0.520   0.540 \n    kidney        0.100     0.110 \n    prostate        0.130 \n    largeintestine  0.050   0.050 \n    spleen          0.140   0.150 \n    lung            0.080   0.090 \n\n---\nlocal CV score\n    [1] [2]\nall    0.77222     0.78857 \nkidney    0.94029     0.94566 \nprostate    0.82681     0.82665 \nlargeintestine    0.88933     0.89904 \nspleen    0.77511     0.81372 \nlung    0.17511     0.22954 \n</code></pre>\n<p>we note that kidney and large intestine are correct.<br>\n(although there are tiliing artifacts from the prediction mask you shown)</p>\n<p>your prostate score is very low. i suspect you simply just tile the prostate test images. this is wrong. the hidden hubmap prostate images at at 6 um (train hubmap is 0.4 um).<br>\nyou should rescale test (about 15 to 16x) prostate images before tiling.<br>\nyou can probe the prostate test images. i think they are about 160x160 only </p>\n<p>your local CV for spleen is low<br>\nfrom you spleen test image, you have over predicted. there are too many FTUs,<br>\nthis could be reason of your low score.<br>\nthis is because spleen FTU are larger (and less distinctive) than that of kidney and largeintestine.<br>\nthey probably require larger context (larger tile window) for accurate prediction.</p>\n<p>to prove this make a separate model for spleen using smaller resize image (while keeping your current tile window the same) for train and test</p>\n<p>for lung, i think you use the same threshold as the rest of the organ. you can lower it by learning a new threshold</p>\n<hr>\n<p>further,</p>\n<p>but your model and mine are using efficientB7.</p>\n<p>you are using tiling  and I am using whole image.<br>\nit is obvious that for the earlier FTU (kidney and large intestine) your score is better becuase you have larger resolution. (and maybe ASPP of deeplab can capture more \"complex mask\" where as unet mask are \"smooth\", i.e. without sharp turns, curvy outlines, …)</p>\n<hr>\n<p>yet further,</p>\n<ul>\n<li>kidney and large intestine are the easiest. most kagglers have differents core.</li>\n<li>spleen is difficult. this differentiate the gold medals from the rest.</li>\n<li>lung: a bit of \"by luck\" prediction. but nevertheless, the gold medal still produce better results than the rest.</li>\n</ul>\n<hr>\n<p>\"lung and prostate hugely failed to be generalized to HUBMAP\". this statement is wrong.</p>\n<p>prostate is a pre-processing bug.</p>\n<p>lung and spleen is modeling bug. (becuase your performance is already low in your local hpa domain). it should be problem of scale and context</p>\n<p>to verify you may want to waste another 5 submissions to probe your test HPA score</p>\n<hr>\n<p>maybe best of both world - local and global context?</p>\n<ul>\n<li>whole image prediction output some feature map</li>\n<li>concat this feature map to input rgb, then perform tiling</li>\n</ul>\n<p>but the annotation is already bad, so you maybe gain much after much efforts.</p>\n<hr>\n<p>actually many kagglers have problems. but they don't provide enough information in the post. hence there is little we can help.<br>\nthe first step in algorithm debug is to generate some results :)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1901363,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-16T15:47:28.067000",
          "content": "<p>Thanks, following your advice, much improved.. but i have to narrow the gap more!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1901385,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-08-16T16:03:40.737000",
          "content": "<p>if you can post some results and Lb scores, I can help you to anaylse what is wrong</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1893731,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-08-11T03:32:53.157000",
      "content": "<p>best LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.<br>\ni suppose local CV (HPA) is also about 0.80.</p>\n<p>(if you are getting 0.90, there is something wrong with your evaluation. note that evaluation on tiles is different from evaluation full image. you should always compute dice for the whole image after any post processing)</p>\n<p>So there is really not so much domain difference as i have initially expected.<br>\n(separate probe shows that LB-humap, LB-HPA are smiliar in public  test set)</p>\n<p>both HPA and Human have a common stain: the blue color.<br>\njust that one is blue+brown and other is blue+pink</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1894028,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-11T08:25:21.377000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Do you mean only training \"HPA images\" can bring us LB 0.8 without (any domain adaptation or pixel size adjusting when training?<br>\nDo you mean only reading tiff images and reshaping the images to some sizes, after normalizing the images to imagenet mean and std or dividing it 255, after training some segmentation models, it can bring us LB 0.8? with the method of tiling or whole images</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1899074,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-08-15T03:55:51.570000",
          "content": "<p>check my inference framwork here:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb-0-78-coat-with-no-decoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-78-coat-with-no-decoder</a></p>\n<p>for tiling, refer to<br>\n<a href=\"https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds\" target=\"_blank\">https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1900279,
          "author_name": "Nurkhan Laiyk",
          "author_url": "",
          "post_date": "2022-08-15T21:35:56.360000",
          "content": "<p>Hi, do you think that training with augmented images(as shown here <a href=\"https://www.kaggle.com/code/gunesevitan/pixel-size-and-tissue-thickness-domain-adaptation/notebook\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/pixel-size-and-tissue-thickness-domain-adaptation/notebook</a>) would make any difference? I think it will give lower score on LB, but it seems to me it will have quite decent score on private set, as all of them contain HuBMAP data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1903350,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-17T10:48:02.507000",
          "content": "<p><a href=\"https://www.kaggle.com/nurkhanlaiyk\" target=\"_blank\">@nurkhanlaiyk</a> in my case, augmentation made differences in terms of HUBMAP but not much in HPA  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1903378,
          "author_name": "Nurkhan Laiyk",
          "author_url": "",
          "post_date": "2022-08-17T11:26:32.867000",
          "content": "<p>Did you you use tiled images? Or took it as whole?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1903419,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-17T12:24:06.777000",
          "content": "<p>I tried both, but only tiling worked</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1893386,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-08-10T18:47:38.397000",
      "content": "<p>I don't know what's weird. Domain adaptation is a very common problem in real world computer vision projects.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1894025,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-11T08:24:33.720000",
          "content": "<p>I will study your domain adatation notebook, thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1893338,
      "author_name": "Szymon Ożóg",
      "author_url": "",
      "post_date": "2022-08-10T18:04:50.163000",
      "content": "<p>That is because the LB consists of HuBMAP + HPA while the train set only consists of HuBMAP, you can read more about it here: <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/333704\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/333704</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1894029,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-11T08:25:37.550000",
          "content": "<p>Thanks for the good link!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1930585,
      "author_name": "Zhongkai Shangguan",
      "author_url": "",
      "post_date": "2022-09-08T04:06:56.327000",
      "content": "<p>May I ask what kind of augmentations did you use?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1944841,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-09-18T16:08:30.263000",
          "content": "<p><a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> sorry to late reply  <br>\nI mostly apply flip, rotate, hsv, grid distortion </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1901565,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-08-16T18:43:28.460000",
      "content": "<p>\"Hubmap 51 , HPA 21, LB73 \"</p>\n<p>this is reasonable score.<br>\nHPA 32 can gives Hubmap 52 to 54 depends on encoder, decoder.<br>\nyou can try different combination once you fixed your tiling and whole image bug. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1901555,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-08-16T18:31:09.280000",
      "content": "<p>\"why on earth prostate local score decreased?\"</p>\n<p>there are only a few prostate validation images. <br>\nchanges in just one image or one FTU can causes large change in scores.</p>\n<p>the label are not consistent (label noise). lower score can mean \"better results\".<br>\nit is better that you inspect the results and find the reason for low score, e.g.</p>\n<p>high score = prediction missed a non labelled FTU.<br>\nlow score = predicted detects the  non labelled FTU.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1901032,
      "author_name": "跟着丞相割麦子",
      "author_url": "",
      "post_date": "2022-08-16T12:38:27.813000",
      "content": "<p>I am same to you when I began this game, and the lb-score is rising when I give the more epochs and mode parameters, try more SOTA models like <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I learn much  from his work</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1901035,
          "author_name": "跟着丞相割麦子",
          "author_url": "",
          "post_date": "2022-08-16T12:39:04.650000",
          "content": "<p>more parameters, my english is not good</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1901370,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-16T15:52:22.380000",
          "content": "<p>it seems like <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> made models trained on full images, not by tiliing<br>\ndo you also use full images when training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1901398,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-08-16T16:08:12.837000",
          "content": "<p>whether to use tiling or not as your final solution, you should perform both whole image and tiling.<br>\ne.g. say I want to make a system of 1000x1000, tile at 500x500, stride at 250.</p>\n<p>the full image image system at 1000x1000 is for me to debug the post processing of my tiling system.<br>\nthe accuracy should not differs by too much.</p>\n<p>but then there is one question, how do you train a 1000x1000 system?<br>\nuse tiling as pretrain and whole image as finetune</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1901424,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-16T16:32:39.453000",
          "content": "<p>At this step, I trained full image(resized from 3000x3000 to 768x768) and tile image(patches cropped with 1024 size and resized to 768 sizes) \"at the same time\" and check cv scores for tiled patch and  full images  <br>\nI mean when I define dataset in PyTorch, sometimes we get patches, sometimes we get full images<br>\nThe resultant cv score of the tile method is around 0.78~0.79 but the cv score of the full image is around 0.4 dice score(i am debugging where my training method is wrong)  <br>\nalso, I tried to train the full image based on the pre-trained patch-based model, but I can't score more than 0.5 dice score. it's weird. Considering having seen some notebooks just using TensorFlow, efficientnetb0+unet can get 0.58 and efficientnetb7+deeplab can score 0.68. so I am debugging the process</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1903083,
          "author_name": "跟着丞相割麦子",
          "author_url": "",
          "post_date": "2022-08-17T05:01:27.320000",
          "content": "<p>I think the data-distribution  for training and evaluation needs to be same, such as pixel size, or object scale .etc, which is the best state of input for model processing</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1904173,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-08-18T02:57:27.557000",
          "content": "<p>\"data-distribution for training and evaluation needs to be same\"</p>\n<p>it is probility true if you are talking about the  loss (e.g bce).</p>\n<p>but when it comes to performance metric, it may not be true.<br>\ne.g single crop classification is better if you enlarge your image for evaluation in imagenet papers.<br>\ngraph of \" train_loss vs valid_loss  \" and  \" train_loss vs valid_dice  \" also reveal that the two grph don't coincide</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1893705,
      "author_name": "Ksana",
      "author_url": "",
      "post_date": "2022-08-11T02:37:39.100000",
      "content": "<p>Try more augmentation, it could be better</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1894024,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-11T08:24:11.757000",
          "content": "<p>I will try, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1895879,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2022-08-12T12:18:51.657000",
          "content": "<p>Using heavy augmentations improved my LB score from ~0.50 to ~0.60, it definitely seems to help!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1897062,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-08-13T12:02:58.310000",
          "content": "<p>Thanks a lot</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1944068,
          "author_name": "TheMinimalist",
          "author_url": "",
          "post_date": "2022-09-18T03:31:51.370000",
          "content": "<p>Hi, I am a beginner, can I ask you what kind of heavy data augmentation you use to reduce the gap between CV and LB during training and testing, thank you very much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1893809,
      "author_name": "Uros Jarc",
      "author_url": "",
      "post_date": "2022-08-11T05:03:47.683000",
      "content": "<p>I'm sorry guys for my comment that \"this competition is a joke\", it joust a little bit frustrating to put so much work in to get only 50%.<br>\nI'm quite puzzled by my results.</p>\n<pre><code>best LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.\ni suppose local CV (HPA) is also about 0.80.\n</code></pre>\n<p>I'm even more puzzled by this😄<br>\nHow long did you train your model to get such a result?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1893813,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-08-11T05:14:09.563000",
          "content": "<p>check my comment and code at <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332941</a><br>\njust one fold with training about 2 hr can get you ~0.78</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1894040,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-08-11T08:34:02.300000",
          "content": "<p>some new notebook by  <a href=\"https://www.kaggle.com/w3579628328\" target=\"_blank\">@w3579628328</a> </p>\n<p>mmsegmentation LB0.78 inference (1/5folds)<br>\n<a href=\"https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds\" target=\"_blank\">https://www.kaggle.com/code/w3579628328/mmsegmentation-lb0-78-inference-1-5folds</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1927956,
      "author_name": "Szymon Ożóg",
      "author_url": "",
      "post_date": "2022-09-06T05:51:26.347000",
      "content": "<p>How did you solve the issue in the end? I keep having CV 0.8 LB 0.6 even with augmentatios</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1944842,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-09-18T16:09:14.743000",
          "content": "<p>very long epoch training</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1894194,
      "author_name": "Alexis",
      "author_url": "",
      "post_date": "2022-08-11T10:13:43.697000",
      "content": "<p>I meet the same condition, <a href=\"https://www.kaggle.com/befunny\" target=\"_blank\">@befunny</a> Hi,wangkui ,I know you use the mmsegmentation in the competition ,i use it published too,i resize the whole img to (1536,1536) to train and test in (1536,1536)size,it is perform good in my CV(0.75),but i just got 0.52 in  my LB,I don't know why, did i  <br>\nmiss some details ?can you help me?which really bothering me for a long time，i will really appreciate it!🙏</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1895055,
          "author_name": "Ksana",
          "author_url": "",
          "post_date": "2022-08-11T23:36:15.740000",
          "content": "<p>I haven't used mmseg yet，I think a possible problem is that the resolution is too large, you should notice the hidden test images's size. </p>\n<blockquote>\n  <p>The HuBMAP images range in size from 4500x4500 down to 160x160 pixels.<br>\n  I just use 768 follow <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1893359,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-10T18:25:32.777000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1893308": "My Efficientnet7 + DeeplabV3Plus can score around 0.7837 with **tiling**, 100epoch in Local CV(single fold) but when i submit it, I just get only 0.60\n2022/8/12 \none fold tiling cv score : **0.7837**\n**(HUBMAP LB 39 , HPA LB 21)**  \n**spleen val dice score : 0.7772**\n**lung val dice score : 0.2246**\n**kidney val dice score : 0.9474**\n**largeintestine val dice score : 0.9093**  \n**prostate val dice score : 0.8610**\n\n2022/8/14  \nI had investigated the Hubmap scores of each organ\n**spleen hubmap score  : 0.11**  \n**lung hubmap score : 0.03**  \n**large-intestine score : 0.05**  \n**kidney hubmap score : 0.11**  \n**prostate hubmap score : 0.06**  \n\nLocal Image prediction  \n![](https://i.imgur.com/DDdnYwT.png)\nPrediction on the test image\n![](https://i.imgur.com/7EPLSG8.png)\n\nlung and prostate scores seem to bad compared to other competitors\nlung and prostate hugely failed to be generalized to HUBMAP\n\nI have been working day and night for this comp.. but i am very depressed by the fact that i can't cross even 0.6\n\n2022/8/15  \nI just added more augmentations than before, no external data\n**(Hubmap LB 51 <- super increased , HPA LB 21 <- cv Up, but not reflected in LB)**\nas @hengck23 said in his topic, not much domain miss match between Hubmap and HPA?   \ntommorrow(2022/8/16, i will investigate scores of each organ)\none fold tiling cv score : **0.7885**\n**spleen val dice score : 0.8092**\n**lung val dice score : 0.2186**\n**kidney val dice score : 0.9471**\n**largeintestine val dice score : 0.9073**  \n**prostate val dice score : 0.8677**\n\nLocal Image prediction  \n![](https://i.imgur.com/sszC68I.png)\nPrediction on the test image \n![](https://i.imgur.com/9UnkJJF.png)\n\n2022/8/16\nMore Epochs improved the cv score but not LB\n**(Hubmap 51  , HPA 21, LB73 )**\none fold tiling cv score : **0.7939**\n**spleen val dice score : 0.8249**\n**lung val dice score : 0.2455**\n**kidney val dice score : 0.9549**\n**largeintestine val dice score : 0.9048**  \n**prostate val dice score : 0.8541**  \nwhy on earth prostate local score decreased?  \n  \n**Hubmap prostate LB : 0.13**  \n**Spleen LB : 0.12**  \n**Lung LB : 0.07**  \n**Kidney LB : 0.11**  \n**LargeIntestine LB : 0.05**  \nMy improvements are from spleen and lung, prostate!\nIt seems like the spleen can't be generalized well with the method of tiling since we need to see a large area\n\n\n2022/8/18  \nLocal Image Prediction  \n![](https://i.imgur.com/pDiq3dQ.png)  \nI have started to train the whole images.  CV dropped from 0.7885 -> 0.7635 but the inference became more realistic.  \n\n2022/8/19  \nimproved with some tuning .. **CV 80, LB 75, HUBMAP LB 53**  \nOnly using provided competition dataset, same model, my model has improved. awesome!\n",
    "1899047": "now that you have provide more information, it is easier to debug\n\nyour scores\n```\n2022/8/14 \nI had investigated the Hubmap scores of each organ\nspleen hubmap score : 0.11 \nlung hubmap score : 0.03 \nlarge-intestine score : 0.05 \nkidney hubmap score : 0.11 \nprostate hubmap score : 0.06\n\n\n```\n\nmy scores\n```\n\t\tunet        unet\n\t\tresnet101d\teffnetb7 \n\t\t[1]\t        [2]\n \t\t\n\t                  CV-HPA  0.7722 \t0.7886 \n       0.278350515\tLB-HPA   0.754 \t0.754 \n       0.721649485\tLB-Hubmap \t0.721 \t0.748 \n\t                  LB-public\t  0.730 \t0.750 \n\t\t\t\nHubmap\tall\t       \t0.520 \t0.540 \n\tkidney\t  \t  0.100 \t0.110 \n\tprostate\t\t0.130 \n\tlargeintestine\t0.050 \t0.050 \n\tspleen\t\t    0.140 \t0.150 \n\tlung\t\t\t0.080 \t0.090 \n\n---\nlocal CV score\n\t[1]\t[2]\nall\t0.77222 \t0.78857 \nkidney\t0.94029 \t0.94566 \nprostate\t0.82681 \t0.82665 \nlargeintestine\t0.88933 \t0.89904 \nspleen\t0.77511 \t0.81372 \nlung\t0.17511 \t0.22954 \n\n\n```\n\nwe note that kidney and large intestine are correct.\n(although there are tiliing artifacts from the prediction mask you shown)\n\nyour prostate score is very low. i suspect you simply just tile the prostate test images. this is wrong. the hidden hubmap prostate images at at 6 um (train hubmap is 0.4 um).\nyou should rescale test (about 15 to 16x) prostate images before tiling.\nyou can probe the prostate test images. i think they are about 160x160 only \n\n\nyour local CV for spleen is low\nfrom you spleen test image, you have over predicted. there are too many FTUs,\nthis could be reason of your low score.\nthis is because spleen FTU are larger (and less distinctive) than that of kidney and largeintestine.\nthey probably require larger context (larger tile window) for accurate prediction.\n\nto prove this make a separate model for spleen using smaller resize image (while keeping your current tile window the same) for train and test\n\nfor lung, i think you use the same threshold as the rest of the organ. you can lower it by learning a new threshold\n\n---\nfurther,\n\nbut your model and mine are using efficientB7.\n\nyou are using tiling  and I am using whole image.\nit is obvious that for the earlier FTU (kidney and large intestine) your score is better becuase you have larger resolution. (and maybe ASPP of deeplab can capture more \"complex mask\" where as unet mask are \"smooth\", i.e. without sharp turns, curvy outlines, ...)\n\n\n---\n\nyet further,\n- kidney and large intestine are the easiest. most kagglers have differents core.\n- spleen is difficult. this differentiate the gold medals from the rest.\n- lung: a bit of \"by luck\" prediction. but nevertheless, the gold medal still produce better results than the rest.\n\n\n---\n\n\"lung and prostate hugely failed to be generalized to HUBMAP\". this statement is wrong.\n\nprostate is a pre-processing bug.\n\nlung and spleen is modeling bug. (becuase your performance is already low in your local hpa domain). it should be problem of scale and context\n\nto verify you may want to waste another 5 submissions to probe your test HPA score\n\n---\nmaybe best of both world - local and global context?\n- whole image prediction output some feature map\n- concat this feature map to input rgb, then perform tiling\n\nbut the annotation is already bad, so you maybe gain much after much efforts.\n\n---\n\nactually many kagglers have problems. but they don't provide enough information in the post. hence there is little we can help.\nthe first step in algorithm debug is to generate some results :)",
    "1893731": "best LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.\ni suppose local CV (HPA) is also about 0.80.\n\n(if you are getting 0.90, there is something wrong with your evaluation. note that evaluation on tiles is different from evaluation full image. you should always compute dice for the whole image after any post processing)\n\nSo there is really not so much domain difference as i have initially expected.\n(separate probe shows that LB-humap, LB-HPA are smiliar in public  test set)\n\nboth HPA and Human have a common stain: the blue color.\njust that one is blue+brown and other is blue+pink\n\n\n\n",
    "1893386": "I don't know what's weird. Domain adaptation is a very common problem in real world computer vision projects.",
    "1893338": "That is because the LB consists of HuBMAP + HPA while the train set only consists of HuBMAP, you can read more about it here: https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/333704",
    "1930585": "May I ask what kind of augmentations did you use?",
    "1901565": "\"Hubmap 51 , HPA 21, LB73 \"\n\nthis is reasonable score.\nHPA 32 can gives Hubmap 52 to 54 depends on encoder, decoder.\nyou can try different combination once you fixed your tiling and whole image bug. \n\n ",
    "1901555": "\"why on earth prostate local score decreased?\"\n\nthere are only a few prostate validation images. \nchanges in just one image or one FTU can causes large change in scores.\n\nthe label are not consistent (label noise). lower score can mean \"better results\".\nit is better that you inspect the results and find the reason for low score, e.g.\n\nhigh score = prediction missed a non labelled FTU.\nlow score = predicted detects the  non labelled FTU.",
    "1901032": "I am same to you when I began this game, and the lb-score is rising when I give the more epochs and mode parameters, try more SOTA models like @hengck23, I learn much  from his work",
    "1893705": "Try more augmentation, it could be better",
    "1893809": "I'm sorry guys for my comment that \"this competition is a joke\", it joust a little bit frustrating to put so much work in to get only 50%.\nI'm quite puzzled by my results.\n\n```\nbest LB score without external data /domain adaption/other tricks (HPA, Hubmap) is ~0.80.\ni suppose local CV (HPA) is also about 0.80.\n```\n\nI'm even more puzzled by this😄\nHow long did you train your model to get such a result?",
    "1927956": "How did you solve the issue in the end? I keep having CV 0.8 LB 0.6 even with augmentatios",
    "1894194": "I meet the same condition, @befunny Hi,wangkui ,I know you use the mmsegmentation in the competition ,i use it published too,i resize the whole img to (1536,1536) to train and test in (1536,1536)size,it is perform good in my CV(0.75),but i just got 0.52 in  my LB,I don't know why, did i  \nmiss some details ?can you help me?which really bothering me for a long time，i will really appreciate it!🙏",
    "1893359": ""
  }
}