{
  "id": 235941,
  "title": "Should there be a shake-up?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/235941",
  "author_name": "",
  "post_date": "2021-05-02T00:38:28.700142400Z",
  "votes": 11,
  "comment_count": 42,
  "views": 0,
  "content": "<p>I think there should be a shake-up in this competition.<br>\nThe reasons are as follows:</p>\n<ul>\n<li>Manually pseudo labelled image in public test set (public LB score should be higher)</li>\n<li>Deepflash 2 notebook can be overfitted on the private leaderboard (the public kernel doesn't use Kfold cross-validation)</li>\n<li>There are two kernels with high scores on the public leaderboard (public csv file can be submitted)</li>\n</ul>",
  "messages": [
    {
      "id": "1290426",
      "postDate": "05/02/2021 00:38:28",
      "content": "<p>I think there should be a shake-up in this competition.<br>\nThe reasons are as follows:</p>\n<ul>\n<li>Manually pseudo labelled image in public test set (public LB score should be higher)</li>\n<li>Deepflash 2 notebook can be overfitted on the private leaderboard (the public kernel doesn't use Kfold cross-validation)</li>\n<li>There are two kernels with high scores on the public leaderboard (public csv file can be submitted)</li>\n</ul>",
      "rawMarkdown": "I think there should be a shake-up in this competition.\nThe reasons are as follows:\n\n- Manually pseudo labelled image in public test set (public LB score should be higher)\n- Deepflash 2 notebook can be overfitted on the private leaderboard (the public kernel doesn't use Kfold cross-validation)\n- There are two kernels with high scores on the public leaderboard (public csv file can be submitted)",
      "votes": null
    },
    {
      "id": "1290461",
      "postDate": "05/02/2021 02:35:33",
      "content": "<p>Is the Deepflash 2 only gets lb 922?</p>",
      "rawMarkdown": "Is the Deepflash 2 only gets lb 922?",
      "votes": null
    },
    {
      "id": "1290466",
      "postDate": "05/02/2021 02:43:42",
      "content": "<p>I got 92.8 on public LB and CV is around 94.5 using pseudo labeling. You can also read this discussion and comments <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/234575\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/234575</a><br>\nThere is someone got more than 93.5 on public LB. </p>",
      "rawMarkdown": "I got 92.8 on public LB and CV is around 94.5 using pseudo labeling. You can also read this discussion and comments https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/234575\nThere is someone got more than 93.5 on public LB.",
      "votes": null
    },
    {
      "id": "1290528",
      "postDate": "05/02/2021 05:05:25",
      "content": "<p>Would u like to join us ? We use keras models.</p>",
      "rawMarkdown": "Would u like to join us ? We use keras models.",
      "votes": null
    },
    {
      "id": "1290534",
      "postDate": "05/02/2021 05:11:44",
      "content": "<p>Thank you for asking. I'm sorry I can't join your team. </p>",
      "rawMarkdown": "Thank you for asking. I'm sorry I can't join your team.",
      "votes": null
    },
    {
      "id": "1290696",
      "postDate": "05/02/2021 09:41:54",
      "content": "<p>There will be a ridiculous shakeup at the top because I suspect most people towards the top have used the D48 labels (which may or may not have been a good idea). Furthermore, anyone who's done extensive test set labelling will inevitably be very high at the moment.  Their scores will certainly diverge, either in a good way or a very bad way, from the others. It's going to be very exciting.</p>",
      "rawMarkdown": "There will be a ridiculous shakeup at the top because I suspect most people towards the top have used the D48 labels (which may or may not have been a good idea). Furthermore, anyone who's done extensive test set labelling will inevitably be very high at the moment.  Their scores will certainly diverge, either in a good way or a very bad way, from the others. It's going to be very exciting.",
      "votes": null
    },
    {
      "id": "1290698",
      "postDate": "05/02/2021 09:50:47",
      "content": "<p>I totally agree with you</p>",
      "rawMarkdown": "I totally agree with you",
      "votes": null
    },
    {
      "id": "1290717",
      "postDate": "05/02/2021 10:12:02",
      "content": "<p>Aha,have you done these?</p>",
      "rawMarkdown": "Aha,have you done these?",
      "votes": null
    },
    {
      "id": "1290721",
      "postDate": "05/02/2021 10:22:51",
      "content": "<p>I've used Zhao's D48 labels, but no further labelling.</p>",
      "rawMarkdown": "I've used Zhao's D48 labels, but no further labelling.",
      "votes": null
    },
    {
      "id": "1290785",
      "postDate": "05/02/2021 12:07:10",
      "content": "<p>This is the problem with pseudo labeling. When I inspect my model's output (0.919 LB), I can't find any glomeruli that it hasn't identified! In total I can find 15 cases where it has either missed or hasn't created a perfect mask. And 15 misses does not explain the 0.081 dice score miss from a full score. Given that in cross validation I've achieved dice scores as high as 0.96, it has come to labeling issues. So if the top LB scores are obtained not because they overfit, but because the pseudo/manual labels they used are better, then they could score higher in the private test too. And that just sucks.</p>",
      "rawMarkdown": "This is the problem with pseudo labeling. When I inspect my model's output (0.919 LB), I can't find any glomeruli that it hasn't identified! In total I can find 15 cases where it has either missed or hasn't created a perfect mask. And 15 misses does not explain the 0.081 dice score miss from a full score. Given that in cross validation I've achieved dice scores as high as 0.96, it has come to labeling issues. So if the top LB scores are obtained not because they overfit, but because the pseudo/manual labels they used are better, then they could score higher in the private test too. And that just sucks.",
      "votes": null
    },
    {
      "id": "1290999",
      "postDate": "05/02/2021 16:07:20",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> Question about the D48…How do you use it then?</p>\n<ol>\n<li>Do you use it to retrain a model?</li>\n<li>Or only use it to modify your final submission?</li>\n</ol>\n<p>I wonder how Kagglers put it to use where it would benefit both Public and Private LB. If it is only used to modify the submission file…then what's the use at all…beside looking good at the Public LB? It won't help for the Private LB in that case.</p>\n<p>In that case there could indeed be nice shakeup.</p>",
      "rawMarkdown": "Hey @jamesphoward Question about the D48...How do you use it then?\n\n1. Do you use it to retrain a model?\n2. Or only use it to modify your final submission?\n\nI wonder how Kagglers put it to use where it would benefit both Public and Private LB. If it is only used to modify the submission file...then what's the use at all...beside looking good at the Public LB? It won't help for the Private LB in that case.\n\nIn that case there could indeed be nice shakeup.",
      "votes": null
    },
    {
      "id": "1291032",
      "postDate": "05/02/2021 16:35:13",
      "content": "<p>I did (1).</p>\n<p>I think (2) is incredibly dangerous (although (1) might be, too) if you do it in a simple way, at least, because unless you've trained on D48, your model will make very unconfident predictions on those types of glomeruli. If you then use the D48 labels to calibrate your predictions, then you'll end up with an incredibly low prediction threshold for all other slides.</p>",
      "rawMarkdown": "I did (1).\n\nI think (2) is incredibly dangerous (although (1) might be, too) if you do it in a simple way, at least, because unless you've trained on D48, your model will make very unconfident predictions on those types of glomeruli. If you then use the D48 labels to calibrate your predictions, then you'll end up with an incredibly low prediction threshold for all other slides.",
      "votes": null
    },
    {
      "id": "1291062",
      "postDate": "05/02/2021 17:08:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> Thanks for your answer. <br>\nInterresting choice to make  there …. i you use it to train a model then there is not that much difference I suppose between automatically pseudo labelled data or manually labelled data.</p>\n<p>Given a proper validation strategy you should be able to get a fair indication I guess of whether it is 'safe' or not.</p>\n<p>That said….the final week and finish will be interresting for sure :-)</p>",
      "rawMarkdown": "Hi @jamesphoward Thanks for your answer. \nInterresting choice to make  there .... i you use it to train a model then there is not that much difference I suppose between automatically pseudo labelled data or manually labelled data.\n\nGiven a proper validation strategy you should be able to get a fair indication I guess of whether it is 'safe' or not.\n\nThat said....the final week and finish will be interresting for sure :-)",
      "votes": null
    },
    {
      "id": "1292140",
      "postDate": "05/03/2021 17:40:15",
      "content": "<p><a href=\"https://www.kaggle.com/lhagiimn\" target=\"_blank\">@lhagiimn</a>  Deepflash use randomized tiles for training and another set of tiles for validation not mixing with train  as such. SO just trying to understand reason for overfit.<br>\nkfold also does similar the scheme that is based not on Tiff ids </p>",
      "rawMarkdown": "lhagiimn  Deepflash use randomized tiles for training and another set of tiles for validation not mixing with train  as such. SO just trying to understand reason for overfit.\nkfold also does similar the scheme that is based not on Tiff ids",
      "votes": null
    },
    {
      "id": "1292148",
      "postDate": "05/03/2021 17:43:28",
      "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> u have not used PL of other test tiff ids ?</p>",
      "rawMarkdown": "jamesphoward u have not used PL of other test tiff ids ?",
      "votes": null
    },
    {
      "id": "1292182",
      "postDate": "05/03/2021 18:08:04",
      "content": "<p>Nope I haven't</p>",
      "rawMarkdown": "Nope I haven't",
      "votes": null
    },
    {
      "id": "1292204",
      "postDate": "05/03/2021 18:32:26",
      "content": "<p>Did u find a big diff in LB score  between with and without D4 label training(  meaning setting just low th during inference no training on using labels)</p>",
      "rawMarkdown": "Did u find a big diff in LB score  between with and without D4 label training(  meaning setting just low th during inference no training on using labels)",
      "votes": null
    },
    {
      "id": "1292205",
      "postDate": "05/03/2021 18:34:54",
      "content": "<p>I think my best model I moved from 0.926 to 0.934 by adding in D4 labels.</p>",
      "rawMarkdown": "I think my best model I moved from 0.926 to 0.934 by adding in D4 labels.",
      "votes": null
    },
    {
      "id": "1292264",
      "postDate": "05/03/2021 19:40:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> that's an interresting improvement. So the remainder of your current score is based on TTA and ensemble likely.</p>\n<p>With 2 submissions that can be selected I think I should at least try it out…given that my current single fold best is over 0.931..so would be interresting to see what the D4 labels would do in that case. Combining in an ensemble some models trained with D4 and without D4 should keep it on the safe side. Interresting…</p>\n<p>Good luck with the remainder of the competition!</p>",
      "rawMarkdown": "Hi @jamesphoward that's an interresting improvement. So the remainder of your current score is based on TTA and ensemble likely.\n\nWith 2 submissions that can be selected I think I should at least try it out...given that my current single fold best is over 0.931..so would be interresting to see what the D4 labels would do in that case. Combining in an ensemble some models trained with D4 and without D4 should keep it on the safe side. Interresting...\n\nGood luck with the remainder of the competition!",
      "votes": null
    },
    {
      "id": "1292309",
      "postDate": "05/03/2021 20:18:32",
      "content": "<p>I certainly think it's wise to have a 2nd submission pipeline <em>without</em> D4, given the controversy around the glomeruli it contains.</p>",
      "rawMarkdown": "I certainly think it's wise to have a 2nd submission pipeline _without_ D4, given the controversy around the glomeruli it contains.",
      "votes": null
    },
    {
      "id": "1292583",
      "postDate": "05/04/2021 05:53:00",
      "content": "<p>There is no separation of which images are used for train or validation, both use all the files in deepflash public notebooks. Train does random tiles considering pdfs &gt; random.random [0.0-1.0] to choose centers.  Validation randomly samples the designated number of tiles.  There is no guarantee of no overlap of tiles or sections of tiles.  The use of random in various places tend to make results hard to reproduce even with seeds.  </p>\n<p>Using stratified kfold e.g.  based on mask coverage for tiles like from the iafoss datasets, or hold out of different ids for validation sampling, would ensure folds have different tiles (or images) for train or validation.</p>",
      "rawMarkdown": "There is no separation of which images are used for train or validation, both use all the files in deepflash public notebooks. Train does random tiles considering pdfs > random.random [0.0-1.0] to choose centers.  Validation randomly samples the designated number of tiles.  There is no guarantee of no overlap of tiles or sections of tiles.  The use of random in various places tend to make results hard to reproduce even with seeds.  \n\nUsing stratified kfold e.g.  based on mask coverage for tiles like from the iafoss datasets, or hold out of different ids for validation sampling, would ensure folds have different tiles (or images) for train or validation.",
      "votes": null
    },
    {
      "id": "1292602",
      "postDate": "05/04/2021 06:26:31",
      "content": "<p>But when I use kfold,the lb decrease.</p>",
      "rawMarkdown": "But when I use kfold,the lb decrease.",
      "votes": null
    },
    {
      "id": "1292643",
      "postDate": "05/04/2021 07:01:02",
      "content": "<p>Hello,is the whole test set 5 images?</p>",
      "rawMarkdown": "Hello,is the whole test set 5 images?",
      "votes": null
    },
    {
      "id": "1292654",
      "postDate": "05/04/2021 07:17:03",
      "content": "<p>No, it should be 15 images. </p>",
      "rawMarkdown": "No, it should be 15 images.",
      "votes": null
    },
    {
      "id": "1292656",
      "postDate": "05/04/2021 07:23:40",
      "content": "<p>Yes, your right. I just said that it could be overfitted. So it couldn’t be overfitted as well. <br>\nMy opinion is that training and validation sets can be come from same image for current public notebook because validation set is random sample tiles. So trained model could easily be overfitted. The test set is different from training set, I think GroupKFold could be good evaluation strategy. </p>",
      "rawMarkdown": "Yes, your right. I just said that it could be overfitted. So it couldn’t be overfitted as well. \nMy opinion is that training and validation sets can be come from same image for current public notebook because validation set is random sample tiles. So trained model could easily be overfitted. The test set is different from training set, I think GroupKFold could be good evaluation strategy.",
      "votes": null
    },
    {
      "id": "1292785",
      "postDate": "05/04/2021 09:24:33",
      "content": "<p>I tried  kfold split based on tiff ids , gkf results were uncorrelated completely. Model possibly overfits to position of ftu given all tiles of same tiff in one set </p>",
      "rawMarkdown": "I tried  kfold split based on tiff ids , gkf results were uncorrelated completely. Model possibly overfits to position of ftu given all tiles of same tiff in one set",
      "votes": null
    },
    {
      "id": "1293769",
      "postDate": "05/05/2021 06:42:13",
      "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a>  could u help me interpret 7.26 to 7.34. Assuming your base score .900  .so improvement is 908 ?</p>",
      "rawMarkdown": "jamesphoward  could u help me interpret 7.26 to 7.34. Assuming your base score .900  .so improvement is 908 ?",
      "votes": null
    },
    {
      "id": "1293792",
      "postDate": "05/05/2021 07:03:49",
      "content": "<p>Sorry that was a bizarre typo - I meant to say 0.926 to 0.934.</p>",
      "rawMarkdown": "Sorry that was a bizarre typo - I meant to say 0.926 to 0.934.",
      "votes": null
    },
    {
      "id": "1293823",
      "postDate": "05/05/2021 07:28:38",
      "content": "<blockquote>\n  <p>Sorry that was a bizarre typo - I meant to say 0.924 to 0.936.</p>\n</blockquote>\n<p>do you mean 0.926 to 0.934 ;)</p>",
      "rawMarkdown": "> Sorry that was a bizarre typo - I meant to say 0.924 to 0.936.\n\ndo you mean 0.926 to 0.934 ;)",
      "votes": null
    },
    {
      "id": "1293828",
      "postDate": "05/05/2021 07:33:32",
      "content": "<p>😳 yes, sorry (again)</p>",
      "rawMarkdown": "😳 yes, sorry (again)",
      "votes": null
    },
    {
      "id": "1295240",
      "postDate": "05/06/2021 10:08:06",
      "content": "<p>The final shakeup probably wont be too different from this<br>\n<img src=\"https://i.ibb.co/Xz0q1Q7/Screenshot-2021-05-06-at-11-04-31.png\" alt=\"The final shakeup probably wont be too different from this\"></p>",
      "rawMarkdown": "The final shakeup probably wont be too different from this\n![The final shakeup probably wont be too different from this](https://i.ibb.co/Xz0q1Q7/Screenshot-2021-05-06-at-11-04-31.png)",
      "votes": null
    },
    {
      "id": "1295303",
      "postDate": "05/06/2021 11:06:16",
      "content": "<p>I think so. I think gold zone would be safe.</p>",
      "rawMarkdown": "I think so. I think gold zone would be safe.",
      "votes": null
    },
    {
      "id": "1296082",
      "postDate": "05/07/2021 00:48:13",
      "content": "<p>Nooooo! It's sort <strong>ascending</strong> 😏</p>",
      "rawMarkdown": "Nooooo! It's sort **ascending** 😏",
      "votes": null
    },
    {
      "id": "1296628",
      "postDate": "05/07/2021 11:41:41",
      "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> I trained my models with hand labelled D4 tfrecord (using Zhao's dataset) but instead of giving me a boost on cv and lb it dropped the score from 0.920 to 0.744 on public lb. Also, another thing I noticed was my models started facing exploding gradients problem on appending the D4 tfrecord to the train set (which was fixed using clipnorm) but still the result was unexpected. Is there something wrong with the generation of D4 tfrecords ? or is that the models fails to train and learn on completely different data (D4) compared to the initial training set ?</p>",
      "rawMarkdown": "jamesphoward I trained my models with hand labelled D4 tfrecord (using Zhao's dataset) but instead of giving me a boost on cv and lb it dropped the score from 0.920 to 0.744 on public lb. Also, another thing I noticed was my models started facing exploding gradients problem on appending the D4 tfrecord to the train set (which was fixed using clipnorm) but still the result was unexpected. Is there something wrong with the generation of D4 tfrecords ? or is that the models fails to train and learn on completely different data (D4) compared to the initial training set ?",
      "votes": null
    },
    {
      "id": "1296641",
      "postDate": "05/07/2021 11:51:39",
      "content": "<p>Strange. I don't use tfrecords so I can't say, but I found my efficientnet linknet models trained almost identically when the D4 data was added in.</p>",
      "rawMarkdown": "Strange. I don't use tfrecords so I can't say, but I found my efficientnet linknet models trained almost identically when the D4 data was added in.",
      "votes": null
    },
    {
      "id": "1297173",
      "postDate": "05/07/2021 19:09:05",
      "content": "<p>I would prefer no shake-up 😉<br>\nBut I must admit that probability for strong shake-up is not null. </p>",
      "rawMarkdown": "I would prefer no shake-up 😉\nBut I must admit that probability for strong shake-up is not null.",
      "votes": null
    },
    {
      "id": "1297222",
      "postDate": "05/07/2021 20:02:26",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> did you guys use the infamous <strong>D48</strong> ? 😜</p>",
      "rawMarkdown": "mpware did you guys use the infamous **D48** ? 😜",
      "votes": null
    },
    {
      "id": "1297699",
      "postDate": "05/08/2021 08:28:01",
      "content": "<p>Zhao's label is not the key. By probing LB i noticed that my CV prediction of d4 is more accurate than Zhao's hand-label</p>",
      "rawMarkdown": "Zhao's label is not the key. By probing LB i noticed that my CV prediction of d4 is more accurate than Zhao's hand-label",
      "votes": null
    },
    {
      "id": "1297700",
      "postDate": "05/08/2021 08:29:54",
      "content": "<p>I would rather trust CV than LB. believe you me. I lost two golds medals 😨</p>",
      "rawMarkdown": "I would rather trust CV than LB. believe you me. I lost two golds medals 😨",
      "votes": null
    },
    {
      "id": "1297701",
      "postDate": "05/08/2021 08:30:24",
      "content": "<p>my d4 is more accurate than Zhao's too but I dont think it is useful for my model. I am still afraid to use that.</p>",
      "rawMarkdown": "my d4 is more accurate than Zhao's too but I dont think it is useful for my model. I am still afraid to use that.",
      "votes": null
    },
    {
      "id": "1297708",
      "postDate": "05/08/2021 08:40:50",
      "content": "<p>Nice to see that you came to the same conclusion</p>",
      "rawMarkdown": "Nice to see that you came to the same conclusion",
      "votes": null
    },
    {
      "id": "1301244",
      "postDate": "05/11/2021 01:55:12",
      "content": "<p>It happened :(</p>",
      "rawMarkdown": "It happened :(",
      "votes": null
    },
    {
      "id": "1301580",
      "postDate": "05/11/2021 06:37:26",
      "content": "<p>Well that aged like a fine wine.</p>",
      "rawMarkdown": "Well that aged like a fine wine.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1290461,
      "author_name": "zekunn",
      "author_url": "",
      "post_date": "05/02/2021 02:35:33",
      "content": "<p>Is the Deepflash 2 only gets lb 922?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1290466,
          "author_name": "lhagiimn",
          "author_url": "",
          "post_date": "05/02/2021 02:43:42",
          "content": "<p>I got 92.8 on public LB and CV is around 94.5 using pseudo labeling. You can also read this discussion and comments <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/234575\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/234575</a><br>\nThere is someone got more than 93.5 on public LB. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1290528,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "05/02/2021 05:05:25",
          "content": "<p>Would u like to join us ? We use keras models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1290534,
          "author_name": "lhagiimn",
          "author_url": "",
          "post_date": "05/02/2021 05:11:44",
          "content": "<p>Thank you for asking. I'm sorry I can't join your team. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1290696,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "05/02/2021 09:41:54",
      "content": "<p>There will be a ridiculous shakeup at the top because I suspect most people towards the top have used the D48 labels (which may or may not have been a good idea). Furthermore, anyone who's done extensive test set labelling will inevitably be very high at the moment.  Their scores will certainly diverge, either in a good way or a very bad way, from the others. It's going to be very exciting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1290698,
          "author_name": "lhagiimn",
          "author_url": "",
          "post_date": "05/02/2021 09:50:47",
          "content": "<p>I totally agree with you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1290717,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "05/02/2021 10:12:02",
          "content": "<p>Aha,have you done these?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1290721,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/02/2021 10:22:51",
          "content": "<p>I've used Zhao's D48 labels, but no further labelling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1290785,
          "author_name": "erikdali",
          "author_url": "",
          "post_date": "05/02/2021 12:07:10",
          "content": "<p>This is the problem with pseudo labeling. When I inspect my model's output (0.919 LB), I can't find any glomeruli that it hasn't identified! In total I can find 15 cases where it has either missed or hasn't created a perfect mask. And 15 misses does not explain the 0.081 dice score miss from a full score. Given that in cross validation I've achieved dice scores as high as 0.96, it has come to labeling issues. So if the top LB scores are obtained not because they overfit, but because the pseudo/manual labels they used are better, then they could score higher in the private test too. And that just sucks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1290999,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "05/02/2021 16:07:20",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> Question about the D48…How do you use it then?</p>\n<ol>\n<li>Do you use it to retrain a model?</li>\n<li>Or only use it to modify your final submission?</li>\n</ol>\n<p>I wonder how Kagglers put it to use where it would benefit both Public and Private LB. If it is only used to modify the submission file…then what's the use at all…beside looking good at the Public LB? It won't help for the Private LB in that case.</p>\n<p>In that case there could indeed be nice shakeup.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1291032,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/02/2021 16:35:13",
          "content": "<p>I did (1).</p>\n<p>I think (2) is incredibly dangerous (although (1) might be, too) if you do it in a simple way, at least, because unless you've trained on D48, your model will make very unconfident predictions on those types of glomeruli. If you then use the D48 labels to calibrate your predictions, then you'll end up with an incredibly low prediction threshold for all other slides.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1291062,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "05/02/2021 17:08:55",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> Thanks for your answer. <br>\nInterresting choice to make  there …. i you use it to train a model then there is not that much difference I suppose between automatically pseudo labelled data or manually labelled data.</p>\n<p>Given a proper validation strategy you should be able to get a fair indication I guess of whether it is 'safe' or not.</p>\n<p>That said….the final week and finish will be interresting for sure :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292148,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "05/03/2021 17:43:28",
          "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> u have not used PL of other test tiff ids ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292182,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/03/2021 18:08:04",
          "content": "<p>Nope I haven't</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292204,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "05/03/2021 18:32:26",
          "content": "<p>Did u find a big diff in LB score  between with and without D4 label training(  meaning setting just low th during inference no training on using labels)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292205,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/03/2021 18:34:54",
          "content": "<p>I think my best model I moved from 0.926 to 0.934 by adding in D4 labels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292264,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "05/03/2021 19:40:37",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> that's an interresting improvement. So the remainder of your current score is based on TTA and ensemble likely.</p>\n<p>With 2 submissions that can be selected I think I should at least try it out…given that my current single fold best is over 0.931..so would be interresting to see what the D4 labels would do in that case. Combining in an ensemble some models trained with D4 and without D4 should keep it on the safe side. Interresting…</p>\n<p>Good luck with the remainder of the competition!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292309,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/03/2021 20:18:32",
          "content": "<p>I certainly think it's wise to have a 2nd submission pipeline <em>without</em> D4, given the controversy around the glomeruli it contains.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292643,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "05/04/2021 07:01:02",
          "content": "<p>Hello,is the whole test set 5 images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292654,
          "author_name": "lhagiimn",
          "author_url": "",
          "post_date": "05/04/2021 07:17:03",
          "content": "<p>No, it should be 15 images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293769,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "05/05/2021 06:42:13",
          "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a>  could u help me interpret 7.26 to 7.34. Assuming your base score .900  .so improvement is 908 ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293792,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/05/2021 07:03:49",
          "content": "<p>Sorry that was a bizarre typo - I meant to say 0.926 to 0.934.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293823,
          "author_name": "fabiendaniel",
          "author_url": "",
          "post_date": "05/05/2021 07:28:38",
          "content": "<blockquote>\n  <p>Sorry that was a bizarre typo - I meant to say 0.924 to 0.936.</p>\n</blockquote>\n<p>do you mean 0.926 to 0.934 ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1293828,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/05/2021 07:33:32",
          "content": "<p>😳 yes, sorry (again)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1296628,
          "author_name": "thakurudit",
          "author_url": "",
          "post_date": "05/07/2021 11:41:41",
          "content": "<p><a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a> I trained my models with hand labelled D4 tfrecord (using Zhao's dataset) but instead of giving me a boost on cv and lb it dropped the score from 0.920 to 0.744 on public lb. Also, another thing I noticed was my models started facing exploding gradients problem on appending the D4 tfrecord to the train set (which was fixed using clipnorm) but still the result was unexpected. Is there something wrong with the generation of D4 tfrecords ? or is that the models fails to train and learn on completely different data (D4) compared to the initial training set ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1296641,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/07/2021 11:51:39",
          "content": "<p>Strange. I don't use tfrecords so I can't say, but I found my efficientnet linknet models trained almost identically when the D4 data was added in.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297699,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "05/08/2021 08:28:01",
          "content": "<p>Zhao's label is not the key. By probing LB i noticed that my CV prediction of d4 is more accurate than Zhao's hand-label</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297701,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "05/08/2021 08:30:24",
          "content": "<p>my d4 is more accurate than Zhao's too but I dont think it is useful for my model. I am still afraid to use that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297708,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "05/08/2021 08:40:50",
          "content": "<p>Nice to see that you came to the same conclusion</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1292140,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "05/03/2021 17:40:15",
      "content": "<p><a href=\"https://www.kaggle.com/lhagiimn\" target=\"_blank\">@lhagiimn</a>  Deepflash use randomized tiles for training and another set of tiles for validation not mixing with train  as such. SO just trying to understand reason for overfit.<br>\nkfold also does similar the scheme that is based not on Tiff ids </p>",
      "votes": null,
      "replies": [
        {
          "id": 1292583,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "05/04/2021 05:53:00",
          "content": "<p>There is no separation of which images are used for train or validation, both use all the files in deepflash public notebooks. Train does random tiles considering pdfs &gt; random.random [0.0-1.0] to choose centers.  Validation randomly samples the designated number of tiles.  There is no guarantee of no overlap of tiles or sections of tiles.  The use of random in various places tend to make results hard to reproduce even with seeds.  </p>\n<p>Using stratified kfold e.g.  based on mask coverage for tiles like from the iafoss datasets, or hold out of different ids for validation sampling, would ensure folds have different tiles (or images) for train or validation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292602,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "05/04/2021 06:26:31",
          "content": "<p>But when I use kfold,the lb decrease.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292656,
          "author_name": "lhagiimn",
          "author_url": "",
          "post_date": "05/04/2021 07:23:40",
          "content": "<p>Yes, your right. I just said that it could be overfitted. So it couldn’t be overfitted as well. <br>\nMy opinion is that training and validation sets can be come from same image for current public notebook because validation set is random sample tiles. So trained model could easily be overfitted. The test set is different from training set, I think GroupKFold could be good evaluation strategy. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292785,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "05/04/2021 09:24:33",
          "content": "<p>I tried  kfold split based on tiff ids , gkf results were uncorrelated completely. Model possibly overfits to position of ftu given all tiles of same tiff in one set </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297700,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "05/08/2021 08:29:54",
          "content": "<p>I would rather trust CV than LB. believe you me. I lost two golds medals 😨</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1295240,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "05/06/2021 10:08:06",
      "content": "<p>The final shakeup probably wont be too different from this<br>\n<img src=\"https://i.ibb.co/Xz0q1Q7/Screenshot-2021-05-06-at-11-04-31.png\" alt=\"The final shakeup probably wont be too different from this\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1295303,
          "author_name": "lhagiimn",
          "author_url": "",
          "post_date": "05/06/2021 11:06:16",
          "content": "<p>I think so. I think gold zone would be safe.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1296082,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/07/2021 00:48:13",
          "content": "<p>Nooooo! It's sort <strong>ascending</strong> 😏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297173,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "05/07/2021 19:09:05",
          "content": "<p>I would prefer no shake-up 😉<br>\nBut I must admit that probability for strong shake-up is not null. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297222,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "05/07/2021 20:02:26",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> did you guys use the infamous <strong>D48</strong> ? 😜</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301580,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "05/11/2021 06:37:26",
          "content": "<p>Well that aged like a fine wine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301244,
      "author_name": "lhagiimn",
      "author_url": "",
      "post_date": "05/11/2021 01:55:12",
      "content": "<p>It happened :(</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1290426": "I think there should be a shake-up in this competition.\nThe reasons are as follows:\n\n- Manually pseudo labelled image in public test set (public LB score should be higher)\n- Deepflash 2 notebook can be overfitted on the private leaderboard (the public kernel doesn't use Kfold cross-validation)\n- There are two kernels with high scores on the public leaderboard (public csv file can be submitted)",
    "1290461": "Is the Deepflash 2 only gets lb 922?",
    "1290466": "I got 92.8 on public LB and CV is around 94.5 using pseudo labeling. You can also read this discussion and comments https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/234575\nThere is someone got more than 93.5 on public LB.",
    "1290528": "Would u like to join us ? We use keras models.",
    "1290534": "Thank you for asking. I'm sorry I can't join your team.",
    "1290696": "There will be a ridiculous shakeup at the top because I suspect most people towards the top have used the D48 labels (which may or may not have been a good idea). Furthermore, anyone who's done extensive test set labelling will inevitably be very high at the moment.  Their scores will certainly diverge, either in a good way or a very bad way, from the others. It's going to be very exciting.",
    "1290698": "I totally agree with you",
    "1290717": "Aha,have you done these?",
    "1290721": "I've used Zhao's D48 labels, but no further labelling.",
    "1290785": "This is the problem with pseudo labeling. When I inspect my model's output (0.919 LB), I can't find any glomeruli that it hasn't identified! In total I can find 15 cases where it has either missed or hasn't created a perfect mask. And 15 misses does not explain the 0.081 dice score miss from a full score. Given that in cross validation I've achieved dice scores as high as 0.96, it has come to labeling issues. So if the top LB scores are obtained not because they overfit, but because the pseudo/manual labels they used are better, then they could score higher in the private test too. And that just sucks.",
    "1290999": "Hey @jamesphoward Question about the D48...How do you use it then?\n\n1. Do you use it to retrain a model?\n2. Or only use it to modify your final submission?\n\nI wonder how Kagglers put it to use where it would benefit both Public and Private LB. If it is only used to modify the submission file...then what's the use at all...beside looking good at the Public LB? It won't help for the Private LB in that case.\n\nIn that case there could indeed be nice shakeup.",
    "1291032": "I did (1).\n\nI think (2) is incredibly dangerous (although (1) might be, too) if you do it in a simple way, at least, because unless you've trained on D48, your model will make very unconfident predictions on those types of glomeruli. If you then use the D48 labels to calibrate your predictions, then you'll end up with an incredibly low prediction threshold for all other slides.",
    "1291062": "Hi @jamesphoward Thanks for your answer. \nInterresting choice to make  there .... i you use it to train a model then there is not that much difference I suppose between automatically pseudo labelled data or manually labelled data.\n\nGiven a proper validation strategy you should be able to get a fair indication I guess of whether it is 'safe' or not.\n\nThat said....the final week and finish will be interresting for sure :-)",
    "1292140": "lhagiimn  Deepflash use randomized tiles for training and another set of tiles for validation not mixing with train  as such. SO just trying to understand reason for overfit.\nkfold also does similar the scheme that is based not on Tiff ids",
    "1292148": "jamesphoward u have not used PL of other test tiff ids ?",
    "1292182": "Nope I haven't",
    "1292204": "Did u find a big diff in LB score  between with and without D4 label training(  meaning setting just low th during inference no training on using labels)",
    "1292205": "I think my best model I moved from 0.926 to 0.934 by adding in D4 labels.",
    "1292264": "Hi @jamesphoward that's an interresting improvement. So the remainder of your current score is based on TTA and ensemble likely.\n\nWith 2 submissions that can be selected I think I should at least try it out...given that my current single fold best is over 0.931..so would be interresting to see what the D4 labels would do in that case. Combining in an ensemble some models trained with D4 and without D4 should keep it on the safe side. Interresting...\n\nGood luck with the remainder of the competition!",
    "1292309": "I certainly think it's wise to have a 2nd submission pipeline _without_ D4, given the controversy around the glomeruli it contains.",
    "1292583": "There is no separation of which images are used for train or validation, both use all the files in deepflash public notebooks. Train does random tiles considering pdfs > random.random [0.0-1.0] to choose centers.  Validation randomly samples the designated number of tiles.  There is no guarantee of no overlap of tiles or sections of tiles.  The use of random in various places tend to make results hard to reproduce even with seeds.  \n\nUsing stratified kfold e.g.  based on mask coverage for tiles like from the iafoss datasets, or hold out of different ids for validation sampling, would ensure folds have different tiles (or images) for train or validation.",
    "1292602": "But when I use kfold,the lb decrease.",
    "1292643": "Hello,is the whole test set 5 images?",
    "1292654": "No, it should be 15 images.",
    "1292656": "Yes, your right. I just said that it could be overfitted. So it couldn’t be overfitted as well. \nMy opinion is that training and validation sets can be come from same image for current public notebook because validation set is random sample tiles. So trained model could easily be overfitted. The test set is different from training set, I think GroupKFold could be good evaluation strategy.",
    "1292785": "I tried  kfold split based on tiff ids , gkf results were uncorrelated completely. Model possibly overfits to position of ftu given all tiles of same tiff in one set",
    "1293769": "jamesphoward  could u help me interpret 7.26 to 7.34. Assuming your base score .900  .so improvement is 908 ?",
    "1293792": "Sorry that was a bizarre typo - I meant to say 0.926 to 0.934.",
    "1293823": "> Sorry that was a bizarre typo - I meant to say 0.924 to 0.936.\n\ndo you mean 0.926 to 0.934 ;)",
    "1293828": "😳 yes, sorry (again)",
    "1295240": "The final shakeup probably wont be too different from this\n![The final shakeup probably wont be too different from this](https://i.ibb.co/Xz0q1Q7/Screenshot-2021-05-06-at-11-04-31.png)",
    "1295303": "I think so. I think gold zone would be safe.",
    "1296082": "Nooooo! It's sort **ascending** 😏",
    "1296628": "jamesphoward I trained my models with hand labelled D4 tfrecord (using Zhao's dataset) but instead of giving me a boost on cv and lb it dropped the score from 0.920 to 0.744 on public lb. Also, another thing I noticed was my models started facing exploding gradients problem on appending the D4 tfrecord to the train set (which was fixed using clipnorm) but still the result was unexpected. Is there something wrong with the generation of D4 tfrecords ? or is that the models fails to train and learn on completely different data (D4) compared to the initial training set ?",
    "1296641": "Strange. I don't use tfrecords so I can't say, but I found my efficientnet linknet models trained almost identically when the D4 data was added in.",
    "1297173": "I would prefer no shake-up 😉\nBut I must admit that probability for strong shake-up is not null.",
    "1297222": "mpware did you guys use the infamous **D48** ? 😜",
    "1297699": "Zhao's label is not the key. By probing LB i noticed that my CV prediction of d4 is more accurate than Zhao's hand-label",
    "1297700": "I would rather trust CV than LB. believe you me. I lost two golds medals 😨",
    "1297701": "my d4 is more accurate than Zhao's too but I dont think it is useful for my model. I am still afraid to use that.",
    "1297708": "Nice to see that you came to the same conclusion",
    "1301244": "It happened :(",
    "1301580": "Well that aged like a fine wine."
  },
  "source": "meta"
}