{
  "id": 221808,
  "title": "Duplicated images from RANZCR and CHESTX",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/221808",
  "author_name": "Mohamed abdelrazik",
  "post_date": "2021-02-24T05:50:54.020000",
  "votes": 48,
  "comment_count": 67,
  "views": 0,
  "content": "<p>i see many people suggest to use additional data set and many found chestx data set the most properly one and thanks to <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> who use chestx dataset to find good starting point but i was so curious to figure out if there is any overlapped data between chestx and RANZCR original data set and the answer is yes this idea shine in my head after this <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221159\" target=\"_blank\">post</a> thank you  <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> i use imagehash library to find dublicated images and found about 29000 duplicated <br>\n i uploaded 6 examples <a href=\"https://www.kaggle.com/mohamed3abdelrazik/4-examples-for-duplication\" target=\"_blank\">Here</a><br>\nand a csv file contains image paths for both original and dublicated images cn be found <a href=\"https://www.kaggle.com/mohamed3abdelrazik/paths-for-duplicated-images-on-chestx-and-ranczr/settings\" target=\"_blank\">Here</a></p>",
  "messages": [
    {
      "id": 1216107,
      "postDate": "2021-02-24T05:50:54.020Z",
      "content": "<p>i see many people suggest to use additional data set and many found chestx data set the most properly one and thanks to <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> who use chestx dataset to find good starting point but i was so curious to figure out if there is any overlapped data between chestx and RANZCR original data set and the answer is yes this idea shine in my head after this <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221159\" target=\"_blank\">post</a> thank you  <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> i use imagehash library to find dublicated images and found about 29000 duplicated <br>\n i uploaded 6 examples <a href=\"https://www.kaggle.com/mohamed3abdelrazik/4-examples-for-duplication\" target=\"_blank\">Here</a><br>\nand a csv file contains image paths for both original and dublicated images cn be found <a href=\"https://www.kaggle.com/mohamed3abdelrazik/paths-for-duplicated-images-on-chestx-and-ranczr/settings\" target=\"_blank\">Here</a></p>",
      "rawMarkdown": "i see many people suggest to use additional data set and many found chestx data set the most properly one and thanks to @ammarali32 who use chestx dataset to find good starting point but i was so curious to figure out if there is any overlapped data between chestx and RANZCR original data set and the answer is yes this idea shine in my head after this [post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221159) thank you  @virilo i use imagehash library to find dublicated images and found about 29000 duplicated \n i uploaded 6 examples [Here](https://www.kaggle.com/mohamed3abdelrazik/4-examples-for-duplication)\nand a csv file contains image paths for both original and dublicated images cn be found [Here](https://www.kaggle.com/mohamed3abdelrazik/paths-for-duplicated-images-on-chestx-and-ranczr/settings)",
      "votes": 48
    },
    {
      "id": 1216234,
      "postDate": "2021-02-24T06:46:39.733Z",
      "content": "<p><a href=\"https://www.kaggle.com/mohamed3abdelrazik\" target=\"_blank\">@mohamed3abdelrazik</a> As per the acknowledgements page - this dataset was created by relabelling the publicly available CXR14 dataset from NIH, therefore there will be 100% overlap in image data. These labels (concerning lines and tubes placement) are not part of the original dataset.</p>",
      "rawMarkdown": "@mohamed3abdelrazik As per the acknowledgements page - this dataset was created by relabelling the publicly available CXR14 dataset from NIH, therefore there will be 100% overlap in image data. These labels (concerning lines and tubes placement) are not part of the original dataset.",
      "votes": 19,
      "replies": [
        {
          "id": 1216258,
          "postDate": "2021-02-24T07:00:27.713Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1216265,
          "postDate": "2021-02-24T07:08:39.063Z",
          "content": "<p>Thanks a lot for pointing to that! I joined late and must have missed this. It is clear now.</p>",
          "rawMarkdown": "Thanks a lot for pointing to that! I joined late and must have missed this. It is clear now."
        },
        {
          "id": 1216333,
          "postDate": "2021-02-24T08:05:16.860Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        },
        {
          "id": 1216343,
          "postDate": "2021-02-24T08:11:33.853Z",
          "rawMarkdown": "",
          "votes": -3,
          "isDeleted": true
        },
        {
          "id": 1216607,
          "postDate": "2021-02-24T11:10:04.373Z",
          "content": "<p>\"this dataset\" means the \"training set\" or  \"training set + testing set\" of the competition? </p>",
          "rawMarkdown": "\"this dataset\" means the \"training set\" or  \"training set + testing set\" of the competition? ",
          "votes": 1
        },
        {
          "id": 1216628,
          "postDate": "2021-02-24T11:36:19.937Z",
          "content": "<p><a href=\"https://www.kaggle.com/dldmw579\" target=\"_blank\">@dldmw579</a> my understanding is that it means \"training set + public test + private test\", i.e. the entire competition data is a subset of CXR14. But it would be nice to verify this to be sure.</p>",
          "rawMarkdown": "@dldmw579 my understanding is that it means \"training set + public test + private test\", i.e. the entire competition data is a subset of CXR14. But it would be nice to verify this to be sure."
        },
        {
          "id": 1217492,
          "postDate": "2021-02-25T05:54:43.070Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 1218391,
          "postDate": "2021-02-25T19:51:26.543Z",
          "content": "<p>good to know</p>",
          "rawMarkdown": "good to know"
        }
      ]
    },
    {
      "id": 1216262,
      "postDate": "2021-02-24T07:03:40.497Z",
      "content": "<p>\" imagehash library to find dublicated images and found about 29000 duplicated\"</p>\n<p>self-supervised learning is the key?<br>\nthe private data has turned from \"black box\" to \"white box\"</p>\n<p><strong>kaggle need to issue a rule that hand labeling of external data is allowed or not</strong></p>",
      "rawMarkdown": "\" imagehash library to find dublicated images and found about 29000 duplicated\"\n\nself-supervised learning is the key?\nthe private data has turned from \"black box\" to \"white box\"\n\n**kaggle need to issue a rule that hand labeling of external data is allowed or not**",
      "votes": 8,
      "replies": [
        {
          "id": 1216350,
          "postDate": "2021-02-24T08:15:11.227Z",
          "content": "<p>I didn't use imagehash to label images i just use to figure out if there are any similarity or not</p>",
          "rawMarkdown": "I didn't use imagehash to label images i just use to figure out if there are any similarity or not"
        },
        {
          "id": 1216481,
          "postDate": "2021-02-24T09:37:26.253Z",
          "content": "<p>I think host should clarify hand labeling, pseudo-labelling and the usage of external dataset.</p>",
          "rawMarkdown": "I think host should clarify hand labeling, pseudo-labelling and the usage of external dataset.",
          "votes": 1
        },
        {
          "id": 1216701,
          "postDate": "2021-02-24T12:39:36.283Z",
          "content": "<p>+1 to <strong>kaggle need to issue a rule that hand labeling of external data is allowed or not</strong></p>\n<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a> this is particularly important since it could \"break\" the entire competition</p>",
          "rawMarkdown": "+1 to **kaggle need to issue a rule that hand labeling of external data is allowed or not**\n\n@jarrelscy this is particularly important since it could \"break\" the entire competition",
          "votes": 1
        }
      ]
    },
    {
      "id": 1219517,
      "postDate": "2021-02-26T22:53:36.177Z",
      "content": "<p>For anyone thinking of using this, please consider that out of 29000 matches, there are over 800 false positives.</p>",
      "rawMarkdown": "For anyone thinking of using this, please consider that out of 29000 matches, there are over 800 false positives.",
      "votes": 6,
      "replies": [
        {
          "id": 1220320,
          "postDate": "2021-02-27T21:03:15.673Z",
          "content": "<p>Sorry, not sure what you mean by false positives? They showed up as duplicates but they are not actually matches?</p>",
          "rawMarkdown": "Sorry, not sure what you mean by false positives? They showed up as duplicates but they are not actually matches?",
          "votes": 2
        },
        {
          "id": 1221018,
          "postDate": "2021-02-28T15:58:53.530Z",
          "content": "<p>exactly. imagehash is not perfect</p>",
          "rawMarkdown": "exactly. imagehash is not perfect",
          "votes": 2
        },
        {
          "id": 1221804,
          "postDate": "2021-03-01T09:51:01.353Z",
          "content": "<p>Hello,is your lb score better after you rectify them?</p>",
          "rawMarkdown": "Hello,is your lb score better after you rectify them?"
        },
        {
          "id": 1221823,
          "postDate": "2021-03-01T10:25:13.970Z",
          "content": "<p>We do not use NIH dataset yet - so don't know. Time will tell :)</p>",
          "rawMarkdown": "We do not use NIH dataset yet - so don't know. Time will tell :)",
          "votes": 1
        },
        {
          "id": 1222176,
          "postDate": "2021-03-01T15:55:16.997Z",
          "content": "<p>I generated pseudo labels for the NIH dataset (after removing duplicates) and trained on them before fine tuning on the labeled data given to us; I saw a nice improvement. My cv results are:</p>\n<ul>\n<li>AUC: 0.9619 -&gt; 0.9636</li>\n<li>Loss (vanilla BCE): 0.1241 -&gt; 0.1192</li>\n</ul>\n<p>I am now concatenating pseudo + labeled and training them jointly with a modified loss function. I expect similar, if not better, results. </p>\n<p>Update 3/2/2021: I have reached public leaderboard <code>.965</code> with a model trained exclusively on NIH Chest XRay images (without any duplicates). </p>",
          "rawMarkdown": "I generated pseudo labels for the NIH dataset (after removing duplicates) and trained on them before fine tuning on the labeled data given to us; I saw a nice improvement. My cv results are:\n\n* AUC: 0.9619 -> 0.9636\n* Loss (vanilla BCE): 0.1241 -> 0.1192\n\nI am now concatenating pseudo + labeled and training them jointly with a modified loss function. I expect similar, if not better, results. \n\nUpdate 3/2/2021: I have reached public leaderboard `.965` with a model trained exclusively on NIH Chest XRay images (without any duplicates). ",
          "votes": 3
        },
        {
          "id": 1222532,
          "postDate": "2021-03-01T21:47:07.977Z",
          "content": "<p>One potential issue with this procedure is if you generated the pseudolabels by training on the validation data then you might potentially leak in information about the validation set in the pseudolabels of the test set. A bit obtuse, but I would just be careful with metrics of that. </p>",
          "rawMarkdown": "One potential issue with this procedure is if you generated the pseudolabels by training on the validation data then you might potentially leak in information about the validation set in the pseudolabels of the test set. A bit obtuse, but I would just be careful with metrics of that. ",
          "votes": 1
        },
        {
          "id": 1222547,
          "postDate": "2021-03-01T22:10:15.783Z",
          "content": "<p>Very good point. I made sure to generate pseudo-labels from models that had not seen the fold I validated against in the final labeled training stage to avoid this potential leakage. </p>",
          "rawMarkdown": "Very good point. I made sure to generate pseudo-labels from models that had not seen the fold I validated against in the final labeled training stage to avoid this potential leakage. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1217735,
      "postDate": "2021-02-25T09:30:01.563Z",
      "content": "<p>Isn't it necessary to prohibit participants from hand-labelling external data and stipulate it in the rules? It is true that NIH CXR data is huge and it's hard to hand annotate them completely in manner of the annotators of this competition, but it seems to be worth clarifying the rules. Of course, whether it is possible to detect participants who use hand-labelling is also important. Can we properly eliminate them?</p>",
      "rawMarkdown": "Isn't it necessary to prohibit participants from hand-labelling external data and stipulate it in the rules? It is true that NIH CXR data is huge and it's hard to hand annotate them completely in manner of the annotators of this competition, but it seems to be worth clarifying the rules. Of course, whether it is possible to detect participants who use hand-labelling is also important. Can we properly eliminate them?",
      "votes": 1
    },
    {
      "id": 1216141,
      "postDate": "2021-02-24T06:08:20.773Z",
      "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> this it the link <br>\n<a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">https://www.kaggle.com/nih-chest-xrays/data</a> and thank you for your notebook it state of art </p>",
      "rawMarkdown": "@underwearfitting this it the link \nhttps://www.kaggle.com/nih-chest-xrays/data and thank you for your notebook it state of art ",
      "votes": 1
    },
    {
      "id": 1216381,
      "postDate": "2021-02-24T08:32:32.567Z",
      "content": "<p>Déjà vu! all previous chest x-ray competitions had the same discovery of competition data to NIH 14 data relation :) Unsurprising </p>",
      "rawMarkdown": "Déjà vu! all previous chest x-ray competitions had the same discovery of competition data to NIH 14 data relation :) Unsurprising ",
      "votes": 2
    },
    {
      "id": 1239730,
      "postDate": "2021-03-16T00:55:47.630Z",
      "content": "<p>Has anyone found additional duplicates from the NIH datasets besides these 29000? Thanks!</p>",
      "rawMarkdown": "Has anyone found additional duplicates from the NIH datasets besides these 29000? Thanks!"
    },
    {
      "id": 1216712,
      "postDate": "2021-02-24T12:48:56.217Z",
      "content": "<p><a href=\"https://www.kaggle.com/arc144\" target=\"_blank\">@arc144</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a><br>\nThe rules around external data and hand-labelling are specified in the competition rules under Section A2. </p>\n<p>\"2. EXTERNAL DATA.<br>\nPublicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"</p>",
      "rawMarkdown": "@arc144 @hengck23 @steamedsheep\nThe rules around external data and hand-labelling are specified in the competition rules under Section A2. \n\n\"2. EXTERNAL DATA.\nPublicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"",
      "replies": [
        {
          "id": 1216757,
          "postDate": "2021-02-24T13:20:27.920Z",
          "content": "<p>I think people are mostly interested if any external data used for training is considered \"training set\". Now it is unclear if \"training set\" is only organizer's provided data, or encompasses any other training data used</p>",
          "rawMarkdown": "I think people are mostly interested if any external data used for training is considered \"training set\". Now it is unclear if \"training set\" is only organizer's provided data, or encompasses any other training data used",
          "votes": 2
        },
        {
          "id": 1216777,
          "postDate": "2021-02-24T13:33:28.757Z",
          "content": "<p>And obviously if test data is part of CXR14 then you are not allowed to hand-label it. Also, what does \"re-annotate\" mean? I assume this only encompasses hand labeling.</p>",
          "rawMarkdown": "And obviously if test data is part of CXR14 then you are not allowed to hand-label it. Also, what does \"re-annotate\" mean? I assume this only encompasses hand labeling.",
          "votes": 3
        },
        {
          "id": 1216784,
          "postDate": "2021-02-24T13:36:28.293Z",
          "content": "<p>re-annotate would most likely mean fixing label errors. there have been multiple posts on that. Not that much of a label noise to be concerned tho - found like 30 instances myself - not using them in my models.</p>",
          "rawMarkdown": "re-annotate would most likely mean fixing label errors. there have been multiple posts on that. Not that much of a label noise to be concerned tho - found like 30 instances myself - not using them in my models.",
          "votes": 2
        },
        {
          "id": 1216785,
          "postDate": "2021-02-24T13:36:36.627Z",
          "content": "<p>Yeah, I think it comes down to semantics… </p>\n<blockquote>\n  <p>Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"</p>\n</blockquote>\n<p>I guess If you are manually annotating images in an external dataset for training purposes but some of those images are also present in the private test set, you are essentially violating the rules, right?</p>",
          "rawMarkdown": "Yeah, I think it comes down to semantics... \n\n> Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"\n\nI guess If you are manually annotating images in an external dataset for training purposes but some of those images are also present in the private test set, you are essentially violating the rules, right?",
          "votes": 2
        },
        {
          "id": 1216792,
          "postDate": "2021-02-24T13:42:14.760Z",
          "content": "<blockquote>\n  <p>I guess If you are manually annotating images in an external dataset for training purposes but some of those images are also present in the private test set, you are essentially violating the rules, right?</p>\n</blockquote>\n<p>If hosts make it clear, yes. So I think they should make it clear. Otherwise, how can you know? That said, it is written on the competition page, so I guess this is not allowed for sure, but doesn't hurt to clarify.</p>",
          "rawMarkdown": "> I guess If you are manually annotating images in an external dataset for training purposes but some of those images are also present in the private test set, you are essentially violating the rules, right?\n\nIf hosts make it clear, yes. So I think they should make it clear. Otherwise, how can you know? That said, it is written on the competition page, so I guess this is not allowed for sure, but doesn't hurt to clarify."
        },
        {
          "id": 1216812,
          "postDate": "2021-02-24T13:53:19.020Z",
          "content": "<p>I already tried hand-labelling some dataset (like 300 images) as I thought I had some qualification to do so. And I say it is very very hard to mimic the labelling process done in this competition. model metrics on my set labels were quite terrible!</p>\n<p>then I asked 2 radiologists, friends of mine to do the same - it ended in same result. </p>\n<p>So I am not too worried about hand-labelling problem at all.</p>",
          "rawMarkdown": "I already tried hand-labelling some dataset (like 300 images) as I thought I had some qualification to do so. And I say it is very very hard to mimic the labelling process done in this competition. model metrics on my set labels were quite terrible!\n\nthen I asked 2 radiologists, friends of mine to do the same - it ended in same result. \n\nSo I am not too worried about hand-labelling problem at all.",
          "votes": 7
        },
        {
          "id": 1216813,
          "postDate": "2021-02-24T13:57:02.513Z",
          "content": "<p>I agree, but you never know what crazy things people plan to do :D</p>",
          "rawMarkdown": "I agree, but you never know what crazy things people plan to do :D",
          "votes": 3
        },
        {
          "id": 1216867,
          "postDate": "2021-02-24T14:53:38.487Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1216985,
          "postDate": "2021-02-24T17:18:13.133Z",
          "content": "<p>If I annote the position of tubes based on the segmentation model's output. Must I publish these annotation to the official competition forum prior to the Entry Deadline? </p>",
          "rawMarkdown": "If I annote the position of tubes based on the segmentation model's output. Must I publish these annotation to the official competition forum prior to the Entry Deadline? ",
          "votes": 1
        },
        {
          "id": 1217099,
          "postDate": "2021-02-24T18:42:03.187Z",
          "content": "<p>this is called pseudolabelling. You do not have to publish it as long as the model can reproduce the exact same labels</p>",
          "rawMarkdown": "this is called pseudolabelling. You do not have to publish it as long as the model can reproduce the exact same labels",
          "votes": 2
        },
        {
          "id": 1217104,
          "postDate": "2021-02-24T18:50:42.207Z",
          "content": "<p>I mean adjust and correct the segmentation model's output by hand. Time to give up this crazy idea. Thank you!</p>",
          "rawMarkdown": "I mean adjust and correct the segmentation model's output by hand. Time to give up this crazy idea. Thank you!"
        },
        {
          "id": 1217408,
          "postDate": "2021-02-25T03:14:55.103Z",
          "content": "<p>note this:</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data</a><br>\n\"This is a code-only competition so there is a hidden test set (approximately 4x larger, with ~14k images) as well.\"</p>\n<p>now we have:</p>\n<ul>\n<li>hand annotate train data is ok</li>\n<li>hand annotate test data is not ok</li>\n<li>hand annotate external data is  ok</li>\n<li>we can find external data == test public data</li>\n<li>we do not know which external data == test private data</li>\n</ul>",
          "rawMarkdown": "note this:\n\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data\n\"This is a code-only competition so there is a hidden test set (approximately 4x larger, with ~14k images) as well.\"\n\nnow we have:\n- hand annotate train data is ok\n- hand annotate test data is not ok\n- hand annotate external data is  ok\n- we can find external data == test public data\n- we do not know which external data == test private data\n\n",
          "votes": 5
        },
        {
          "id": 1217411,
          "postDate": "2021-02-25T03:18:40.643Z",
          "content": "<p>Is it leakage?</p>",
          "rawMarkdown": "Is it leakage?",
          "votes": -1
        },
        {
          "id": 1217415,
          "postDate": "2021-02-25T03:22:26.700Z",
          "content": "<p>\"So I am not too worried about hand-labelling problem at all.\"</p>\n<p>I think a top solution will be pseudo label external dataset, then some human in the loop to manual correct obvious mistake. hence it is important for the host to make it clear that any external hand label is allowed or not.</p>\n<p>even said that the kaggle do not check integrity of all submitted solutions. only the top solutions that are eligible for prize will be check for hand labeling or not. This is something that needs to be improved.</p>",
          "rawMarkdown": "\"So I am not too worried about hand-labelling problem at all.\"\n\nI think a top solution will be pseudo label external dataset, then some human in the loop to manual correct obvious mistake. hence it is important for the host to make it clear that any external hand label is allowed or not.\n\neven said that the kaggle do not check integrity of all submitted solutions. only the top solutions that are eligible for prize will be check for hand labeling or not. This is something that needs to be improved.",
          "votes": 3
        },
        {
          "id": 1217418,
          "postDate": "2021-02-25T03:27:57.973Z",
          "content": "<p>\"I mean adjust and correct the segmentation model's output by hand. Time to give up this crazy idea. Thank you!\"</p>\n<p>there are papers that mention that just by manual correction of 5% of the error (of algorithmicaly selected predicted cases), you can reduce error by more than 50%.</p>\n<p>Hence this is not a crazy paper. you can google for human in the loop papers or active learning</p>",
          "rawMarkdown": "\"I mean adjust and correct the segmentation model's output by hand. Time to give up this crazy idea. Thank you!\"\n\nthere are papers that mention that just by manual correction of 5% of the error (of algorithmicaly selected predicted cases), you can reduce error by more than 50%.\n\nHence this is not a crazy paper. you can google for human in the loop papers or active learning",
          "votes": 3
        },
        {
          "id": 1217424,
          "postDate": "2021-02-25T03:44:48.953Z",
          "content": "<p>'there are papers that mention that just by correctly 5% of the error (of algorithmically selected predicted cases), you can reduce error by more than 50%.'</p>\n<p>Yes, pseudo label mask give me pretty large boost, so hand correct hard sample may have a high chance to help. <br>\nCrazy means spend lot's of time and have to public the annotation to public. This maybe good at the start of the competition, but not this time point.<br>\nThanks for your paper.</p>",
          "rawMarkdown": "'there are papers that mention that just by correctly 5% of the error (of algorithmically selected predicted cases), you can reduce error by more than 50%.'\n\nYes, pseudo label mask give me pretty large boost, so hand correct hard sample may have a high chance to help. \nCrazy means spend lot's of time and have to public the annotation to public. This maybe good at the start of the competition, but not this time point.\nThanks for your paper.",
          "votes": 3
        },
        {
          "id": 1217609,
          "postDate": "2021-02-25T07:29:20.400Z",
          "content": "<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a> Do you think nobody would hand annotate it and can lie?</p>",
          "rawMarkdown": "@jarrelscy Do you think nobody would hand annotate it and can lie?",
          "votes": 1
        },
        {
          "id": 1218000,
          "postDate": "2021-02-25T13:31:34.473Z",
          "content": "<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a><br>\nThank you very much for your efforts to make the competition a success.</p>\n<p>Does re-annotation include pseudo-labeling as well as hand labeling?<br>\nSuppose it is possible to label chestx. We are allowed to re-label the external data, but according to the rules, we have to make those labels available to all participants by March 8, the entry deadline. Whether it is hand-labeling or pseudo-labeling using trained models, it has to be released as well, correct?</p>",
          "rawMarkdown": "@jarrelscy\nThank you very much for your efforts to make the competition a success.\n\nDoes re-annotation include pseudo-labeling as well as hand labeling?\nSuppose it is possible to label chestx. We are allowed to re-label the external data, but according to the rules, we have to make those labels available to all participants by March 8, the entry deadline. Whether it is hand-labeling or pseudo-labeling using trained models, it has to be released as well, correct?",
          "votes": 1
        },
        {
          "id": 1218014,
          "postDate": "2021-02-25T13:40:17.893Z",
          "content": "<p><a href=\"https://www.kaggle.com/YYama\" target=\"_blank\">@YYama</a> please read the rules again. pseudo-labelling is not hand-labelling.</p>",
          "rawMarkdown": "@YYama please read the rules again. pseudo-labelling is not hand-labelling.",
          "votes": 5
        },
        {
          "id": 1218020,
          "postDate": "2021-02-25T13:41:58.430Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1218035,
          "postDate": "2021-02-25T14:02:43.620Z",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a><br>\nThank you for quick reply!</p>\n<p>Sorry, I am a beginner of machine learning. After reading the rules, I thought that if the 're-annotated data' included labels, then it had to be made public regardless of the means.<br>\nSo, generally, re-annotation does not include pseudo-labeling? If so, I would love to start pseudo labeling of external data.</p>",
          "rawMarkdown": "@raddar and @nguyenbadung\nThank you for quick reply!\n\nSorry, I am a beginner of machine learning. After reading the rules, I thought that if the 're-annotated data' included labels, then it had to be made public regardless of the means.\nSo, generally, re-annotation does not include pseudo-labeling? If so, I would love to start pseudo labeling of external data.",
          "votes": 1
        },
        {
          "id": 1220426,
          "postDate": "2021-02-28T01:34:02.323Z",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>\"I say it is very very hard to mimic the labelling process done in this competition. model metrics on my set labels were quite terrible!\"</p>\n<hr>\n<p>BOOKMARK THIS!!!!</p>\n<p>the trick is not hand label, cvc-normal/abnormal, ett-normal/abnormal etc</p>\n<p>rather the trick is to mark the endpoints (or lines if you have lots of time) in CXR14. then feed the  rendered image into a classifier. The classifer will label the rendered images as cvc-normal/abnormal, ett-normal/abnormal etc.</p>\n<p>the classifier is trained on kaggle extra annotation (rendered images) and it will learn the style of annotation.</p>",
          "rawMarkdown": "@raddar \n\n\"I say it is very very hard to mimic the labelling process done in this competition. model metrics on my set labels were quite terrible!\"\n\n---\n\nBOOKMARK THIS!!!!\n\nthe trick is not hand label, cvc-normal/abnormal, ett-normal/abnormal etc\n\nrather the trick is to mark the endpoints (or lines if you have lots of time) in CXR14. then feed the  rendered image into a classifier. The classifer will label the rendered images as cvc-normal/abnormal, ett-normal/abnormal etc.\n\nthe classifier is trained on kaggle extra annotation (rendered images) and it will learn the style of annotation.",
          "votes": 1
        },
        {
          "id": 1220427,
          "postDate": "2021-02-28T01:35:09.573Z",
          "content": "<p>Is it permitted?</p>",
          "rawMarkdown": "Is it permitted?"
        },
        {
          "id": 1220428,
          "postDate": "2021-02-28T01:36:49.977Z",
          "content": "<p>if you can't do it for CXR14, then at least do it for the rest of the non-annotated train images in kaggle<br>\nsee my post <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243</a></p>\n<p>but one think of how to do it carefully to share or not to share via \"(i) ensure the re-annotated data is available to use by all participants\"</p>\n<p>it seems that is rule is specific to RANZCR? the winning solutions recent rfcx audio classification rely heaving on hand label. but I don't see the same rule applies there.</p>\n<hr>\n<p>in my experiment, the effects of just few annotations is huge.</p>",
          "rawMarkdown": "if you can't do it for CXR14, then at least do it for the rest of the non-annotated train images in kaggle\nsee my post https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\n\nbut one think of how to do it carefully to share or not to share via \"(i) ensure the re-annotated data is available to use by all participants\"\n\nit seems that is rule is specific to RANZCR? the winning solutions recent rfcx audio classification rely heaving on hand label. but I don't see the same rule applies there.\n\n---\n\nin my experiment, the effects of just few annotations is huge."
        },
        {
          "id": 1220445,
          "postDate": "2021-02-28T02:09:18.893Z",
          "content": "<p>So do us, Annotations boost a lot.</p>",
          "rawMarkdown": "So do us, Annotations boost a lot.",
          "votes": 1
        },
        {
          "id": 1220447,
          "postDate": "2021-02-28T02:13:25.860Z",
          "content": "<p>so the question now is how to do it smartly (fast, accurate, and with ease) …. by hand, pseudolabel, etc … automatic, semi-automatic, human in the loop, synthetic …</p>",
          "rawMarkdown": "so the question now is how to do it smartly (fast, accurate, and with ease) .... by hand, pseudolabel, etc ... automatic, semi-automatic, human in the loop, synthetic ..."
        },
        {
          "id": 1220449,
          "postDate": "2021-02-28T02:17:01.973Z",
          "content": "<p>HI,How to make some annotations?</p>",
          "rawMarkdown": "HI,How to make some annotations?"
        },
        {
          "id": 1220454,
          "postDate": "2021-02-28T02:26:32.203Z",
          "content": "<p>I feel that before the merge ddl, there will be many hand labelled data 's releasing.<br>\n'HI,How to make some annotations?', <br>\nBasicly you use labelling tool if you understande the problem. Or some Instence segmentation model to generate  pseudolabel, or other method mentioned by 🐸.</p>",
          "rawMarkdown": "I feel that before the merge ddl, there will be many hand labelled data 's releasing.\n'HI,How to make some annotations?', \nBasicly you use labelling tool if you understande the problem. Or some Instence segmentation model to generate  pseudolabel, or other method mentioned by 🐸."
        },
        {
          "id": 1220458,
          "postDate": "2021-02-28T02:39:08.960Z",
          "content": "<p>there is a possibility that we end up making annotation tool like:<br>\n(this become a marker tool competition)</p>\n<p>PolygonRNN++ or Curve-GCN</p>\n<p>there are some very interesting work on polygon trasnformer</p>\n<p><img src=\"https://amlankar.github.io/img/curve-gcn.jpg\" alt=\"\"></p>\n<p>these are like photoshop line snapping tool</p>",
          "rawMarkdown": "there is a possibility that we end up making annotation tool like:\n(this become a marker tool competition)\n\nPolygonRNN++ or Curve-GCN\n\nthere are some very interesting work on polygon trasnformer\n\n![](https://amlankar.github.io/img/curve-gcn.jpg)\n\nthese are like photoshop line snapping tool",
          "votes": 3
        },
        {
          "id": 1220459,
          "postDate": "2021-02-28T02:39:53.853Z",
          "content": "<blockquote>\n  <p>the trick is to mark the endpoints</p>\n</blockquote>\n<p>I think this idea is great and crucial for this competion.</p>\n<blockquote>\n  <p>it seems that is rule is specific to RANZCR? the winning solutions recent rfcx audio classification rely heaving on hand label. but I don't see familiar rule applies.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a></p>\n<p>It is significant to clarify whether it is permitted or not. Could you inform us of the official view?</p>",
          "rawMarkdown": "> the trick is to mark the endpoints\n\nI think this idea is great and crucial for this competion.\n\n> it seems that is rule is specific to RANZCR? the winning solutions recent rfcx audio classification rely heaving on hand label. but I don't see familiar rule applies.\n\n@jarrelscy\n\nIt is significant to clarify whether it is permitted or not. Could you inform us of the official view?",
          "votes": 1
        },
        {
          "id": 1220466,
          "postDate": "2021-02-28T02:59:26.487Z",
          "content": "<p><a href=\"https://www.kaggle.com/likein12\" target=\"_blank\">@likein12</a> <br>\nI totally agree with you.<br>\nSince the definition of the term \"re-annotate\" is not clear, no one, except the hosts, can make the detailed rules clear.</p>\n<p>The following is just my opinion.<br>\nAs there is no way to confirm that all participants have complied with the rules except for those in the prize zone. I would like the hosts to make an announcement to allow as much freedom for competitors as possible.</p>",
          "rawMarkdown": "@likein12 \nI totally agree with you.\nSince the definition of the term \"re-annotate\" is not clear, no one, except the hosts, can make the detailed rules clear.\n\nThe following is just my opinion.\nAs there is no way to confirm that all participants have complied with the rules except for those in the prize zone. I would like the hosts to make an announcement to allow as much freedom for competitors as possible.",
          "votes": 1
        },
        {
          "id": 1220936,
          "postDate": "2021-02-28T14:13:39.650Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644\" target=\"_blank\">The host announcement about this matter has come.</a></p>",
          "rawMarkdown": "[The host announcement about this matter has come.](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1216277,
      "postDate": "2021-02-24T07:26:03.480Z",
      "content": "<p>Hi, how to find them when they are different size?</p>",
      "rawMarkdown": "Hi, how to find them when they are different size?"
    },
    {
      "id": 1216176,
      "postDate": "2021-02-24T06:20:07.643Z",
      "content": "<p>Thanks a lot for sharing <a href=\"https://www.kaggle.com/mohamed3abdelrazik\" target=\"_blank\">@mohamed3abdelrazik</a>! Wow, 29.000 is a lot! Do you think there might be false positives in your duplicate list? Did you include only train or train+test images of the competition data when identifying duplicates?</p>",
      "rawMarkdown": "Thanks a lot for sharing @mohamed3abdelrazik! Wow, 29.000 is a lot! Do you think there might be false positives in your duplicate list? Did you include only train or train+test images of the competition data when identifying duplicates?",
      "replies": [
        {
          "id": 1216191,
          "postDate": "2021-02-24T06:28:18.233Z",
          "content": "<p>I don't investigate if there are FP or not and i didn't include RANZCR test data i just warm up.But i will try on both test+train of RANZCR and if there is also duplicated image between RANZCR test and CHESTX so if we use additional data or pretrain model on CHESTX for find good starting point this will cause over-fitting  </p>",
          "rawMarkdown": "I don't investigate if there are FP or not and i didn't include RANZCR test data i just warm up.But i will try on both test+train of RANZCR and if there is also duplicated image between RANZCR test and CHESTX so if we use additional data or pretrain model on CHESTX for find good starting point this will cause over-fitting  "
        }
      ]
    },
    {
      "id": 1216121,
      "postDate": "2021-02-24T06:02:24.410Z",
      "content": "<p>Can you also share the link to chestx? Thanks.</p>",
      "rawMarkdown": "Can you also share the link to chestx? Thanks.",
      "replies": [
        {
          "id": 1216137,
          "postDate": "2021-02-24T06:06:47.063Z",
          "content": "<p><a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">https://www.kaggle.com/nih-chest-xrays/data</a></p>",
          "rawMarkdown": "https://www.kaggle.com/nih-chest-xrays/data",
          "votes": 1
        },
        {
          "id": 1216158,
          "postDate": "2021-02-24T06:12:33.460Z",
          "content": "<p>Thanks a lot to you two!</p>",
          "rawMarkdown": "Thanks a lot to you two!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1216331,
      "postDate": "2021-02-24T08:02:16.317Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 1216231,
      "postDate": "2021-02-24T06:45:11.603Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1216115,
      "postDate": "2021-02-24T05:58:56.200Z",
      "content": "<p>Thanks for the information.</p>",
      "rawMarkdown": "Thanks for the information.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1216234,
      "author_name": "Jarrel Seah",
      "author_url": "",
      "post_date": "2021-02-24T06:46:39.733000",
      "content": "<p><a href=\"https://www.kaggle.com/mohamed3abdelrazik\" target=\"_blank\">@mohamed3abdelrazik</a> As per the acknowledgements page - this dataset was created by relabelling the publicly available CXR14 dataset from NIH, therefore there will be 100% overlap in image data. These labels (concerning lines and tubes placement) are not part of the original dataset.</p>",
      "votes": 19,
      "replies": [
        {
          "id": 1216258,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-24T07:00:27.713000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1216265,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-24T07:08:39.063000",
          "content": "<p>Thanks a lot for pointing to that! I joined late and must have missed this. It is clear now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1216333,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-24T08:05:16.860000",
          "content": "",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1216343,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-24T08:11:33.853000",
          "content": "",
          "votes": -3,
          "replies": []
        },
        {
          "id": 1216607,
          "author_name": "Wang Xinliang",
          "author_url": "",
          "post_date": "2021-02-24T11:10:04.373000",
          "content": "<p>\"this dataset\" means the \"training set\" or  \"training set + testing set\" of the competition? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1216628,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-02-24T11:36:19.937000",
          "content": "<p><a href=\"https://www.kaggle.com/dldmw579\" target=\"_blank\">@dldmw579</a> my understanding is that it means \"training set + public test + private test\", i.e. the entire competition data is a subset of CXR14. But it would be nice to verify this to be sure.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1217492,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-25T05:54:43.070000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1218391,
          "author_name": "Holli Wainwright",
          "author_url": "",
          "post_date": "2021-02-25T19:51:26.543000",
          "content": "<p>good to know</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1216262,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-24T07:03:40.497000",
      "content": "<p>\" imagehash library to find dublicated images and found about 29000 duplicated\"</p>\n<p>self-supervised learning is the key?<br>\nthe private data has turned from \"black box\" to \"white box\"</p>\n<p><strong>kaggle need to issue a rule that hand labeling of external data is allowed or not</strong></p>",
      "votes": 8,
      "replies": [
        {
          "id": 1216350,
          "author_name": "Mohamed abdelrazik",
          "author_url": "",
          "post_date": "2021-02-24T08:15:11.227000",
          "content": "<p>I didn't use imagehash to label images i just use to figure out if there are any similarity or not</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1216481,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-02-24T09:37:26.253000",
          "content": "<p>I think host should clarify hand labeling, pseudo-labelling and the usage of external dataset.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1216701,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2021-02-24T12:39:36.283000",
          "content": "<p>+1 to <strong>kaggle need to issue a rule that hand labeling of external data is allowed or not</strong></p>\n<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a> this is particularly important since it could \"break\" the entire competition</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1219517,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2021-02-26T22:53:36.177000",
      "content": "<p>For anyone thinking of using this, please consider that out of 29000 matches, there are over 800 false positives.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1220320,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2021-02-27T21:03:15.673000",
          "content": "<p>Sorry, not sure what you mean by false positives? They showed up as duplicates but they are not actually matches?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1221018,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-02-28T15:58:53.530000",
          "content": "<p>exactly. imagehash is not perfect</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1221804,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-03-01T09:51:01.353000",
          "content": "<p>Hello,is your lb score better after you rectify them?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1221823,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-03-01T10:25:13.970000",
          "content": "<p>We do not use NIH dataset yet - so don't know. Time will tell :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1222176,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-03-01T15:55:16.997000",
          "content": "<p>I generated pseudo labels for the NIH dataset (after removing duplicates) and trained on them before fine tuning on the labeled data given to us; I saw a nice improvement. My cv results are:</p>\n<ul>\n<li>AUC: 0.9619 -&gt; 0.9636</li>\n<li>Loss (vanilla BCE): 0.1241 -&gt; 0.1192</li>\n</ul>\n<p>I am now concatenating pseudo + labeled and training them jointly with a modified loss function. I expect similar, if not better, results. </p>\n<p>Update 3/2/2021: I have reached public leaderboard <code>.965</code> with a model trained exclusively on NIH Chest XRay images (without any duplicates). </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1222532,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2021-03-01T21:47:07.977000",
          "content": "<p>One potential issue with this procedure is if you generated the pseudolabels by training on the validation data then you might potentially leak in information about the validation set in the pseudolabels of the test set. A bit obtuse, but I would just be careful with metrics of that. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1222547,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2021-03-01T22:10:15.783000",
          "content": "<p>Very good point. I made sure to generate pseudo-labels from models that had not seen the fold I validated against in the final labeled training stage to avoid this potential leakage. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1217735,
      "author_name": "rikein12",
      "author_url": "",
      "post_date": "2021-02-25T09:30:01.563000",
      "content": "<p>Isn't it necessary to prohibit participants from hand-labelling external data and stipulate it in the rules? It is true that NIH CXR data is huge and it's hard to hand annotate them completely in manner of the annotators of this competition, but it seems to be worth clarifying the rules. Of course, whether it is possible to detect participants who use hand-labelling is also important. Can we properly eliminate them?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1216141,
      "author_name": "Mohamed abdelrazik",
      "author_url": "",
      "post_date": "2021-02-24T06:08:20.773000",
      "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> this it the link <br>\n<a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">https://www.kaggle.com/nih-chest-xrays/data</a> and thank you for your notebook it state of art </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1216381,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2021-02-24T08:32:32.567000",
      "content": "<p>Déjà vu! all previous chest x-ray competitions had the same discovery of competition data to NIH 14 data relation :) Unsurprising </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1239730,
      "author_name": "Gooooose",
      "author_url": "",
      "post_date": "2021-03-16T00:55:47.630000",
      "content": "<p>Has anyone found additional duplicates from the NIH datasets besides these 29000? Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1216712,
      "author_name": "Jarrel Seah",
      "author_url": "",
      "post_date": "2021-02-24T12:48:56.217000",
      "content": "<p><a href=\"https://www.kaggle.com/arc144\" target=\"_blank\">@arc144</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a><br>\nThe rules around external data and hand-labelling are specified in the competition rules under Section A2. </p>\n<p>\"2. EXTERNAL DATA.<br>\nPublicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1216757,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-02-24T13:20:27.920000",
          "content": "<p>I think people are mostly interested if any external data used for training is considered \"training set\". Now it is unclear if \"training set\" is only organizer's provided data, or encompasses any other training data used</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1216777,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-02-24T13:33:28.757000",
          "content": "<p>And obviously if test data is part of CXR14 then you are not allowed to hand-label it. Also, what does \"re-annotate\" mean? I assume this only encompasses hand labeling.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1216784,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-02-24T13:36:28.293000",
          "content": "<p>re-annotate would most likely mean fixing label errors. there have been multiple posts on that. Not that much of a label noise to be concerned tho - found like 30 instances myself - not using them in my models.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1216785,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2021-02-24T13:36:36.627000",
          "content": "<p>Yeah, I think it comes down to semantics… </p>\n<blockquote>\n  <p>Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"</p>\n</blockquote>\n<p>I guess If you are manually annotating images in an external dataset for training purposes but some of those images are also present in the private test set, you are essentially violating the rules, right?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1216792,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-02-24T13:42:14.760000",
          "content": "<blockquote>\n  <p>I guess If you are manually annotating images in an external dataset for training purposes but some of those images are also present in the private test set, you are essentially violating the rules, right?</p>\n</blockquote>\n<p>If hosts make it clear, yes. So I think they should make it clear. Otherwise, how can you know? That said, it is written on the competition page, so I guess this is not allowed for sure, but doesn't hurt to clarify.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1216812,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-02-24T13:53:19.020000",
          "content": "<p>I already tried hand-labelling some dataset (like 300 images) as I thought I had some qualification to do so. And I say it is very very hard to mimic the labelling process done in this competition. model metrics on my set labels were quite terrible!</p>\n<p>then I asked 2 radiologists, friends of mine to do the same - it ended in same result. </p>\n<p>So I am not too worried about hand-labelling problem at all.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1216813,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-02-24T13:57:02.513000",
          "content": "<p>I agree, but you never know what crazy things people plan to do :D</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1216867,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-24T14:53:38.487000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1216985,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-02-24T17:18:13.133000",
          "content": "<p>If I annote the position of tubes based on the segmentation model's output. Must I publish these annotation to the official competition forum prior to the Entry Deadline? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1217099,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-02-24T18:42:03.187000",
          "content": "<p>this is called pseudolabelling. You do not have to publish it as long as the model can reproduce the exact same labels</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1217104,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-02-24T18:50:42.207000",
          "content": "<p>I mean adjust and correct the segmentation model's output by hand. Time to give up this crazy idea. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1217408,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-25T03:14:55.103000",
          "content": "<p>note this:</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/data</a><br>\n\"This is a code-only competition so there is a hidden test set (approximately 4x larger, with ~14k images) as well.\"</p>\n<p>now we have:</p>\n<ul>\n<li>hand annotate train data is ok</li>\n<li>hand annotate test data is not ok</li>\n<li>hand annotate external data is  ok</li>\n<li>we can find external data == test public data</li>\n<li>we do not know which external data == test private data</li>\n</ul>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1217411,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-25T03:18:40.643000",
          "content": "<p>Is it leakage?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1217415,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-25T03:22:26.700000",
          "content": "<p>\"So I am not too worried about hand-labelling problem at all.\"</p>\n<p>I think a top solution will be pseudo label external dataset, then some human in the loop to manual correct obvious mistake. hence it is important for the host to make it clear that any external hand label is allowed or not.</p>\n<p>even said that the kaggle do not check integrity of all submitted solutions. only the top solutions that are eligible for prize will be check for hand labeling or not. This is something that needs to be improved.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1217418,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-25T03:27:57.973000",
          "content": "<p>\"I mean adjust and correct the segmentation model's output by hand. Time to give up this crazy idea. Thank you!\"</p>\n<p>there are papers that mention that just by manual correction of 5% of the error (of algorithmicaly selected predicted cases), you can reduce error by more than 50%.</p>\n<p>Hence this is not a crazy paper. you can google for human in the loop papers or active learning</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1217424,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-02-25T03:44:48.953000",
          "content": "<p>'there are papers that mention that just by correctly 5% of the error (of algorithmically selected predicted cases), you can reduce error by more than 50%.'</p>\n<p>Yes, pseudo label mask give me pretty large boost, so hand correct hard sample may have a high chance to help. <br>\nCrazy means spend lot's of time and have to public the annotation to public. This maybe good at the start of the competition, but not this time point.<br>\nThanks for your paper.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1217609,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2021-02-25T07:29:20.400000",
          "content": "<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a> Do you think nobody would hand annotate it and can lie?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1218000,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2021-02-25T13:31:34.473000",
          "content": "<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a><br>\nThank you very much for your efforts to make the competition a success.</p>\n<p>Does re-annotation include pseudo-labeling as well as hand labeling?<br>\nSuppose it is possible to label chestx. We are allowed to re-label the external data, but according to the rules, we have to make those labels available to all participants by March 8, the entry deadline. Whether it is hand-labeling or pseudo-labeling using trained models, it has to be released as well, correct?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1218014,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-02-25T13:40:17.893000",
          "content": "<p><a href=\"https://www.kaggle.com/YYama\" target=\"_blank\">@YYama</a> please read the rules again. pseudo-labelling is not hand-labelling.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1218020,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-25T13:41:58.430000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1218035,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2021-02-25T14:02:43.620000",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a><br>\nThank you for quick reply!</p>\n<p>Sorry, I am a beginner of machine learning. After reading the rules, I thought that if the 're-annotated data' included labels, then it had to be made public regardless of the means.<br>\nSo, generally, re-annotation does not include pseudo-labeling? If so, I would love to start pseudo labeling of external data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220426,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-28T01:34:02.323000",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> </p>\n<p>\"I say it is very very hard to mimic the labelling process done in this competition. model metrics on my set labels were quite terrible!\"</p>\n<hr>\n<p>BOOKMARK THIS!!!!</p>\n<p>the trick is not hand label, cvc-normal/abnormal, ett-normal/abnormal etc</p>\n<p>rather the trick is to mark the endpoints (or lines if you have lots of time) in CXR14. then feed the  rendered image into a classifier. The classifer will label the rendered images as cvc-normal/abnormal, ett-normal/abnormal etc.</p>\n<p>the classifier is trained on kaggle extra annotation (rendered images) and it will learn the style of annotation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220427,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-28T01:35:09.573000",
          "content": "<p>Is it permitted?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220428,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-28T01:36:49.977000",
          "content": "<p>if you can't do it for CXR14, then at least do it for the rest of the non-annotated train images in kaggle<br>\nsee my post <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243</a></p>\n<p>but one think of how to do it carefully to share or not to share via \"(i) ensure the re-annotated data is available to use by all participants\"</p>\n<p>it seems that is rule is specific to RANZCR? the winning solutions recent rfcx audio classification rely heaving on hand label. but I don't see the same rule applies there.</p>\n<hr>\n<p>in my experiment, the effects of just few annotations is huge.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220445,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-02-28T02:09:18.893000",
          "content": "<p>So do us, Annotations boost a lot.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220447,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-28T02:13:25.860000",
          "content": "<p>so the question now is how to do it smartly (fast, accurate, and with ease) …. by hand, pseudolabel, etc … automatic, semi-automatic, human in the loop, synthetic …</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220449,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-28T02:17:01.973000",
          "content": "<p>HI,How to make some annotations?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220454,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-02-28T02:26:32.203000",
          "content": "<p>I feel that before the merge ddl, there will be many hand labelled data 's releasing.<br>\n'HI,How to make some annotations?', <br>\nBasicly you use labelling tool if you understande the problem. Or some Instence segmentation model to generate  pseudolabel, or other method mentioned by 🐸.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220458,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-28T02:39:08.960000",
          "content": "<p>there is a possibility that we end up making annotation tool like:<br>\n(this become a marker tool competition)</p>\n<p>PolygonRNN++ or Curve-GCN</p>\n<p>there are some very interesting work on polygon trasnformer</p>\n<p><img src=\"https://amlankar.github.io/img/curve-gcn.jpg\" alt=\"\"></p>\n<p>these are like photoshop line snapping tool</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1220459,
          "author_name": "rikein12",
          "author_url": "",
          "post_date": "2021-02-28T02:39:53.853000",
          "content": "<blockquote>\n  <p>the trick is to mark the endpoints</p>\n</blockquote>\n<p>I think this idea is great and crucial for this competion.</p>\n<blockquote>\n  <p>it seems that is rule is specific to RANZCR? the winning solutions recent rfcx audio classification rely heaving on hand label. but I don't see familiar rule applies.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a></p>\n<p>It is significant to clarify whether it is permitted or not. Could you inform us of the official view?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220466,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2021-02-28T02:59:26.487000",
          "content": "<p><a href=\"https://www.kaggle.com/likein12\" target=\"_blank\">@likein12</a> <br>\nI totally agree with you.<br>\nSince the definition of the term \"re-annotate\" is not clear, no one, except the hosts, can make the detailed rules clear.</p>\n<p>The following is just my opinion.<br>\nAs there is no way to confirm that all participants have complied with the rules except for those in the prize zone. I would like the hosts to make an announcement to allow as much freedom for competitors as possible.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220936,
          "author_name": "rikein12",
          "author_url": "",
          "post_date": "2021-02-28T14:13:39.650000",
          "content": "<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644\" target=\"_blank\">The host announcement about this matter has come.</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1216277,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-02-24T07:26:03.480000",
      "content": "<p>Hi, how to find them when they are different size?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1216176,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2021-02-24T06:20:07.643000",
      "content": "<p>Thanks a lot for sharing <a href=\"https://www.kaggle.com/mohamed3abdelrazik\" target=\"_blank\">@mohamed3abdelrazik</a>! Wow, 29.000 is a lot! Do you think there might be false positives in your duplicate list? Did you include only train or train+test images of the competition data when identifying duplicates?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1216191,
          "author_name": "Mohamed abdelrazik",
          "author_url": "",
          "post_date": "2021-02-24T06:28:18.233000",
          "content": "<p>I don't investigate if there are FP or not and i didn't include RANZCR test data i just warm up.But i will try on both test+train of RANZCR and if there is also duplicated image between RANZCR test and CHESTX so if we use additional data or pretrain model on CHESTX for find good starting point this will cause over-fitting  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1216121,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-02-24T06:02:24.410000",
      "content": "<p>Can you also share the link to chestx? Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1216137,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2021-02-24T06:06:47.063000",
          "content": "<p><a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">https://www.kaggle.com/nih-chest-xrays/data</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1216158,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2021-02-24T06:12:33.460000",
          "content": "<p>Thanks a lot to you two!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1216331,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-24T08:02:16.317000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1216231,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-24T06:45:11.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1216115,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-02-24T05:58:56.200000",
      "content": "<p>Thanks for the information.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1216107": "i see many people suggest to use additional data set and many found chestx data set the most properly one and thanks to @ammarali32 who use chestx dataset to find good starting point but i was so curious to figure out if there is any overlapped data between chestx and RANZCR original data set and the answer is yes this idea shine in my head after this [post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221159) thank you  @virilo i use imagehash library to find dublicated images and found about 29000 duplicated \n i uploaded 6 examples [Here](https://www.kaggle.com/mohamed3abdelrazik/4-examples-for-duplication)\nand a csv file contains image paths for both original and dublicated images cn be found [Here](https://www.kaggle.com/mohamed3abdelrazik/paths-for-duplicated-images-on-chestx-and-ranczr/settings)",
    "1216234": "@mohamed3abdelrazik As per the acknowledgements page - this dataset was created by relabelling the publicly available CXR14 dataset from NIH, therefore there will be 100% overlap in image data. These labels (concerning lines and tubes placement) are not part of the original dataset.",
    "1216262": "\" imagehash library to find dublicated images and found about 29000 duplicated\"\n\nself-supervised learning is the key?\nthe private data has turned from \"black box\" to \"white box\"\n\n**kaggle need to issue a rule that hand labeling of external data is allowed or not**",
    "1219517": "For anyone thinking of using this, please consider that out of 29000 matches, there are over 800 false positives.",
    "1217735": "Isn't it necessary to prohibit participants from hand-labelling external data and stipulate it in the rules? It is true that NIH CXR data is huge and it's hard to hand annotate them completely in manner of the annotators of this competition, but it seems to be worth clarifying the rules. Of course, whether it is possible to detect participants who use hand-labelling is also important. Can we properly eliminate them?",
    "1216141": "@underwearfitting this it the link \nhttps://www.kaggle.com/nih-chest-xrays/data and thank you for your notebook it state of art ",
    "1216381": "Déjà vu! all previous chest x-ray competitions had the same discovery of competition data to NIH 14 data relation :) Unsurprising ",
    "1239730": "Has anyone found additional duplicates from the NIH datasets besides these 29000? Thanks!",
    "1216712": "@arc144 @hengck23 @steamedsheep\nThe rules around external data and hand-labelling are specified in the competition rules under Section A2. \n\n\"2. EXTERNAL DATA.\nPublicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"",
    "1216277": "Hi, how to find them when they are different size?",
    "1216176": "Thanks a lot for sharing @mohamed3abdelrazik! Wow, 29.000 is a lot! Do you think there might be false positives in your duplicate list? Did you include only train or train+test images of the competition data when identifying duplicates?",
    "1216121": "Can you also share the link to chestx? Thanks.",
    "1216331": "",
    "1216231": "",
    "1216115": "Thanks for the information."
  }
}