{
  "id": 225266,
  "title": "Meta-Pseudo Labels: Sharing the idea I tried then gave up !",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/225266",
  "author_name": "",
  "post_date": "2021-03-11T14:00:21.470182300Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As soon as I read <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">this topic</a> and knew the data was obtained  by relabelling the publicly available CXR14 dataset from NIH,  I jumped back to the competition and  had  an idea to use <a href=\"https://arxiv.org/pdf/2003.10580.pdf\" target=\"_blank\">Meta Pseudo Labels</a>  from  the paper backing the current Imagenet SOTA. But I adapted it from multi-class to multi-labels setting. </p>\n<p>The concept was appealing for me after I read in the paper the methods outperformed both  previous SOTA : ViT-L and Bit-L by using <strong>unlabelled</strong>  JFT-300M while the latters used  <strong>labelled</strong>  JFT-300M . This is astonishing to read : Un/semi-supervised method beating a full supervised method while using the same external dataset. </p>\n<p>For my setting,  I used  the entire CXR14  dataset as unlebelled dataset (the paper uses JFT-300M) and the competition data as the labelled one (Imagenet one in the paper ). </p>\n<p>As well as the paper, my teacher model is trained on  the unlabelled data using <br>\n <a href=\"https://arxiv.org/pdf/1904.12848.pdf\" target=\"_blank\">UDA objective</a> (in addition  to the supervised training on the labelled dataset) and the student, supervised by the generated pseudo labels, give feedback to the teacher in order to generate better pseudo labels </p>\n<p>Unfornately and while this shows promising reults, it required substantial compute power that I don't have, due to multi-training objectives, requirement of big images resolution (I couldn't go upper than 384x384),  the large size of CXR14 dataset and the need of very long training [1]<br>\nI end it up abandoning the idea as well as the competition :(</p>\n<p>[1} Guys from google are accustomed to use clusters of hundreds TPUs and other compute power from outer space :xD</p>",
  "messages": [
    {
      "id": "1234697",
      "postDate": "03/11/2021 14:00:21",
      "content": "<p>As soon as I read <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">this topic</a> and knew the data was obtained  by relabelling the publicly available CXR14 dataset from NIH,  I jumped back to the competition and  had  an idea to use <a href=\"https://arxiv.org/pdf/2003.10580.pdf\" target=\"_blank\">Meta Pseudo Labels</a>  from  the paper backing the current Imagenet SOTA. But I adapted it from multi-class to multi-labels setting. </p>\n<p>The concept was appealing for me after I read in the paper the methods outperformed both  previous SOTA : ViT-L and Bit-L by using <strong>unlabelled</strong>  JFT-300M while the latters used  <strong>labelled</strong>  JFT-300M . This is astonishing to read : Un/semi-supervised method beating a full supervised method while using the same external dataset. </p>\n<p>For my setting,  I used  the entire CXR14  dataset as unlebelled dataset (the paper uses JFT-300M) and the competition data as the labelled one (Imagenet one in the paper ). </p>\n<p>As well as the paper, my teacher model is trained on  the unlabelled data using <br>\n <a href=\"https://arxiv.org/pdf/1904.12848.pdf\" target=\"_blank\">UDA objective</a> (in addition  to the supervised training on the labelled dataset) and the student, supervised by the generated pseudo labels, give feedback to the teacher in order to generate better pseudo labels </p>\n<p>Unfornately and while this shows promising reults, it required substantial compute power that I don't have, due to multi-training objectives, requirement of big images resolution (I couldn't go upper than 384x384),  the large size of CXR14 dataset and the need of very long training [1]<br>\nI end it up abandoning the idea as well as the competition :(</p>\n<p>[1} Guys from google are accustomed to use clusters of hundreds TPUs and other compute power from outer space :xD</p>",
      "rawMarkdown": "As soon as I read [this topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808) and knew the data was obtained  by relabelling the publicly available CXR14 dataset from NIH,  I jumped back to the competition and  had  an idea to use [Meta Pseudo Labels](https://arxiv.org/pdf/2003.10580.pdf)  from  the paper backing the current Imagenet SOTA. But I adapted it from multi-class to multi-labels setting. \n\nThe concept was appealing for me after I read in the paper the methods outperformed both  previous SOTA : ViT-L and Bit-L by using **unlabelled**  JFT-300M while the latters used  **labelled**  JFT-300M . This is astonishing to read : Un/semi-supervised method beating a full supervised method while using the same external dataset. \n\nFor my setting,  I used  the entire CXR14  dataset as unlebelled dataset (the paper uses JFT-300M) and the competition data as the labelled one (Imagenet one in the paper ). \n\nAs well as the paper, my teacher model is trained on  the unlabelled data using \n [UDA objective](https://arxiv.org/pdf/1904.12848.pdf) (in addition  to the supervised training on the labelled dataset) and the student, supervised by the generated pseudo labels, give feedback to the teacher in order to generate better pseudo labels \n\n\nUnfornately and while this shows promising reults, it required substantial compute power that I don't have, due to multi-training objectives, requirement of big images resolution (I couldn't go upper than 384x384),  the large size of CXR14 dataset and the need of very long training [1]\nI end it up abandoning the idea as well as the competition :(\n\n[1} Guys from google are accustomed to use clusters of hundreds TPUs and other compute power from outer space :xD",
      "votes": null
    },
    {
      "id": "1236126",
      "postDate": "03/12/2021 19:33:29",
      "content": "<p>Can you share what is the improvement it gave you?</p>",
      "rawMarkdown": "Can you share what is the improvement it gave you?",
      "votes": null
    },
    {
      "id": "1236150",
      "postDate": "03/12/2021 19:59:43",
      "content": "<p>After 20 epoches, The student model got 0.946 on the competition validation fold0 data. While it was only trained on unlabelled CXR14 data with pseudo labels generated by the teacher (which hasn't see the fold0 validation) </p>\n<p>That means it could get much better result with bigger resolution and fine-tuned afterwards on the competition data. </p>\n<p>But just this took me more than a day to train (on single V100 16gb)</p>",
      "rawMarkdown": "After 20 epoches, The student model got 0.946 on the competition validation fold0 data. While it was only trained on unlabelled CXR14 data with pseudo labels generated by the teacher (which hasn't see the fold0 validation) \n\nThat means it could get much better result with bigger resolution and fine-tuned afterwards on the competition data. \n\nBut just this took me more than a day to train (on single V100 16gb)",
      "votes": null
    },
    {
      "id": "1237502",
      "postDate": "03/14/2021 07:54:18",
      "content": "<p>What was the CV score of your teacher model? Also, have you used the whole NIH for training? From other discussions we saw that RANZCR data is actually a subset of NIH. So maybe your student model has seen the validation data during training. Anyway you did an amazing job! One tip I want to share because I am also implementing the same method: If you use a reduced teacher (Multi-Layer-Perceptron) like described on the paper you can use larger resolution (600 x 600) global batch size 64 on the free Colab TPU</p>",
      "rawMarkdown": "What was the CV score of your teacher model? Also, have you used the whole NIH for training? From other discussions we saw that RANZCR data is actually a subset of NIH. So maybe your student model has seen the validation data during training. Anyway you did an amazing job! One tip I want to share because I am also implementing the same method: If you use a reduced teacher (Multi-Layer-Perceptron) like described on the paper you can use larger resolution (600 x 600) global batch size 64 on the free Colab TPU",
      "votes": null
    },
    {
      "id": "1237839",
      "postDate": "03/14/2021 13:23:46",
      "content": "<p>The goal is to teach and make robust the student model, so I didn't compute CV for the teacher.  Neither the teacher nor the student has seen the competition validation data. May be they have seen some overlap on NIH . But I don't think it's an issue, given the data was relabelled by the host, so the teacher will try to teach new labels to the student.</p>\n<p>Reduced Teacher might be a good idea, I was using the same big model for both.</p>\n<p>I use Pytorch, there is still CPU/Dataloading bottleneck to use Pytorch/XLA TPU on colab as efficiently as TF does. </p>\n<p>Right now, V100  on Colab Pro gives better option for Pytorch. </p>",
      "rawMarkdown": "The goal is to teach and make robust the student model, so I didn't compute CV for the teacher.  Neither the teacher nor the student has seen the competition validation data. May be they have seen some overlap on NIH . But I don't think it's an issue, given the data was relabelled by the host, so the teacher will try to teach new labels to the student.\n\nReduced Teacher might be a good idea, I was using the same big model for both.\n\nI use Pytorch, there is still CPU/Dataloading bottleneck to use Pytorch/XLA TPU on colab as efficiently as TF does. \n\nRight now, V100  on Colab Pro gives better option for Pytorch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1236126,
      "author_name": "morizin",
      "author_url": "",
      "post_date": "03/12/2021 19:33:29",
      "content": "<p>Can you share what is the improvement it gave you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1236150,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "03/12/2021 19:59:43",
          "content": "<p>After 20 epoches, The student model got 0.946 on the competition validation fold0 data. While it was only trained on unlabelled CXR14 data with pseudo labels generated by the teacher (which hasn't see the fold0 validation) </p>\n<p>That means it could get much better result with bigger resolution and fine-tuned afterwards on the competition data. </p>\n<p>But just this took me more than a day to train (on single V100 16gb)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1237502,
      "author_name": "nickthenick23",
      "author_url": "",
      "post_date": "03/14/2021 07:54:18",
      "content": "<p>What was the CV score of your teacher model? Also, have you used the whole NIH for training? From other discussions we saw that RANZCR data is actually a subset of NIH. So maybe your student model has seen the validation data during training. Anyway you did an amazing job! One tip I want to share because I am also implementing the same method: If you use a reduced teacher (Multi-Layer-Perceptron) like described on the paper you can use larger resolution (600 x 600) global batch size 64 on the free Colab TPU</p>",
      "votes": null,
      "replies": [
        {
          "id": 1237839,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "03/14/2021 13:23:46",
          "content": "<p>The goal is to teach and make robust the student model, so I didn't compute CV for the teacher.  Neither the teacher nor the student has seen the competition validation data. May be they have seen some overlap on NIH . But I don't think it's an issue, given the data was relabelled by the host, so the teacher will try to teach new labels to the student.</p>\n<p>Reduced Teacher might be a good idea, I was using the same big model for both.</p>\n<p>I use Pytorch, there is still CPU/Dataloading bottleneck to use Pytorch/XLA TPU on colab as efficiently as TF does. </p>\n<p>Right now, V100  on Colab Pro gives better option for Pytorch. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1234697": "As soon as I read [this topic](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808) and knew the data was obtained  by relabelling the publicly available CXR14 dataset from NIH,  I jumped back to the competition and  had  an idea to use [Meta Pseudo Labels](https://arxiv.org/pdf/2003.10580.pdf)  from  the paper backing the current Imagenet SOTA. But I adapted it from multi-class to multi-labels setting. \n\nThe concept was appealing for me after I read in the paper the methods outperformed both  previous SOTA : ViT-L and Bit-L by using **unlabelled**  JFT-300M while the latters used  **labelled**  JFT-300M . This is astonishing to read : Un/semi-supervised method beating a full supervised method while using the same external dataset. \n\nFor my setting,  I used  the entire CXR14  dataset as unlebelled dataset (the paper uses JFT-300M) and the competition data as the labelled one (Imagenet one in the paper ). \n\nAs well as the paper, my teacher model is trained on  the unlabelled data using \n [UDA objective](https://arxiv.org/pdf/1904.12848.pdf) (in addition  to the supervised training on the labelled dataset) and the student, supervised by the generated pseudo labels, give feedback to the teacher in order to generate better pseudo labels \n\n\nUnfornately and while this shows promising reults, it required substantial compute power that I don't have, due to multi-training objectives, requirement of big images resolution (I couldn't go upper than 384x384),  the large size of CXR14 dataset and the need of very long training [1]\nI end it up abandoning the idea as well as the competition :(\n\n[1} Guys from google are accustomed to use clusters of hundreds TPUs and other compute power from outer space :xD",
    "1236126": "Can you share what is the improvement it gave you?",
    "1236150": "After 20 epoches, The student model got 0.946 on the competition validation fold0 data. While it was only trained on unlabelled CXR14 data with pseudo labels generated by the teacher (which hasn't see the fold0 validation) \n\nThat means it could get much better result with bigger resolution and fine-tuned afterwards on the competition data. \n\nBut just this took me more than a day to train (on single V100 16gb)",
    "1237502": "What was the CV score of your teacher model? Also, have you used the whole NIH for training? From other discussions we saw that RANZCR data is actually a subset of NIH. So maybe your student model has seen the validation data during training. Anyway you did an amazing job! One tip I want to share because I am also implementing the same method: If you use a reduced teacher (Multi-Layer-Perceptron) like described on the paper you can use larger resolution (600 x 600) global batch size 64 on the free Colab TPU",
    "1237839": "The goal is to teach and make robust the student model, so I didn't compute CV for the teacher.  Neither the teacher nor the student has seen the competition validation data. May be they have seen some overlap on NIH . But I don't think it's an issue, given the data was relabelled by the host, so the teacher will try to teach new labels to the student.\n\nReduced Teacher might be a good idea, I was using the same big model for both.\n\nI use Pytorch, there is still CPU/Dataloading bottleneck to use Pytorch/XLA TPU on colab as efficiently as TF does. \n\nRight now, V100  on Colab Pro gives better option for Pytorch."
  },
  "source": "meta"
}