{
  "id": 311986,
  "title": "Self Supervised Methods ?",
  "url": "/competitions/ultra-mnist/discussion/311986",
  "author_name": "",
  "post_date": "2022-03-09T20:46:25.307201900Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Using external dataset is not allowed. That means you can't fintune object detectors for predicting the bounding boxes for the digits without putting manual effort of annotating the given dataset (Assuming there is no public available model trained on such dataset). Cost of manually annotating the whole dataset is huge.</p>\n<p>Seems like a perfect opportunity to explore self supervised learning.</p>\n<p>Attaching helpful resources to learn about the same:</p>\n<ul>\n<li><a href=\"https://youtube.com/playlist?list=PLd9i_xMMzZF7QiPZNF7zblpTksTt0DDy6\" target=\"_blank\">Introductory youtube video</a> by Anuj Shah, about using self supervised learning for classification</li>\n</ul>",
  "messages": [
    {
      "id": "1717350",
      "postDate": "03/09/2022 20:46:25",
      "content": "<p>Using external dataset is not allowed. That means you can't fintune object detectors for predicting the bounding boxes for the digits without putting manual effort of annotating the given dataset (Assuming there is no public available model trained on such dataset). Cost of manually annotating the whole dataset is huge.</p>\n<p>Seems like a perfect opportunity to explore self supervised learning.</p>\n<p>Attaching helpful resources to learn about the same:</p>\n<ul>\n<li><a href=\"https://youtube.com/playlist?list=PLd9i_xMMzZF7QiPZNF7zblpTksTt0DDy6\" target=\"_blank\">Introductory youtube video</a> by Anuj Shah, about using self supervised learning for classification</li>\n</ul>",
      "rawMarkdown": "Using external dataset is not allowed. That means you can't fintune object detectors for predicting the bounding boxes for the digits without putting manual effort of annotating the given dataset (Assuming there is no public available model trained on such dataset). Cost of manually annotating the whole dataset is huge.\n\nSeems like a perfect opportunity to explore self supervised learning.\n\nAttaching helpful resources to learn about the same:\n- [Introductory youtube video](https://youtube.com/playlist?list=PLd9i_xMMzZF7QiPZNF7zblpTksTt0DDy6) by Anuj Shah, about using self supervised learning for classification",
      "votes": null
    },
    {
      "id": "1717389",
      "postDate": "03/09/2022 22:21:55",
      "content": "<p>When you mention self-supervised are you considering building a Student-Teacher Models kind of scenario. Where the knowledge from one distill out to another. On the other hand it does make sense to annotate data without much effort but would it be accurate? What are your thoughts(As self-supervised will also lead to a probabilistic result)</p>",
      "rawMarkdown": "When you mention self-supervised are you considering building a Student-Teacher Models kind of scenario. Where the knowledge from one distill out to another. On the other hand it does make sense to annotate data without much effort but would it be accurate? What are your thoughts(As self-supervised will also lead to a probabilistic result)",
      "votes": null
    },
    {
      "id": "1717773",
      "postDate": "03/10/2022 08:00:07",
      "content": "<p>No. What you are talking about is called Knowledge distillation. It is completely different from self supervised learning, as in distillation the teacher model is trained on the given labelled dataset.</p>\n<p>self supervised learning is helpful especially when large portion of dataset is unlabelled and only some small section is labelled.</p>\n<p>Yes, now that we can get labelled data easily, using self supervised learning doesn't make sense.<br>\nI feel results from model trained on generated data will have better results compared to the one trained using self supervised learning. </p>\n<p>One thing you have to ensure if using generated data is that there should not be significant domain shift between generated data and the provided data.<br>\nCommon approaches of generating data that could cause domain shift:</p>\n<ul>\n<li>Keeping bounding boxes of same size.. The dataset has digits of different size.</li>\n<li>Manually annotating only a few samples per digit</li>\n<li>using augmentations like flip, random crop, center crop etc</li>\n</ul>\n<p>How to check if the data distribution of generated dataset is same as that of provided data ?<br>\nThis is called covariate shift. You can read about the existing methods. One simple way is to train a CNN to classify. the data as generated or original. If the classifier is able to distinguish these two, then there is a significant change in the data distribution and the generated data is not good.</p>",
      "rawMarkdown": "No. What you are talking about is called Knowledge distillation. It is completely different from self supervised learning, as in distillation the teacher model is trained on the given labelled dataset.\n\nself supervised learning is helpful especially when large portion of dataset is unlabelled and only some small section is labelled.\n\nYes, now that we can get labelled data easily, using self supervised learning doesn't make sense.\nI feel results from model trained on generated data will have better results compared to the one trained using self supervised learning. \n\nOne thing you have to ensure if using generated data is that there should not be significant domain shift between generated data and the provided data.\nCommon approaches of generating data that could cause domain shift:\n- Keeping bounding boxes of same size.. The dataset has digits of different size.\n- Manually annotating only a few samples per digit\n- using augmentations like flip, random crop, center crop etc\n\nHow to check if the data distribution of generated dataset is same as that of provided data ?\nThis is called covariate shift. You can read about the existing methods. One simple way is to train a CNN to classify. the data as generated or original. If the classifier is able to distinguish these two, then there is a significant change in the data distribution and the generated data is not good.",
      "votes": null
    },
    {
      "id": "1717811",
      "postDate": "03/10/2022 08:36:58",
      "content": "<p>I understand what you mean. I actually mis quoted distillation instead of self supervision. </p>",
      "rawMarkdown": "I understand what you mean. I actually mis quoted distillation instead of self supervision.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1717389,
      "author_name": "shivkumarganesh",
      "author_url": "",
      "post_date": "03/09/2022 22:21:55",
      "content": "<p>When you mention self-supervised are you considering building a Student-Teacher Models kind of scenario. Where the knowledge from one distill out to another. On the other hand it does make sense to annotate data without much effort but would it be accurate? What are your thoughts(As self-supervised will also lead to a probabilistic result)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1717773,
          "author_name": "harshraj22",
          "author_url": "",
          "post_date": "03/10/2022 08:00:07",
          "content": "<p>No. What you are talking about is called Knowledge distillation. It is completely different from self supervised learning, as in distillation the teacher model is trained on the given labelled dataset.</p>\n<p>self supervised learning is helpful especially when large portion of dataset is unlabelled and only some small section is labelled.</p>\n<p>Yes, now that we can get labelled data easily, using self supervised learning doesn't make sense.<br>\nI feel results from model trained on generated data will have better results compared to the one trained using self supervised learning. </p>\n<p>One thing you have to ensure if using generated data is that there should not be significant domain shift between generated data and the provided data.<br>\nCommon approaches of generating data that could cause domain shift:</p>\n<ul>\n<li>Keeping bounding boxes of same size.. The dataset has digits of different size.</li>\n<li>Manually annotating only a few samples per digit</li>\n<li>using augmentations like flip, random crop, center crop etc</li>\n</ul>\n<p>How to check if the data distribution of generated dataset is same as that of provided data ?<br>\nThis is called covariate shift. You can read about the existing methods. One simple way is to train a CNN to classify. the data as generated or original. If the classifier is able to distinguish these two, then there is a significant change in the data distribution and the generated data is not good.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1717811,
          "author_name": "shivkumarganesh",
          "author_url": "",
          "post_date": "03/10/2022 08:36:58",
          "content": "<p>I understand what you mean. I actually mis quoted distillation instead of self supervision. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1717350": "Using external dataset is not allowed. That means you can't fintune object detectors for predicting the bounding boxes for the digits without putting manual effort of annotating the given dataset (Assuming there is no public available model trained on such dataset). Cost of manually annotating the whole dataset is huge.\n\nSeems like a perfect opportunity to explore self supervised learning.\n\nAttaching helpful resources to learn about the same:\n- [Introductory youtube video](https://youtube.com/playlist?list=PLd9i_xMMzZF7QiPZNF7zblpTksTt0DDy6) by Anuj Shah, about using self supervised learning for classification",
    "1717389": "When you mention self-supervised are you considering building a Student-Teacher Models kind of scenario. Where the knowledge from one distill out to another. On the other hand it does make sense to annotate data without much effort but would it be accurate? What are your thoughts(As self-supervised will also lead to a probabilistic result)",
    "1717773": "No. What you are talking about is called Knowledge distillation. It is completely different from self supervised learning, as in distillation the teacher model is trained on the given labelled dataset.\n\nself supervised learning is helpful especially when large portion of dataset is unlabelled and only some small section is labelled.\n\nYes, now that we can get labelled data easily, using self supervised learning doesn't make sense.\nI feel results from model trained on generated data will have better results compared to the one trained using self supervised learning. \n\nOne thing you have to ensure if using generated data is that there should not be significant domain shift between generated data and the provided data.\nCommon approaches of generating data that could cause domain shift:\n- Keeping bounding boxes of same size.. The dataset has digits of different size.\n- Manually annotating only a few samples per digit\n- using augmentations like flip, random crop, center crop etc\n\nHow to check if the data distribution of generated dataset is same as that of provided data ?\nThis is called covariate shift. You can read about the existing methods. One simple way is to train a CNN to classify. the data as generated or original. If the classifier is able to distinguish these two, then there is a significant change in the data distribution and the generated data is not good.",
    "1717811": "I understand what you mean. I actually mis quoted distillation instead of self supervision."
  },
  "source": "meta"
}