{
  "id": 125196,
  "title": "What about PseudoLabelling?",
  "url": "/competitions/bengaliai-cv19/discussion/125196",
  "author_name": "",
  "post_date": "2020-01-09T06:15:36.472058200Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Since the models are pretty consistent when predicting for <code>public</code> test-set as shown <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/124454\">here</a>. I was thinking of using them as pseduolabels and train using <code>train</code> + <code>public test</code> set. Did anybody try this approach? Any luck?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F674454cee189a2ebdb00212a29a79fa6%2Fpseudo.png?generation=1578550435266754&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "714166",
      "postDate": "01/09/2020 06:15:36",
      "content": "<p>Since the models are pretty consistent when predicting for <code>public</code> test-set as shown <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/124454\">here</a>. I was thinking of using them as pseduolabels and train using <code>train</code> + <code>public test</code> set. Did anybody try this approach? Any luck?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F674454cee189a2ebdb00212a29a79fa6%2Fpseudo.png?generation=1578550435266754&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Since the models are pretty consistent when predicting for `public` test-set as shown [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/124454). I was thinking of using them as pseduolabels and train using `train` + `public test` set. Did anybody try this approach? Any luck?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F674454cee189a2ebdb00212a29a79fa6%2Fpseudo.png?generation=1578550435266754&amp;alt=media)",
      "votes": null
    },
    {
      "id": "714171",
      "postDate": "01/09/2020 06:24:26",
      "content": "<p>I guess the organizers provided a small public test dataset of 12 images to safeguard against PseudoLabelling based approaches. \nAt the moment of writing, the topmost entry on the public leaderboard is 0.9875. If they had provided a larger public test dataset, people here would've scored &gt;99.5% on the metric. </p>",
      "rawMarkdown": "I guess the organizers provided a small public test dataset of 12 images to safeguard against PseudoLabelling based approaches. \nAt the moment of writing, the topmost entry on the public leaderboard is 0.9875. If they had provided a larger public test dataset, people here would've scored &gt;99.5% on the metric.",
      "votes": null
    },
    {
      "id": "714345",
      "postDate": "01/09/2020 10:29:03",
      "content": "<p>The same thing got in my mind. We can use external dataset for the remaining dataset (e.g they used most common occuring 1000 out of 1350) to improve accuracy.</p>",
      "rawMarkdown": "The same thing got in my mind. We can use external dataset for the remaining dataset (e.g they used most common occuring 1000 out of 1350) to improve accuracy.",
      "votes": null
    },
    {
      "id": "715941",
      "postDate": "01/11/2020 02:25:34",
      "content": "<p>yea i think it will work</p>",
      "rawMarkdown": "yea i think it will work",
      "votes": null
    },
    {
      "id": "717474",
      "postDate": "01/13/2020 07:07:22",
      "content": "<p>you can treat it as boxing box detection without box annotation. you only have image level annotation ground truth</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1f2f09e3a715b9b0f725e48ec843ddd1%2FSelection_050.png?generation=1578899240672583&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "you can treat it as boxing box detection without box annotation. you only have image level annotation ground truth\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1f2f09e3a715b9b0f725e48ec843ddd1%2FSelection_050.png?generation=1578899240672583&amp;alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 714171,
      "author_name": "mightyrains",
      "author_url": "",
      "post_date": "01/09/2020 06:24:26",
      "content": "<p>I guess the organizers provided a small public test dataset of 12 images to safeguard against PseudoLabelling based approaches. \nAt the moment of writing, the topmost entry on the public leaderboard is 0.9875. If they had provided a larger public test dataset, people here would've scored &gt;99.5% on the metric. </p>",
      "votes": null,
      "replies": [
        {
          "id": 714345,
          "author_name": "karan07",
          "author_url": "",
          "post_date": "01/09/2020 10:29:03",
          "content": "<p>The same thing got in my mind. We can use external dataset for the remaining dataset (e.g they used most common occuring 1000 out of 1350) to improve accuracy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 715941,
      "author_name": "moewie94",
      "author_url": "",
      "post_date": "01/11/2020 02:25:34",
      "content": "<p>yea i think it will work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 717474,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/13/2020 07:07:22",
      "content": "<p>you can treat it as boxing box detection without box annotation. you only have image level annotation ground truth</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1f2f09e3a715b9b0f725e48ec843ddd1%2FSelection_050.png?generation=1578899240672583&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "714166": "Since the models are pretty consistent when predicting for `public` test-set as shown [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/124454). I was thinking of using them as pseduolabels and train using `train` + `public test` set. Did anybody try this approach? Any luck?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F674454cee189a2ebdb00212a29a79fa6%2Fpseudo.png?generation=1578550435266754&amp;alt=media)",
    "714171": "I guess the organizers provided a small public test dataset of 12 images to safeguard against PseudoLabelling based approaches. \nAt the moment of writing, the topmost entry on the public leaderboard is 0.9875. If they had provided a larger public test dataset, people here would've scored &gt;99.5% on the metric.",
    "714345": "The same thing got in my mind. We can use external dataset for the remaining dataset (e.g they used most common occuring 1000 out of 1350) to improve accuracy.",
    "715941": "yea i think it will work",
    "717474": "you can treat it as boxing box detection without box annotation. you only have image level annotation ground truth\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1f2f09e3a715b9b0f725e48ec843ddd1%2FSelection_050.png?generation=1578899240672583&amp;alt=media)"
  },
  "source": "meta"
}