{
  "id": 107212,
  "title": "possible data leak way",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107212",
  "author_name": "",
  "post_date": "2019-09-03T03:34:27.020126300Z",
  "votes": -9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was wondering is there a possibility of data leak which some of the kagglers might have exploited to achieve good ranks.\n1) During run of submission kernel every one knows that private set images would be available on test images path.\nWont some people then could try to train the their model  based on   most confident predictions made and then   with their trained model they can again  do fresh predictions of pubic set and private set ?\nIf this can be a possibility i request organizer try to watch for this and </p>\n\n<p>2) Also some people also might tried psuedo labelling method where they label the public test images based on confident predictions  and then train the model back on those images.\nSince this model meant to take some critical decision so any such method that can make things biased any manner for professionals using the solution  should be strickly avoided.</p>",
  "messages": [
    {
      "id": "616359",
      "postDate": "09/03/2019 03:34:27",
      "content": "<p>I was wondering is there a possibility of data leak which some of the kagglers might have exploited to achieve good ranks.\n1) During run of submission kernel every one knows that private set images would be available on test images path.\nWont some people then could try to train the their model  based on   most confident predictions made and then   with their trained model they can again  do fresh predictions of pubic set and private set ?\nIf this can be a possibility i request organizer try to watch for this and </p>\n\n<p>2) Also some people also might tried psuedo labelling method where they label the public test images based on confident predictions  and then train the model back on those images.\nSince this model meant to take some critical decision so any such method that can make things biased any manner for professionals using the solution  should be strickly avoided.</p>",
      "rawMarkdown": "I was wondering is there a possibility of data leak which some of the kagglers might have exploited to achieve good ranks.\n1) During run of submission kernel every one knows that private set images would be available on test images path.\nWont some people then could try to train the their model  based on   most confident predictions made and then   with their trained model they can again  do fresh predictions of pubic set and private set ?\nIf this can be a possibility i request organizer try to watch for this and \n\n2) Also some people also might tried psuedo labelling method where they label the public test images based on confident predictions  and then train the model back on those images.\nSince this model meant to take some critical decision so any such method that can make things biased any manner for professionals using the solution  should be strickly avoided.",
      "votes": null
    },
    {
      "id": "616381",
      "postDate": "09/03/2019 05:00:34",
      "content": "<p>maybe you're wondering why you get downvoted ... it's simple - both cases (1) and (2) are valid / correct</p>\n\n<p>especially pseudo labelling is most likely used by everyone in top 10 (and that's why they \"overfit\" on the public test set), it's also standard practice when you have large test set</p>\n\n<p>point (1) - guess not many people will do this ... it's already hard to predict the private set in the time limit, don't think you could actually manage to do predict + train on pseudo labels of private + predict again (esp. when you consider that most teams will have a few models, and you can't train them in parallel)</p>",
      "rawMarkdown": "maybe you're wondering why you get downvoted ... it's simple - both cases (1) and (2) are valid / correct\n\nespecially pseudo labelling is most likely used by everyone in top 10 (and that's why they \"overfit\" on the public test set), it's also standard practice when you have large test set\n\npoint (1) - guess not many people will do this ... it's already hard to predict the private set in the time limit, don't think you could actually manage to do predict + train on pseudo labels of private + predict again (esp. when you consider that most teams will have a few models, and you can't train them in parallel)",
      "votes": null
    },
    {
      "id": "616404",
      "postDate": "09/03/2019 05:39:06",
      "content": "<p>I think pseudo labelling can help you impove your score in public, that maybe overfit public data, so should worry about pseudo labelling model would performance bad in private data. But there can commit two final score, one you can commit the pseudo labelling model,  the other you can commit without pseudo labelling model.</p>",
      "rawMarkdown": "I think pseudo labelling can help you impove your score in public, that maybe overfit public data, so should worry about pseudo labelling model would performance bad in private data. But there can commit two final score, one you can commit the pseudo labelling model,  the other you can commit without pseudo labelling model.",
      "votes": null
    },
    {
      "id": "616442",
      "postDate": "09/03/2019 06:19:13",
      "content": "<p>Wel i dont know why i got disliked :)\nyou have 9 hrs...\nit takes 30m-1hr to run   prediction once\nhere how is it can be\n1) you do get preds -  1 hr\n2)  Extract the confident predictions -no time\n4) build data-set out of it-No time\n5) Train for few epochs using test images as  your input path and pseudo labels as target- say got 5k images - 10 min per epoch \n6) Again do get preds  which is your final submission for  scoring...</p>\n\n<p>I am not sure will any one be able to do this.. but it always good there should be investigation done by kaggle team for such possible leak route..</p>\n\n<p>I m not telling this because m lagging behind but just that this solution is going to assist in diagnosing some ones critical eye ilness so any solution built should not be  aimed at  just scoring top ranks and prize ...\nThis striked into my mind when i was thinking about all possible means to push my rank up .. but  such ways are unethical for me so never explored but just thought of raising an alarm if such practice is adopted by any one...</p>",
      "rawMarkdown": "Wel i dont know why i got disliked :)\nyou have 9 hrs...\nit takes 30m-1hr to run   prediction once\nhere how is it can be\n1) you do get preds -  1 hr\n2)  Extract the confident predictions -no time\n4) build data-set out of it-No time\n5) Train for few epochs using test images as  your input path and pseudo labels as target- say got 5k images - 10 min per epoch \n6) Again do get preds  which is your final submission for  scoring...\n\nI am not sure will any one be able to do this.. but it always good there should be investigation done by kaggle team for such possible leak route..\n\nI m not telling this because m lagging behind but just that this solution is going to assist in diagnosing some ones critical eye ilness so any solution built should not be  aimed at  just scoring top ranks and prize ...\nThis striked into my mind when i was thinking about all possible means to push my rank up .. but  such ways are unethical for me so never explored but just thought of raising an alarm if such practice is adopted by any one...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 616381,
      "author_name": "steelrose",
      "author_url": "",
      "post_date": "09/03/2019 05:00:34",
      "content": "<p>maybe you're wondering why you get downvoted ... it's simple - both cases (1) and (2) are valid / correct</p>\n\n<p>especially pseudo labelling is most likely used by everyone in top 10 (and that's why they \"overfit\" on the public test set), it's also standard practice when you have large test set</p>\n\n<p>point (1) - guess not many people will do this ... it's already hard to predict the private set in the time limit, don't think you could actually manage to do predict + train on pseudo labels of private + predict again (esp. when you consider that most teams will have a few models, and you can't train them in parallel)</p>",
      "votes": null,
      "replies": [
        {
          "id": 616442,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/03/2019 06:19:13",
          "content": "<p>Wel i dont know why i got disliked :)\nyou have 9 hrs...\nit takes 30m-1hr to run   prediction once\nhere how is it can be\n1) you do get preds -  1 hr\n2)  Extract the confident predictions -no time\n4) build data-set out of it-No time\n5) Train for few epochs using test images as  your input path and pseudo labels as target- say got 5k images - 10 min per epoch \n6) Again do get preds  which is your final submission for  scoring...</p>\n\n<p>I am not sure will any one be able to do this.. but it always good there should be investigation done by kaggle team for such possible leak route..</p>\n\n<p>I m not telling this because m lagging behind but just that this solution is going to assist in diagnosing some ones critical eye ilness so any solution built should not be  aimed at  just scoring top ranks and prize ...\nThis striked into my mind when i was thinking about all possible means to push my rank up .. but  such ways are unethical for me so never explored but just thought of raising an alarm if such practice is adopted by any one...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 616404,
      "author_name": "jiangkun2",
      "author_url": "",
      "post_date": "09/03/2019 05:39:06",
      "content": "<p>I think pseudo labelling can help you impove your score in public, that maybe overfit public data, so should worry about pseudo labelling model would performance bad in private data. But there can commit two final score, one you can commit the pseudo labelling model,  the other you can commit without pseudo labelling model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "616359": "I was wondering is there a possibility of data leak which some of the kagglers might have exploited to achieve good ranks.\n1) During run of submission kernel every one knows that private set images would be available on test images path.\nWont some people then could try to train the their model  based on   most confident predictions made and then   with their trained model they can again  do fresh predictions of pubic set and private set ?\nIf this can be a possibility i request organizer try to watch for this and \n\n2) Also some people also might tried psuedo labelling method where they label the public test images based on confident predictions  and then train the model back on those images.\nSince this model meant to take some critical decision so any such method that can make things biased any manner for professionals using the solution  should be strickly avoided.",
    "616381": "maybe you're wondering why you get downvoted ... it's simple - both cases (1) and (2) are valid / correct\n\nespecially pseudo labelling is most likely used by everyone in top 10 (and that's why they \"overfit\" on the public test set), it's also standard practice when you have large test set\n\npoint (1) - guess not many people will do this ... it's already hard to predict the private set in the time limit, don't think you could actually manage to do predict + train on pseudo labels of private + predict again (esp. when you consider that most teams will have a few models, and you can't train them in parallel)",
    "616404": "I think pseudo labelling can help you impove your score in public, that maybe overfit public data, so should worry about pseudo labelling model would performance bad in private data. But there can commit two final score, one you can commit the pseudo labelling model,  the other you can commit without pseudo labelling model.",
    "616442": "Wel i dont know why i got disliked :)\nyou have 9 hrs...\nit takes 30m-1hr to run   prediction once\nhere how is it can be\n1) you do get preds -  1 hr\n2)  Extract the confident predictions -no time\n4) build data-set out of it-No time\n5) Train for few epochs using test images as  your input path and pseudo labels as target- say got 5k images - 10 min per epoch \n6) Again do get preds  which is your final submission for  scoring...\n\nI am not sure will any one be able to do this.. but it always good there should be investigation done by kaggle team for such possible leak route..\n\nI m not telling this because m lagging behind but just that this solution is going to assist in diagnosing some ones critical eye ilness so any solution built should not be  aimed at  just scoring top ranks and prize ...\nThis striked into my mind when i was thinking about all possible means to push my rank up .. but  such ways are unethical for me so never explored but just thought of raising an alarm if such practice is adopted by any one..."
  },
  "source": "meta"
}