{
  "id": 43599,
  "title": "Unsupervised learning on the test set",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43599",
  "author_name": "Asier",
  "post_date": "2017-11-16T17:21:51.998000",
  "votes": 12,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Are entries that use the test set for unsupervised learning allowed?</p>",
  "messages": [
    {
      "id": 244657,
      "postDate": "2017-11-16T17:21:51.997Z",
      "content": "<p>Are entries that use the test set for unsupervised learning allowed?</p>",
      "rawMarkdown": "Are entries that use the test set for unsupervised learning allowed?",
      "votes": 12
    },
    {
      "id": 248569,
      "postDate": "2017-11-26T14:01:58.123Z",
      "content": "<p>I did some experiments with using the test set for pseudo-labeling, but so far it did not seem to make my model any better (or worse).</p>",
      "rawMarkdown": "I did some experiments with using the test set for pseudo-labeling, but so far it did not seem to make my model any better (or worse).",
      "votes": 1
    },
    {
      "id": 256512,
      "postDate": "2017-12-12T05:09:47.553Z",
      "content": "<p>Refer to <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239</a></p>\n\n<p>Yes, unsupervised learning on the test set is allowed.</p>",
      "rawMarkdown": "Refer to https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239\n\nYes, unsupervised learning on the test set is allowed."
    },
    {
      "id": 247558,
      "postDate": "2017-11-23T11:00:47.040Z",
      "content": "<p>I don't know if it's allowed, but I don't think it should be. If that was the case, an unlabelled set for unsupervised learning should be provided. If you use the test set for unsupervised learning you will probably get a very good (unrealistic) score in Kaggle, which would probably not be the case for a new set of never seen examples.</p>",
      "rawMarkdown": "I don't know if it's allowed, but I don't think it should be. If that was the case, an unlabelled set for unsupervised learning should be provided. If you use the test set for unsupervised learning you will probably get a very good (unrealistic) score in Kaggle, which would probably not be the case for a new set of never seen examples.",
      "replies": [
        {
          "id": 247563,
          "postDate": "2017-11-23T11:19:09.510Z",
          "content": "<p>I disagree on both points. I doubt unsupervised learning on the test set would make it so easy to get a very good score on the leaderboard, though if done well it is likely to help. And I think unsupervised (or semi-supervised) learning would improve generalisation to never-seen examples, assuming they are from roughly the same distribution as the training and test sets.</p>",
          "rawMarkdown": "I disagree on both points. I doubt unsupervised learning on the test set would make it so easy to get a very good score on the leaderboard, though if done well it is likely to help. And I think unsupervised (or semi-supervised) learning would improve generalisation to never-seen examples, assuming they are from roughly the same distribution as the training and test sets."
        },
        {
          "id": 247565,
          "postDate": "2017-11-23T11:24:50.050Z",
          "content": "<p>I've used Ladder Networks (semi-supervised learning) in the past and it helps getting better results. Some people are also using semi-supervised learning with GANs with success in some problems, so I think it's very likely that someone would get high scores on Kaggle if they used the test set, since it would be the same set you used for the unsupervised part of training.</p>",
          "rawMarkdown": "I've used Ladder Networks (semi-supervised learning) in the past and it helps getting better results. Some people are also using semi-supervised learning with GANs with success in some problems, so I think it's very likely that someone would get high scores on Kaggle if they used the test set, since it would be the same set you used for the unsupervised part of training.",
          "votes": 3
        },
        {
          "id": 247601,
          "postDate": "2017-11-23T11:51:31.030Z",
          "content": "<p>Yep I agree it will help. Hence the judges should weigh in on whether it's allowed!</p>",
          "rawMarkdown": "Yep I agree it will help. Hence the judges should weigh in on whether it's allowed!",
          "votes": 1
        }
      ]
    },
    {
      "id": 247086,
      "postDate": "2017-11-22T09:32:47.847Z",
      "content": "<p>Would be good if we could get a definitive answer from the judges on this one as I think it could help a lot.</p>",
      "rawMarkdown": "Would be good if we could get a definitive answer from the judges on this one as I think it could help a lot."
    },
    {
      "id": 246165,
      "postDate": "2017-11-20T16:10:57.507Z",
      "content": "<p>I would also like to know if it is allowed.</p>",
      "rawMarkdown": "I would also like to know if it is allowed."
    },
    {
      "id": 244885,
      "postDate": "2017-11-17T04:40:35.373Z",
      "content": "<p>That sounds an interesting question. </p>\n\n<p>To me it seems it does not matter though... as long as the result submission format is as instructed.</p>",
      "rawMarkdown": "That sounds an interesting question. \n\nTo me it seems it does not matter though... as long as the result submission format is as instructed."
    },
    {
      "id": 247260,
      "postDate": "2017-11-22T17:12:40.830Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 248569,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2017-11-26T14:01:58.123000",
      "content": "<p>I did some experiments with using the test set for pseudo-labeling, but so far it did not seem to make my model any better (or worse).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 256512,
      "author_name": "Steven Du",
      "author_url": "",
      "post_date": "2017-12-12T05:09:47.553000",
      "content": "<p>Refer to <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239</a></p>\n\n<p>Yes, unsupervised learning on the test set is allowed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 247558,
      "author_name": "Rafael Barbolo",
      "author_url": "",
      "post_date": "2017-11-23T11:00:47.040000",
      "content": "<p>I don't know if it's allowed, but I don't think it should be. If that was the case, an unlabelled set for unsupervised learning should be provided. If you use the test set for unsupervised learning you will probably get a very good (unrealistic) score in Kaggle, which would probably not be the case for a new set of never seen examples.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 247563,
          "author_name": "Liam",
          "author_url": "",
          "post_date": "2017-11-23T11:19:09.510000",
          "content": "<p>I disagree on both points. I doubt unsupervised learning on the test set would make it so easy to get a very good score on the leaderboard, though if done well it is likely to help. And I think unsupervised (or semi-supervised) learning would improve generalisation to never-seen examples, assuming they are from roughly the same distribution as the training and test sets.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 247565,
          "author_name": "Rafael Barbolo",
          "author_url": "",
          "post_date": "2017-11-23T11:24:50.050000",
          "content": "<p>I've used Ladder Networks (semi-supervised learning) in the past and it helps getting better results. Some people are also using semi-supervised learning with GANs with success in some problems, so I think it's very likely that someone would get high scores on Kaggle if they used the test set, since it would be the same set you used for the unsupervised part of training.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 247601,
          "author_name": "Liam",
          "author_url": "",
          "post_date": "2017-11-23T11:51:31.030000",
          "content": "<p>Yep I agree it will help. Hence the judges should weigh in on whether it's allowed!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 247086,
      "author_name": "Liam",
      "author_url": "",
      "post_date": "2017-11-22T09:32:47.847000",
      "content": "<p>Would be good if we could get a definitive answer from the judges on this one as I think it could help a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 246165,
      "author_name": "Cadu",
      "author_url": "",
      "post_date": "2017-11-20T16:10:57.507000",
      "content": "<p>I would also like to know if it is allowed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 244885,
      "author_name": "IsaacSim",
      "author_url": "",
      "post_date": "2017-11-17T04:40:35.373000",
      "content": "<p>That sounds an interesting question. </p>\n\n<p>To me it seems it does not matter though... as long as the result submission format is as instructed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 247260,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-11-22T17:12:40.830000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "244657": "Are entries that use the test set for unsupervised learning allowed?",
    "248569": "I did some experiments with using the test set for pseudo-labeling, but so far it did not seem to make my model any better (or worse).",
    "256512": "Refer to https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239\n\nYes, unsupervised learning on the test set is allowed.",
    "247558": "I don't know if it's allowed, but I don't think it should be. If that was the case, an unlabelled set for unsupervised learning should be provided. If you use the test set for unsupervised learning you will probably get a very good (unrealistic) score in Kaggle, which would probably not be the case for a new set of never seen examples.",
    "247086": "Would be good if we could get a definitive answer from the judges on this one as I think it could help a lot.",
    "246165": "I would also like to know if it is allowed.",
    "244885": "That sounds an interesting question. \n\nTo me it seems it does not matter though... as long as the result submission format is as instructed.",
    "247260": ""
  }
}