{
  "id": 5146,
  "title": "people in the test set",
  "url": "/competitions/multi-modal-gesture-recognition/discussion/5146",
  "author_name": "",
  "post_date": "2013-07-18T23:16:36.420Z",
  "votes": null,
  "comment_count": 3,
  "views": 1113,
  "content": "<p>While I haven't done a systematic check, it appears to me that the people in the Training set are distinct from the people in the Validation set. Could one of the competition admins tell us whether the Test set will contain:</p>\r\n<p></p>\r\n<p>a) only people from the Training set,</p>\r\n<p></p>\r\n<p>b) only people from the validation set,</p>\r\n<p></p>\r\n<p>c) a roughly 50/50 mix of people from both the Training and Validation sets, or</p>\r\n<p></p>\r\n<p>d) something else?</p>\r\n<p></p>\r\n<p>Knowing this will affect the training strategy, without taking away from the challenge of the competition. &nbsp;thanks,&nbsp;<span>wweight</span></p>\r\n<p></p>",
  "messages": [
    {
      "id": "27429",
      "postDate": "07/18/2013 23:16:36",
      "content": "<p>While I haven't done a systematic check, it appears to me that the people in the Training set are distinct from the people in the Validation set. Could one of the competition admins tell us whether the Test set will contain:</p>\r\n<p></p>\r\n<p>a) only people from the Training set,</p>\r\n<p></p>\r\n<p>b) only people from the validation set,</p>\r\n<p></p>\r\n<p>c) a roughly 50/50 mix of people from both the Training and Validation sets, or</p>\r\n<p></p>\r\n<p>d) something else?</p>\r\n<p></p>\r\n<p>Knowing this will affect the training strategy, without taking away from the challenge of the competition. &nbsp;thanks,&nbsp;<span>wweight</span></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27474",
      "postDate": "07/20/2013 15:30:23",
      "content": "<p>Dear wweight,</p>\n<p>the final evaluation data (Test data) will contain people appeared in Training and Validation sets. Take into account that labels for the Validation data will be provided together with the final evaluation data, therefore, you will be able to use both, training and validation data sets in order to train your methods.</p>\n<p>Thanks.</p>\n<p>Xavier</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27493",
      "postDate": "07/21/2013 18:53:19",
      "content": "<p>Thanks for the response&nbsp;Xavier. &nbsp;It's probably too late to change, but here are some arguments against the Test set including only people that we've trained on.</p>\n<p>&nbsp;</p>\n<p>First, this will make the test phase quite different from the initial phase. &nbsp;In the initial phase, where the Training and Validation people are distinct, there is an incentive to create a model that will generalize well to people that are not in the Training set. &nbsp;In the second, final phase, however, there will be no such incentive, and the algorithms and the results might look quite different between these two phases. &nbsp;For example, I estimate that, if my current model were being tested on Training people, then my score would improve from ~0.5 to ~0.2. &nbsp;Such large changes in the incentives and scores between two phases of a competition are a little jarring and unpleasant, although I admit that you have given us plenty of warning time.</p>\n<p>&nbsp;</p>\n<p><span style=\"line-height: 1.4\">The second, similar point is that this competition will reward the models that can best&nbsp;</span>recognize<span style=\"line-height: 1.4\">&nbsp;gestures of people it has already seen, rather than new people, and will therefore have less &quot;real-world&quot; utility. &nbsp;If the current rules continue, I'd be very surprised if the winner doesn't score lower than 0.1, whereas the&nbsp;</span>generalization<span style=\"line-height: 1.4\">&nbsp;error would probably be significantly higher.</span></p>\n<p>&nbsp;</p>\n<p><span style=\"line-height: 1.4\">Two possibilities are (1) do not release the labels for the Validation set, and restrict the Test set to be only people from the Validation set, or (2) release the labels for the Validation set, but use a third, distinct set of people for the Test set (unlikely that such data exists, I know).</span></p>\n<p>&nbsp;</p>\n<p><span style=\"line-height: 1.4\">I may have misunderstood the aim of the competition. &nbsp;Maybe recognizing the gestures of people that have already been seen, rather than new people, is exactly what you're after, in which case the current rules are perfect. &nbsp;Anyway, thanks for a fun competition. &nbsp;-wweight</span></p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27653",
      "postDate": "07/25/2013 16:02:36",
      "content": "<p>Hi wweight,</p>\n<p>I am puzzled by your comment that a score of lower than 0.1 is expected by the winning team. After looking at the video and the gestures, my humble opinion is that accurate classification is not easy at all. This is because of two reasons: (1) many gestures are very similar visually, e.g. raising of the hand and small waving of the palm; (2) many gestures are done very quickly, e.g. less than 20 frames.</p>\n<p>My understanding of the normalized edit distance score is that it is a kind of error metric between the ground truth and the predicted labels. Given that there are 10 gestures in a test video, a score of 0.1 would mean that there is at most only 1 error (can be 1 insertion, 1 deletion, but no more than that). To me this is very difficult. Do I remember correctly that the state of art for an &quot;easy&quot; dataset such as the KTH (only 6 actions, and each input video is with frames and hence rich with features) is only around 95%?</p>\n<p>Thanks for any feedback and comments.</p>\n<p>telepoints</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 27474,
      "author_name": "xbaro100272",
      "author_url": "",
      "post_date": "07/20/2013 15:30:23",
      "content": "<p>Dear wweight,</p>\n<p>the final evaluation data (Test data) will contain people appeared in Training and Validation sets. Take into account that labels for the Validation data will be provided together with the final evaluation data, therefore, you will be able to use both, training and validation data sets in order to train your methods.</p>\n<p>Thanks.</p>\n<p>Xavier</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27493,
      "author_name": "rkeisler",
      "author_url": "",
      "post_date": "07/21/2013 18:53:19",
      "content": "<p>Thanks for the response&nbsp;Xavier. &nbsp;It's probably too late to change, but here are some arguments against the Test set including only people that we've trained on.</p>\n<p>&nbsp;</p>\n<p>First, this will make the test phase quite different from the initial phase. &nbsp;In the initial phase, where the Training and Validation people are distinct, there is an incentive to create a model that will generalize well to people that are not in the Training set. &nbsp;In the second, final phase, however, there will be no such incentive, and the algorithms and the results might look quite different between these two phases. &nbsp;For example, I estimate that, if my current model were being tested on Training people, then my score would improve from ~0.5 to ~0.2. &nbsp;Such large changes in the incentives and scores between two phases of a competition are a little jarring and unpleasant, although I admit that you have given us plenty of warning time.</p>\n<p>&nbsp;</p>\n<p><span style=\"line-height: 1.4\">The second, similar point is that this competition will reward the models that can best&nbsp;</span>recognize<span style=\"line-height: 1.4\">&nbsp;gestures of people it has already seen, rather than new people, and will therefore have less &quot;real-world&quot; utility. &nbsp;If the current rules continue, I'd be very surprised if the winner doesn't score lower than 0.1, whereas the&nbsp;</span>generalization<span style=\"line-height: 1.4\">&nbsp;error would probably be significantly higher.</span></p>\n<p>&nbsp;</p>\n<p><span style=\"line-height: 1.4\">Two possibilities are (1) do not release the labels for the Validation set, and restrict the Test set to be only people from the Validation set, or (2) release the labels for the Validation set, but use a third, distinct set of people for the Test set (unlikely that such data exists, I know).</span></p>\n<p>&nbsp;</p>\n<p><span style=\"line-height: 1.4\">I may have misunderstood the aim of the competition. &nbsp;Maybe recognizing the gestures of people that have already been seen, rather than new people, is exactly what you're after, in which case the current rules are perfect. &nbsp;Anyway, thanks for a fun competition. &nbsp;-wweight</span></p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27653,
      "author_name": "kongwah",
      "author_url": "",
      "post_date": "07/25/2013 16:02:36",
      "content": "<p>Hi wweight,</p>\n<p>I am puzzled by your comment that a score of lower than 0.1 is expected by the winning team. After looking at the video and the gestures, my humble opinion is that accurate classification is not easy at all. This is because of two reasons: (1) many gestures are very similar visually, e.g. raising of the hand and small waving of the palm; (2) many gestures are done very quickly, e.g. less than 20 frames.</p>\n<p>My understanding of the normalized edit distance score is that it is a kind of error metric between the ground truth and the predicted labels. Given that there are 10 gestures in a test video, a score of 0.1 would mean that there is at most only 1 error (can be 1 insertion, 1 deletion, but no more than that). To me this is very difficult. Do I remember correctly that the state of art for an &quot;easy&quot; dataset such as the KTH (only 6 actions, and each input video is with frames and hence rich with features) is only around 95%?</p>\n<p>Thanks for any feedback and comments.</p>\n<p>telepoints</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "27429": "",
    "27474": "",
    "27493": "",
    "27653": ""
  },
  "source": "meta"
}