{
  "id": 1183,
  "title": "devel*_test.csv vs. devel*_train.csv Question",
  "url": "/competitions/GestureChallenge/discussion/1183",
  "author_name": "",
  "post_date": "2011-12-20T05:30:50.057Z",
  "votes": null,
  "comment_count": 5,
  "views": 2951,
  "content": "<p>Hello,&nbsp;</p>\r\n<p>I was just looking at the data and noticed these two .csv files.&nbsp;</p>\r\n<p>My question is: Are we only allowed to train our classifier with the examples specified in devel*_train.csv? The reason why I'm asking is that the examples in devel*_test.csv seem to combine multiple gestures and also repeat the same gestures for different\r\n data points, which I thought defeated the purpose of the competition.</p>\r\n<p>The examples in devel*_train.csv are for testing/cross-validation?</p>\r\n<p>I think that our classifier is supposed to use each gesture only once - that's why I was a little confused.&nbsp;</p>",
  "messages": [
    {
      "id": "7359",
      "postDate": "12/20/2011 05:30:50",
      "content": "<p>Hello,&nbsp;</p>\r\n<p>I was just looking at the data and noticed these two .csv files.&nbsp;</p>\r\n<p>My question is: Are we only allowed to train our classifier with the examples specified in devel*_train.csv? The reason why I'm asking is that the examples in devel*_test.csv seem to combine multiple gestures and also repeat the same gestures for different\r\n data points, which I thought defeated the purpose of the competition.</p>\r\n<p>The examples in devel*_train.csv are for testing/cross-validation?</p>\r\n<p>I think that our classifier is supposed to use each gesture only once - that's why I was a little confused.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7362",
      "postDate": "12/20/2011 07:01:58",
      "content": "<p>This competition consists of three main datasets, two of which have been released. &nbsp;These are:</p>\r\n<ul>\r\n<li>Development Set (Currently Released)\r\n<ul>\r\n<li>devel01-devel20 folders (additional development data will be released January 7)\r\n</li><li>each folder corresponds to a different person </li><li>This set is used to develop your models </li><li>Contains videos identifying unique gestures (as labeled in develXX_train.csv)\r\n</li><li>Contains videos with sequences of one or more gestures (as labeled in develXX_test.csv)\r\n</li></ul>\r\n</li><li>Validation Set (Currently Relased)\r\n<ul>\r\n<li>valid01-valid20 folders </li><li>each folder corresponds to a different person </li><li>This data is used to form the public leaderboard </li><li>Video sequences identifying unique getures are labeled in validXX_train.csv </li><li>The remaining video sequences contain one or more gestures, which should be identified by your model (and then you should upload your predictions to Kaggle)\r\n</li></ul>\r\n</li><li>Final Evaluation Set (to be released on April 7, 2012)\r\n<ul>\r\n<li>this set will be in the same format as the validation set, only with different folder names\r\n</li><li>the results of predictions on this set will be used for the final leaderboard\r\n</li></ul>\r\n</li></ul>\r\n<div>Thanks for your interest in the competition, and let me know if you have any further questions!</div>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7383",
      "postDate": "12/20/2011 19:28:34",
      "content": "<p>[quote] Video sequences identifying unique getures are labeled in validXX_train.csv [/quote]</p>\r\n<p>For the valid folders, are you supposed to train your model with the unique gestures labeled in validXX_train.csv and then predict the gestures for the rest of the data?&nbsp;</p>\r\n<p>I think that's what you are supposed to do but I just wanted to make sure.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7384",
      "postDate": "12/20/2011 19:31:11",
      "content": "<p>[quote=CihanBaran;7383]</p>\r\n<p>[quote] Video sequences identifying unique getures are labeled in validXX_train.csv [/quote]</p>\r\n<p>For the valid folders, are you supposed to train your model with the unique gestures labeled in validXX_train.csv and then predict the gestures for the rest of the data?&nbsp;</p>\r\n<p>I think that's what you are supposed to do but I just wanted to make sure.&nbsp;</p>\r\n<p>[/quote]</p>\r\nYes, that's correct",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7413",
      "postDate": "12/21/2011 16:54:28",
      "content": "<p>[quote=Ben Hamner;7384]</p>\r\n<p>[quote=CihanBaran;7383]</p>\r\n<p>[quote] Video sequences identifying unique getures are labeled in validXX_train.csv [/quote]</p>\r\n<p>For the valid folders, are you supposed to train your model with the unique gestures labeled in validXX_train.csv and then predict the gestures for the rest of the data?&nbsp;</p>\r\n<p>I think that's what you are supposed to do but I just wanted to make sure.&nbsp;</p>\r\n<p>[/quote]</p>\r\n<p>Yes, that's correct</p>\r\n<p>[/quote]</p>\r\n<p>==&gt; I confirm that for the validation data batches as well as for the final evaluation batches, you have only one labeled example of video of each gesture to train.&nbsp;</p>\r\n<p>==&gt; The development data cannot be used as extra labeled data for the validation and fnal evaluation tasks. However, they can be used to train a preprocessor, in the spirit of &quot;transfer learning&quot;. For example, you can develop features using the development\r\n data that you can then use to solve the validation and final evaluation tasks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7475",
      "postDate": "12/23/2011 22:54:54",
      "content": "<p>I think the whole point of having a validation set is to be able to train the model with training set or development set first and then decide upon the best bias or variance factors by drawing learning curves over the validation set. I hope I am getting\r\n your question correctly. Sorry if I am making it sound more confusing.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 7362,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "12/20/2011 07:01:58",
      "content": "<p>This competition consists of three main datasets, two of which have been released. &nbsp;These are:</p>\r\n<ul>\r\n<li>Development Set (Currently Released)\r\n<ul>\r\n<li>devel01-devel20 folders (additional development data will be released January 7)\r\n</li><li>each folder corresponds to a different person </li><li>This set is used to develop your models </li><li>Contains videos identifying unique gestures (as labeled in develXX_train.csv)\r\n</li><li>Contains videos with sequences of one or more gestures (as labeled in develXX_test.csv)\r\n</li></ul>\r\n</li><li>Validation Set (Currently Relased)\r\n<ul>\r\n<li>valid01-valid20 folders </li><li>each folder corresponds to a different person </li><li>This data is used to form the public leaderboard </li><li>Video sequences identifying unique getures are labeled in validXX_train.csv </li><li>The remaining video sequences contain one or more gestures, which should be identified by your model (and then you should upload your predictions to Kaggle)\r\n</li></ul>\r\n</li><li>Final Evaluation Set (to be released on April 7, 2012)\r\n<ul>\r\n<li>this set will be in the same format as the validation set, only with different folder names\r\n</li><li>the results of predictions on this set will be used for the final leaderboard\r\n</li></ul>\r\n</li></ul>\r\n<div>Thanks for your interest in the competition, and let me know if you have any further questions!</div>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7383,
      "author_name": "cihanb",
      "author_url": "",
      "post_date": "12/20/2011 19:28:34",
      "content": "<p>[quote] Video sequences identifying unique getures are labeled in validXX_train.csv [/quote]</p>\r\n<p>For the valid folders, are you supposed to train your model with the unique gestures labeled in validXX_train.csv and then predict the gestures for the rest of the data?&nbsp;</p>\r\n<p>I think that's what you are supposed to do but I just wanted to make sure.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7384,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "12/20/2011 19:31:11",
      "content": "<p>[quote=CihanBaran;7383]</p>\r\n<p>[quote] Video sequences identifying unique getures are labeled in validXX_train.csv [/quote]</p>\r\n<p>For the valid folders, are you supposed to train your model with the unique gestures labeled in validXX_train.csv and then predict the gestures for the rest of the data?&nbsp;</p>\r\n<p>I think that's what you are supposed to do but I just wanted to make sure.&nbsp;</p>\r\n<p>[/quote]</p>\r\nYes, that's correct",
      "votes": null,
      "replies": []
    },
    {
      "id": 7413,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "12/21/2011 16:54:28",
      "content": "<p>[quote=Ben Hamner;7384]</p>\r\n<p>[quote=CihanBaran;7383]</p>\r\n<p>[quote] Video sequences identifying unique getures are labeled in validXX_train.csv [/quote]</p>\r\n<p>For the valid folders, are you supposed to train your model with the unique gestures labeled in validXX_train.csv and then predict the gestures for the rest of the data?&nbsp;</p>\r\n<p>I think that's what you are supposed to do but I just wanted to make sure.&nbsp;</p>\r\n<p>[/quote]</p>\r\n<p>Yes, that's correct</p>\r\n<p>[/quote]</p>\r\n<p>==&gt; I confirm that for the validation data batches as well as for the final evaluation batches, you have only one labeled example of video of each gesture to train.&nbsp;</p>\r\n<p>==&gt; The development data cannot be used as extra labeled data for the validation and fnal evaluation tasks. However, they can be used to train a preprocessor, in the spirit of &quot;transfer learning&quot;. For example, you can develop features using the development\r\n data that you can then use to solve the validation and final evaluation tasks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7475,
      "author_name": "anup13958",
      "author_url": "",
      "post_date": "12/23/2011 22:54:54",
      "content": "<p>I think the whole point of having a validation set is to be able to train the model with training set or development set first and then decide upon the best bias or variance factors by drawing learning curves over the validation set. I hope I am getting\r\n your question correctly. Sorry if I am making it sound more confusing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "7359": "",
    "7362": "",
    "7383": "",
    "7384": "",
    "7413": "",
    "7475": ""
  },
  "source": "meta"
}