{
  "id": 1150,
  "title": "Getting started",
  "url": "/competitions/GestureChallenge/discussion/1150",
  "author_name": "",
  "post_date": "2011-12-12T19:21:45.390Z",
  "votes": null,
  "comment_count": 30,
  "views": 16705,
  "content": "<p><span><strong>I added an entry to the Help to help you getting started:</strong></span></p>\r\n<p><span>There are essentially 2 approches that can be taken for data representation:</span></p>\r\n<ul>\r\n<li>Extracting a &quot;bag&quot; of low level spatio-temporal features. This approach is often taken by the researchers working on activity recognition. An example is the&nbsp;<a href=\"http://www.irisa.fr/vista/actions/\">bag of STIP features</a>.\r\n</li><li>Tracking the position of body parts. This approach is used in most games. One popular method was introduced by Microsoft with their&nbsp;<a href=\"http://research.microsoft.com/apps/pubs/default.aspx?id=145347\">skeleton tracker</a>, which is part of their&nbsp;<a href=\"http://kinectforwindows.org/download/\">SDK</a>.\r\n</li></ul>\r\n<p>There is&nbsp;<a href=\"http://www.amazon.com/Visual-Analysis-Humans-Looking-People/dp/0857299964/ref=sr_1_1?ie=UTF8&amp;qid=1321321690&amp;sr=8-1\">an excellent book</a>&nbsp;on the subject.</p>\r\n<p>Some approaches require separating the gesture sequences into isolated gestures first, which is relatively easy in this dataset because the users return their hands to a resting position between gestures.&nbsp;Once you have a vector representation of isolated\r\n gestures, to do the &quot;one-shot-learning&quot;, the simplest method is the&nbsp;<a href=\"http://en.wikipedia.org/wiki/K-nearest_neighbor_algorithm\">nearest neighbor</a>&nbsp;method. But you may also look for the best match between temporal sequences directly without isolating\r\n gestures using&nbsp;<a href=\"http://en.wikipedia.org/wiki/Dynamic_time_warping\">dynamic time warping</a>.</p>\r\n<p>Isabelle</p>",
  "messages": [
    {
      "id": "7092",
      "postDate": "12/12/2011 19:21:45",
      "content": "<p><span><strong>I added an entry to the Help to help you getting started:</strong></span></p>\r\n<p><span>There are essentially 2 approches that can be taken for data representation:</span></p>\r\n<ul>\r\n<li>Extracting a &quot;bag&quot; of low level spatio-temporal features. This approach is often taken by the researchers working on activity recognition. An example is the&nbsp;<a href=\"http://www.irisa.fr/vista/actions/\">bag of STIP features</a>.\r\n</li><li>Tracking the position of body parts. This approach is used in most games. One popular method was introduced by Microsoft with their&nbsp;<a href=\"http://research.microsoft.com/apps/pubs/default.aspx?id=145347\">skeleton tracker</a>, which is part of their&nbsp;<a href=\"http://kinectforwindows.org/download/\">SDK</a>.\r\n</li></ul>\r\n<p>There is&nbsp;<a href=\"http://www.amazon.com/Visual-Analysis-Humans-Looking-People/dp/0857299964/ref=sr_1_1?ie=UTF8&amp;qid=1321321690&amp;sr=8-1\">an excellent book</a>&nbsp;on the subject.</p>\r\n<p>Some approaches require separating the gesture sequences into isolated gestures first, which is relatively easy in this dataset because the users return their hands to a resting position between gestures.&nbsp;Once you have a vector representation of isolated\r\n gestures, to do the &quot;one-shot-learning&quot;, the simplest method is the&nbsp;<a href=\"http://en.wikipedia.org/wiki/K-nearest_neighbor_algorithm\">nearest neighbor</a>&nbsp;method. But you may also look for the best match between temporal sequences directly without isolating\r\n gestures using&nbsp;<a href=\"http://en.wikipedia.org/wiki/Dynamic_time_warping\">dynamic time warping</a>.</p>\r\n<p>Isabelle</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7094",
      "postDate": "12/13/2011 05:37:51",
      "content": "<p>For the prize, how are you going you deal with equal scores? I mean, what if the top 10 all get the same score (like everybody get 1 wrong, or even 0 wrong prediction). I feel this problem can be done with very very high accuracy, as high as a human can\r\n do.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7097",
      "postDate": "12/13/2011 07:18:44",
      "content": "<p>Is it allowed to use the skeleton tracker from Microsoft to preprocess the data or would we have to write our own tracker?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7125",
      "postDate": "12/14/2011 10:03:34",
      "content": "<p>[quote=woshialex;7094]</p>\r\n<p>I feel this problem can be done with very very high accuracy, as high as a human can do.</p>\r\n<p>[/quote]</p>\r\n<p>I am not sure that state-of-the-art techniques in gesture recognition can perform as accurately as a human can do...</p>\r\n<p>Anyway, do we have to perform better than the human performance benchmark to be elligible for the prize?</p>\r\n<p>It could be a good idea to submit a state-of-the-art benchmark rather than just a result of a template matching...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7149",
      "postDate": "12/14/2011 21:29:46",
      "content": "<p>There is no requirement to do better than human performance to win the prizes. The human performance is provided for information only. It is based on the results of a single person. Most errors are due to inattention and do not correspond to intrinsically\r\n ambiguous gestures. Hence, it is possible to do better than human performance.</p>\r\n<p>The organizers will post new benchmark results from time to time.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7151",
      "postDate": "12/14/2011 21:34:48",
      "content": "<p>You can use any other third party software, as long as you have the right to use such software. There is no contraint on the originality of the code to enter the challenge and win the prizes. However, if you are interested in licensing your algorithms if\r\n you win, originality will play a role. Microsoft will review the submissions of interested participants and grant up to $100,000 in contracts. See the rules.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7152",
      "postDate": "12/14/2011 21:35:56",
      "content": "<p>For the ties, it is stipulated in the rules:</p>\r\n<p>In the event of a tie between any eligible entries, the tie will be broken in the quantitative evaluation by giving preference to the earliest submission.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7162",
      "postDate": "12/15/2011 02:06:40",
      "content": "<p>Isabelle,</p>\r\n<p>Could you elaborate on this rule?</p>\r\n<p style=\"padding-left:30px\">In addition, to enter the qualitative evaluation, the participants will be asked to certify that they use a Kinect sensor and the Microsoft Software Development Kit (SDK).\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7508",
      "postDate": "12/26/2011 06:41:36",
      "content": "<p>Bonjour :-)</p>\r\n<p>In addition to Jose's question, I have another question before I start coding. I suppose I should avoid using GPL-licensed libraries since they are not Microsoft-friendly?</p>\r\n<p>Thanks,</p>\r\n<p>Ali</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7520",
      "postDate": "12/27/2011 16:32:46",
      "content": "<p>Dear Isabelle,</p>\r\n<p>Thank you for your tips.</p>\r\n<p>Could you elaborate or direct to information on how to read the video into Microsoft SDK?</p>\r\n<p>&nbsp;</p>\r\n<p>I would like to use its skeleton output.</p>\r\n<p>&nbsp;</p>\r\n<p>thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7524",
      "postDate": "12/28/2011 00:09:23",
      "content": "<p>[quote=shuky;7520]</p>\r\n<p>Dear Isabelle,</p>\r\n<p>Thank you for your tips.</p>\r\n<p>Could you elaborate or direct to information on how to read the video into Microsoft SDK?</p>\r\n<p>&nbsp;</p>\r\n<p>I would like to use its skeleton output.</p>\r\n<p>&nbsp;</p>\r\n<p>thanks</p>\r\n<p>[/quote]</p>\r\n<p>As far as I know, this is not trivial to do. We are working on providing you with skeleton data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7525",
      "postDate": "12/28/2011 00:27:33",
      "content": "<p>[quote=Ali Hassaïne;7508]</p>\r\n<p>Bonjour :-)</p>\r\n<p>In addition to Jose's question, I have another question before I start coding. I suppose I should avoid using GPL-licensed libraries since they are not Microsoft-friendly?</p>\r\n<p>Thanks,</p>\r\n<p>Ali</p>\r\n<p>[/quote]</p>\r\n<p>My recommendation is to go by the official rules, which imply that:</p>\r\n<p>- You may use any third party software that you have a legal right to use for the present challenge (first quantitative evaluation), because it uses pre-recorded data.</p>\r\n<p>- For the demonstration competition (qualitative evaluation) you will have to use Microsoft hardware (Kinect camera) and the drivers provided by Microsoft with their development kit.</p>\r\n<p>- If you are interested in potentially licensing your code to Microsoft (and benefit from the opportunity offered by Microsoft who may license code for up to $100,000 in this challenge), you need to have some original code of your own. Obviously you may\r\n also use third party libraries that you have a legal right to use and that Microsoft will be allowed to license. But your own original code is what will count. It is possible (and even likely) that Microsoft will prefer code running on a Microsoft platform.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7526",
      "postDate": "12/28/2011 00:31:40",
      "content": "<p>[quote=Jose H. Solorzano;7162]</p>\r\n<p>Isabelle,</p>\r\n<p>Could you elaborate on this rule?</p>\r\n<p style=\"padding-left:30px\">In addition, to enter the qualitative evaluation, the participants will be asked to certify that they use a Kinect sensor and the Microsoft Software Development Kit (SDK).</p>\r\n<p>[/quote]</p>\r\n<p>Microsoft wants to make sure that all the demonstrators in the demonstration competition (qualitative evaluation) use Microsoft hardware and software for motion capture. See my more detailed answer to Ali. This does not concern the present challenge, which\r\n uses pre-recorded data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7555",
      "postDate": "12/29/2011 03:09:34",
      "content": "I've been reviewing the related rules.\r\n<p>Apparently there are 2 conferences: CVPR and ICPR. </p>\r\n<p>I have no idea what these are about, as their web pages are opaque :0 </p>\r\n<p>For each contest there are 2 challenges: Quantitative and Qualitative (4 challenges total).\r\n</p>\r\n<p>The quantitative challenge is where your program munches on pre-recorded videos and attempts to decipher gestures.\r\n</p>\r\n<p>The qualitative challenge is where your program runs live at a conference, taking video directly from a Kinect and using Microsoft's API. I haven't read the details yet (perhaps they haven't been finalized), but it looks like your program will have to recognize\r\n gestures live in competition with other challengers.</p>\r\n<p>For each of the 4 challenges, the prizes are the same: $5K/$3K/$2K for 1st, 2nd, 3rd place.\r\n</p>\r\n<p>It would appear that for the quantitative challenge, all you need is a program that can read videos and generate a csv file.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7572",
      "postDate": "12/30/2011 03:42:39",
      "content": "<p>[quote=shuky;7520]</p>\r\n<p>Dear Isabelle,</p>\r\n<p>Thank you for your tips.</p>\r\n<p>Could you elaborate or direct to information on how to read the video into Microsoft SDK?</p>\r\n<p>I would like to use its skeleton output.</p>\r\n<p>thanks</p>\r\n<p>[/quote]</p>\r\n<p>I think most people are waiting for the skeleton output, its probably the best foundation to attack this problem with.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7578",
      "postDate": "12/30/2011 17:48:08",
      "content": "<p>From the description in their SDK (here for example: http://kinectforwindows.org/documents/SkeletalViewer_Walkthrough.pdf ), it seems that the Kinect skeleton tracker can only detect joints if the entire body is present within the frame, which is not true\r\n for the samples in this contest. So I don't think it's directly usable here. Perhaps this is why we have this competition?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "7607",
      "postDate": "01/01/2012 09:44:36",
      "content": "<p><span class=\"x_profilelink\">Isabelle,</span></p>\r\n<p>&nbsp;</p>\r\n<p>Could you elaborate more on the Skeleton data - </p>\r\n<p>When will it be available?</p>\r\n<p>Is there skeleton data for the videos with persons which are sitting (from what I understand skeleton view can only be created when the entire body can be seen)?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8046",
      "postDate": "01/10/2012 19:04:55",
      "content": "<p>Isabelle,</p>\r\n<p>Are there any rules against modifying the Kinect with additional hardware?<br>\r\nAlso will you give us a new batch of gestures to train on for final evaluation and then validating those? Or will we be validating data that we have previously trained on earlier in the competition?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8051",
      "postDate": "01/10/2012 23:14:11",
      "content": "<p>[quote=Rajstennaj Barrabas;7555]</p>\r\n<p>I've been reviewing the related rules.</p>\r\n<p>Apparently there are 2 conferences: CVPR and ICPR.</p>\r\n<p>I have no idea what these are about, as their web pages are opaque :0</p>\r\n<p>For each contest there are 2 challenges: Quantitative and Qualitative (4 challenges total).</p>\r\n<p>The quantitative challenge is where your program munches on pre-recorded videos and attempts to decipher gestures.</p>\r\n<p>The qualitative challenge is where your program runs live at a conference, taking video directly from a Kinect and using Microsoft's API. I haven't read the details yet (perhaps they haven't been finalized), but it looks like your program will have to recognize\r\n gestures live in competition with other challengers.</p>\r\n<p>For each of the 4 challenges, the prizes are the same: $5K/$3K/$2K for 1st, 2nd, 3rd place.</p>\r\n<p>It would appear that for the quantitative challenge, all you need is a program that can read videos and generate a csv file.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>CVPR is a computer vision conference and ICPR a pattern recognition conference. Both will be a great opportunity to showcase your technology.</p>\r\n<p>==&gt; The quantitative evaluations are on pre-recorded data. We now released all the data, except the final evaluation data.</p>\r\n<p>==&gt; The qualitative evaluations are demonstration competitions &quot;free style&quot;: you will not be given data, you just have to show a great application of gesture recognition using Kinect.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8052",
      "postDate": "01/10/2012 23:19:00",
      "content": "<p>[quote=Ehouse;8046]</p>\r\n<p>Isabelle,</p>\r\n<p>Are there any rules against modifying the Kinect with additional hardware?<br>\r\nAlso will you give us a new batch of gestures to train on for final evaluation and then validating those? Or will we be validating data that we have previously trained on earlier in the competition?</p>\r\n<p>[/quote]</p>\r\n<p>The quantitative evaluation presently on-going is on pre-recorded data, so your first question does not apply. For the demo competition, I will get back to you later.</p>\r\n<p>For the final evaluation of the&nbsp;quantitative evaluation presently on-going you will get different batches, for new gesture vocabularies, but organized in the same way as the validation data. For each batch you have one labeled example of each gesture to\r\n train. The development data is not really training data. You can use it to design your system and do unsupervised learning or transfer learning to learn data representations. But the supervised learning part will take place only with one example of each gesture\r\n in each batch (one-shot-learning) when you get the final evaluation data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8054",
      "postDate": "01/10/2012 23:23:44",
      "content": "<p>[quote=redstr;7578]</p>\r\n<p>From the description in their SDK (here for example: http://kinectforwindows.org/documents/SkeletalViewer_Walkthrough.pdf ), it seems that the Kinect skeleton tracker can only detect joints if the entire body is present within the frame, which is not true\r\n for the samples in this contest. So I don't think it's directly usable here. Perhaps this is why we have this competition?</p>\r\n<p>[/quote]</p>\r\n<p>1) The SDK skeleton tracker works best for full body, this is true.</p>\r\n<p>2) Having the skeleton may help but this is unclear. Many gestures involve hand and finger motion or posture that are not captured by the skeleton tracker. My guess is that the RGB image is going to play an important role.</p>\r\n<p>3) Having the skeleton is not the end of the problem, you still need to figure out how to match skeleton trajectories and perform one-shot-learning.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8055",
      "postDate": "01/11/2012 00:17:05",
      "content": "<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8061",
      "postDate": "01/11/2012 18:08:43",
      "content": "<p>[quote=Isabelle;8054]2) Having the skeleton may help but this is unclear. Many gestures involve hand and finger motion or posture that are not captured by the skeleton tracker. My guess is that the RGB image is going to play an important role.[/quote]</p>\r\n<p>I for one (and maybe others) want the skeleton data not just for measuring trajectories, but also to reliably identify points of interest in the RGB image. If we know where the hands are in the image, it'll be much easier to learn gestures that are dependent\r\n on hand and finger posture.</p>\r\n<p>I'm working on my own homebrew hand tracker, but who knows how reliable it'll be. It seems like most other attempts at hand tracking involve some kind of initialization gesture, which isn't really possible with the current data set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8063",
      "postDate": "01/11/2012 18:17:19",
      "content": "<p>[quote=Isabelle;8052]</p>\r\n<p>The quantitative evaluation presently on-going is on pre-recorded data, so your first question does not apply. For the demo competition, I will get back to you later.</p>\r\n<p>[/quote]</p>\r\n<p>Perhaps I phrased my question wrong. Could we add additional components to work alongside the Kinect sensor as a hardware excelerator?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8077",
      "postDate": "01/12/2012 21:14:54",
      "content": "<p>[quote=Jose H. Solorzano;8055]</p>\r\n<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>\r\n<p>[/quote]</p>\r\n<p>Hi Jose - the IR camera is slightly offset from the RGB camera in the Kinect</p>\r\n<p>&nbsp;</p>\r\n<p><img title=\"Kinect Sensor\" src=\"http://qph.cf.quoracdn.net/main-qimg-90d9a2ceb96f836e0b724027c2aba723\" alt=\"Kinect Sensor\" width=\"676\" height=\"666\"></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8079",
      "postDate": "01/12/2012 21:49:21",
      "content": "<p>[quote=Ben Hamner;8077]</p>\r\n<p>[quote=Jose H. Solorzano;8055]</p>\r\n<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>\r\n<p>[/quote]</p>\r\n<p>Hi Jose - the IR camera is slightly offset from the RGB camera in the Kinect</p>\r\n<p>[/quote]</p>\r\n<p>Thanks, Ben, that's what I thought. There's probably no official/standard conversion, is there? It has to depend on the depth of the objects and other factors.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8171",
      "postDate": "01/19/2012 19:47:56",
      "content": "<p>Still haven't heard back from you if we are allowed to add additional components to work alongside the Kinect sensor as a hardware excelerator?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8178",
      "postDate": "01/20/2012 04:44:18",
      "content": "<p>[quote=Jose H. Solorzano;8079]</p>\r\n<p>[quote=Ben Hamner;8077]</p>\r\n<p>[quote=Jose H. Solorzano;8055]</p>\r\n<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>\r\n<p>[/quote]</p>\r\n<p>Hi Jose - the IR camera is slightly offset from the RGB camera in the Kinect</p>\r\n<p>[/quote]</p>\r\n<p>Thanks, Ben, that's what I thought. There's probably no official/standard conversion, is there? It has to depend on the depth of the objects and other factors.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>There is a conversion built into the kinect sdk, but since we're using pre-recorded video, and not a live kinect, the kinect sdk is useless.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8188",
      "postDate": "01/20/2012 20:10:25",
      "content": "<p>[quote=Ehouse;8171]</p>\r\n<p>Still haven't heard back from you if we are allowed to add additional components to work alongside the Kinect sensor as a hardware excelerator?</p>\r\n<p>[/quote]</p>\r\n<p>You can use anything you want (hardware or software) as long as you have the legal right to use it. If you are one of the winners and interested in licensing your methods to Microsoft, novelty/originality will play an important role in the decision and you\r\n will need to use only components that Micosoft can license. But to win the small prizes (1st place $5000, 2nd place $3000, 3rd place $2000), there are no restrictions.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8581",
      "postDate": "02/15/2012 20:30:48",
      "content": "<p>I have a question regarding the rules for training. <br>\r\nThe system I am using to train, loads up the gesture video i want to train on. I train it on that gesture, then play the validation data to see if it recognizes it. If the system is having a hard time recognizing it, I will load up the original gesture video\r\n I trained on, and train it with further analysis on that original video. Is this against the rules of &quot;One shot learning&quot;.<br>\r\nI am still only using that one example video to train on the gesture, I just needed to fine tune the analysis I did in the training aspect of said gesture.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "8585",
      "postDate": "02/15/2012 22:31:59",
      "content": "<p>[quote=Ehouse;8581]</p>\r\n<p>I have a question regarding the rules for training. <br>\r\nThe system I am using to train, loads up the gesture video i want to train on. I train it on that gesture, then play the validation data to see if it recognizes it. If the system is having a hard time recognizing it, I will load up the original gesture video\r\n I trained on, and train it with further analysis on that original video. Is this against the rules of &quot;One shot learning&quot;.<br>\r\nI am still only using that one example video to train on the gesture, I just needed to fine tune the analysis I did in the training aspect of said gesture.</p>\r\n<p>[/quote]</p>\r\n<p>As long as you do not add any human-made information (like additional labels that were not provided), you can revisit the data in your procedure. By &quot;one-shot-learning&quot; we just mean that you have only one labeled example of each gesture in a given batch.\r\n It is also permitted to use the unlabeled examples as part of training (e.g. you can cluster the gestures if that helps you recognize them).</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 7094,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "12/13/2011 05:37:51",
      "content": "<p>For the prize, how are you going you deal with equal scores? I mean, what if the top 10 all get the same score (like everybody get 1 wrong, or even 0 wrong prediction). I feel this problem can be done with very very high accuracy, as high as a human can\r\n do.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7097,
      "author_name": "julian1",
      "author_url": "",
      "post_date": "12/13/2011 07:18:44",
      "content": "<p>Is it allowed to use the skeleton tracker from Microsoft to preprocess the data or would we have to write our own tracker?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7125,
      "author_name": "ahassaine",
      "author_url": "",
      "post_date": "12/14/2011 10:03:34",
      "content": "<p>[quote=woshialex;7094]</p>\r\n<p>I feel this problem can be done with very very high accuracy, as high as a human can do.</p>\r\n<p>[/quote]</p>\r\n<p>I am not sure that state-of-the-art techniques in gesture recognition can perform as accurately as a human can do...</p>\r\n<p>Anyway, do we have to perform better than the human performance benchmark to be elligible for the prize?</p>\r\n<p>It could be a good idea to submit a state-of-the-art benchmark rather than just a result of a template matching...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7149,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "12/14/2011 21:29:46",
      "content": "<p>There is no requirement to do better than human performance to win the prizes. The human performance is provided for information only. It is based on the results of a single person. Most errors are due to inattention and do not correspond to intrinsically\r\n ambiguous gestures. Hence, it is possible to do better than human performance.</p>\r\n<p>The organizers will post new benchmark results from time to time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7151,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "12/14/2011 21:34:48",
      "content": "<p>You can use any other third party software, as long as you have the right to use such software. There is no contraint on the originality of the code to enter the challenge and win the prizes. However, if you are interested in licensing your algorithms if\r\n you win, originality will play a role. Microsoft will review the submissions of interested participants and grant up to $100,000 in contracts. See the rules.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7152,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "12/14/2011 21:35:56",
      "content": "<p>For the ties, it is stipulated in the rules:</p>\r\n<p>In the event of a tie between any eligible entries, the tie will be broken in the quantitative evaluation by giving preference to the earliest submission.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7162,
      "author_name": "solorzano",
      "author_url": "",
      "post_date": "12/15/2011 02:06:40",
      "content": "<p>Isabelle,</p>\r\n<p>Could you elaborate on this rule?</p>\r\n<p style=\"padding-left:30px\">In addition, to enter the qualitative evaluation, the participants will be asked to certify that they use a Kinect sensor and the Microsoft Software Development Kit (SDK).\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7508,
      "author_name": "ahassaine",
      "author_url": "",
      "post_date": "12/26/2011 06:41:36",
      "content": "<p>Bonjour :-)</p>\r\n<p>In addition to Jose's question, I have another question before I start coding. I suppose I should avoid using GPL-licensed libraries since they are not Microsoft-friendly?</p>\r\n<p>Thanks,</p>\r\n<p>Ali</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7520,
      "author_name": "shuky27958",
      "author_url": "",
      "post_date": "12/27/2011 16:32:46",
      "content": "<p>Dear Isabelle,</p>\r\n<p>Thank you for your tips.</p>\r\n<p>Could you elaborate or direct to information on how to read the video into Microsoft SDK?</p>\r\n<p>&nbsp;</p>\r\n<p>I would like to use its skeleton output.</p>\r\n<p>&nbsp;</p>\r\n<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7524,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "12/28/2011 00:09:23",
      "content": "<p>[quote=shuky;7520]</p>\r\n<p>Dear Isabelle,</p>\r\n<p>Thank you for your tips.</p>\r\n<p>Could you elaborate or direct to information on how to read the video into Microsoft SDK?</p>\r\n<p>&nbsp;</p>\r\n<p>I would like to use its skeleton output.</p>\r\n<p>&nbsp;</p>\r\n<p>thanks</p>\r\n<p>[/quote]</p>\r\n<p>As far as I know, this is not trivial to do. We are working on providing you with skeleton data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7525,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "12/28/2011 00:27:33",
      "content": "<p>[quote=Ali Hassaïne;7508]</p>\r\n<p>Bonjour :-)</p>\r\n<p>In addition to Jose's question, I have another question before I start coding. I suppose I should avoid using GPL-licensed libraries since they are not Microsoft-friendly?</p>\r\n<p>Thanks,</p>\r\n<p>Ali</p>\r\n<p>[/quote]</p>\r\n<p>My recommendation is to go by the official rules, which imply that:</p>\r\n<p>- You may use any third party software that you have a legal right to use for the present challenge (first quantitative evaluation), because it uses pre-recorded data.</p>\r\n<p>- For the demonstration competition (qualitative evaluation) you will have to use Microsoft hardware (Kinect camera) and the drivers provided by Microsoft with their development kit.</p>\r\n<p>- If you are interested in potentially licensing your code to Microsoft (and benefit from the opportunity offered by Microsoft who may license code for up to $100,000 in this challenge), you need to have some original code of your own. Obviously you may\r\n also use third party libraries that you have a legal right to use and that Microsoft will be allowed to license. But your own original code is what will count. It is possible (and even likely) that Microsoft will prefer code running on a Microsoft platform.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7526,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "12/28/2011 00:31:40",
      "content": "<p>[quote=Jose H. Solorzano;7162]</p>\r\n<p>Isabelle,</p>\r\n<p>Could you elaborate on this rule?</p>\r\n<p style=\"padding-left:30px\">In addition, to enter the qualitative evaluation, the participants will be asked to certify that they use a Kinect sensor and the Microsoft Software Development Kit (SDK).</p>\r\n<p>[/quote]</p>\r\n<p>Microsoft wants to make sure that all the demonstrators in the demonstration competition (qualitative evaluation) use Microsoft hardware and software for motion capture. See my more detailed answer to Ali. This does not concern the present challenge, which\r\n uses pre-recorded data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7555,
      "author_name": "rajstennajbarrabas",
      "author_url": "",
      "post_date": "12/29/2011 03:09:34",
      "content": "I've been reviewing the related rules.\r\n<p>Apparently there are 2 conferences: CVPR and ICPR. </p>\r\n<p>I have no idea what these are about, as their web pages are opaque :0 </p>\r\n<p>For each contest there are 2 challenges: Quantitative and Qualitative (4 challenges total).\r\n</p>\r\n<p>The quantitative challenge is where your program munches on pre-recorded videos and attempts to decipher gestures.\r\n</p>\r\n<p>The qualitative challenge is where your program runs live at a conference, taking video directly from a Kinect and using Microsoft's API. I haven't read the details yet (perhaps they haven't been finalized), but it looks like your program will have to recognize\r\n gestures live in competition with other challengers.</p>\r\n<p>For each of the 4 challenges, the prizes are the same: $5K/$3K/$2K for 1st, 2nd, 3rd place.\r\n</p>\r\n<p>It would appear that for the quantitative challenge, all you need is a program that can read videos and generate a csv file.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7572,
      "author_name": "vthiru",
      "author_url": "",
      "post_date": "12/30/2011 03:42:39",
      "content": "<p>[quote=shuky;7520]</p>\r\n<p>Dear Isabelle,</p>\r\n<p>Thank you for your tips.</p>\r\n<p>Could you elaborate or direct to information on how to read the video into Microsoft SDK?</p>\r\n<p>I would like to use its skeleton output.</p>\r\n<p>thanks</p>\r\n<p>[/quote]</p>\r\n<p>I think most people are waiting for the skeleton output, its probably the best foundation to attack this problem with.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7578,
      "author_name": "redstr",
      "author_url": "",
      "post_date": "12/30/2011 17:48:08",
      "content": "<p>From the description in their SDK (here for example: http://kinectforwindows.org/documents/SkeletalViewer_Walkthrough.pdf ), it seems that the Kinect skeleton tracker can only detect joints if the entire body is present within the frame, which is not true\r\n for the samples in this contest. So I don't think it's directly usable here. Perhaps this is why we have this competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 7607,
      "author_name": "shuky27958",
      "author_url": "",
      "post_date": "01/01/2012 09:44:36",
      "content": "<p><span class=\"x_profilelink\">Isabelle,</span></p>\r\n<p>&nbsp;</p>\r\n<p>Could you elaborate more on the Skeleton data - </p>\r\n<p>When will it be available?</p>\r\n<p>Is there skeleton data for the videos with persons which are sitting (from what I understand skeleton view can only be created when the entire body can be seen)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8046,
      "author_name": "ehouse1",
      "author_url": "",
      "post_date": "01/10/2012 19:04:55",
      "content": "<p>Isabelle,</p>\r\n<p>Are there any rules against modifying the Kinect with additional hardware?<br>\r\nAlso will you give us a new batch of gestures to train on for final evaluation and then validating those? Or will we be validating data that we have previously trained on earlier in the competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8051,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "01/10/2012 23:14:11",
      "content": "<p>[quote=Rajstennaj Barrabas;7555]</p>\r\n<p>I've been reviewing the related rules.</p>\r\n<p>Apparently there are 2 conferences: CVPR and ICPR.</p>\r\n<p>I have no idea what these are about, as their web pages are opaque :0</p>\r\n<p>For each contest there are 2 challenges: Quantitative and Qualitative (4 challenges total).</p>\r\n<p>The quantitative challenge is where your program munches on pre-recorded videos and attempts to decipher gestures.</p>\r\n<p>The qualitative challenge is where your program runs live at a conference, taking video directly from a Kinect and using Microsoft's API. I haven't read the details yet (perhaps they haven't been finalized), but it looks like your program will have to recognize\r\n gestures live in competition with other challengers.</p>\r\n<p>For each of the 4 challenges, the prizes are the same: $5K/$3K/$2K for 1st, 2nd, 3rd place.</p>\r\n<p>It would appear that for the quantitative challenge, all you need is a program that can read videos and generate a csv file.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>CVPR is a computer vision conference and ICPR a pattern recognition conference. Both will be a great opportunity to showcase your technology.</p>\r\n<p>==&gt; The quantitative evaluations are on pre-recorded data. We now released all the data, except the final evaluation data.</p>\r\n<p>==&gt; The qualitative evaluations are demonstration competitions &quot;free style&quot;: you will not be given data, you just have to show a great application of gesture recognition using Kinect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8052,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "01/10/2012 23:19:00",
      "content": "<p>[quote=Ehouse;8046]</p>\r\n<p>Isabelle,</p>\r\n<p>Are there any rules against modifying the Kinect with additional hardware?<br>\r\nAlso will you give us a new batch of gestures to train on for final evaluation and then validating those? Or will we be validating data that we have previously trained on earlier in the competition?</p>\r\n<p>[/quote]</p>\r\n<p>The quantitative evaluation presently on-going is on pre-recorded data, so your first question does not apply. For the demo competition, I will get back to you later.</p>\r\n<p>For the final evaluation of the&nbsp;quantitative evaluation presently on-going you will get different batches, for new gesture vocabularies, but organized in the same way as the validation data. For each batch you have one labeled example of each gesture to\r\n train. The development data is not really training data. You can use it to design your system and do unsupervised learning or transfer learning to learn data representations. But the supervised learning part will take place only with one example of each gesture\r\n in each batch (one-shot-learning) when you get the final evaluation data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8054,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "01/10/2012 23:23:44",
      "content": "<p>[quote=redstr;7578]</p>\r\n<p>From the description in their SDK (here for example: http://kinectforwindows.org/documents/SkeletalViewer_Walkthrough.pdf ), it seems that the Kinect skeleton tracker can only detect joints if the entire body is present within the frame, which is not true\r\n for the samples in this contest. So I don't think it's directly usable here. Perhaps this is why we have this competition?</p>\r\n<p>[/quote]</p>\r\n<p>1) The SDK skeleton tracker works best for full body, this is true.</p>\r\n<p>2) Having the skeleton may help but this is unclear. Many gestures involve hand and finger motion or posture that are not captured by the skeleton tracker. My guess is that the RGB image is going to play an important role.</p>\r\n<p>3) Having the skeleton is not the end of the problem, you still need to figure out how to match skeleton trajectories and perform one-shot-learning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8055,
      "author_name": "solorzano",
      "author_url": "",
      "post_date": "01/11/2012 00:17:05",
      "content": "<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8061,
      "author_name": "mmolek",
      "author_url": "",
      "post_date": "01/11/2012 18:08:43",
      "content": "<p>[quote=Isabelle;8054]2) Having the skeleton may help but this is unclear. Many gestures involve hand and finger motion or posture that are not captured by the skeleton tracker. My guess is that the RGB image is going to play an important role.[/quote]</p>\r\n<p>I for one (and maybe others) want the skeleton data not just for measuring trajectories, but also to reliably identify points of interest in the RGB image. If we know where the hands are in the image, it'll be much easier to learn gestures that are dependent\r\n on hand and finger posture.</p>\r\n<p>I'm working on my own homebrew hand tracker, but who knows how reliable it'll be. It seems like most other attempts at hand tracking involve some kind of initialization gesture, which isn't really possible with the current data set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8063,
      "author_name": "ehouse1",
      "author_url": "",
      "post_date": "01/11/2012 18:17:19",
      "content": "<p>[quote=Isabelle;8052]</p>\r\n<p>The quantitative evaluation presently on-going is on pre-recorded data, so your first question does not apply. For the demo competition, I will get back to you later.</p>\r\n<p>[/quote]</p>\r\n<p>Perhaps I phrased my question wrong. Could we add additional components to work alongside the Kinect sensor as a hardware excelerator?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8077,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "01/12/2012 21:14:54",
      "content": "<p>[quote=Jose H. Solorzano;8055]</p>\r\n<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>\r\n<p>[/quote]</p>\r\n<p>Hi Jose - the IR camera is slightly offset from the RGB camera in the Kinect</p>\r\n<p>&nbsp;</p>\r\n<p><img title=\"Kinect Sensor\" src=\"http://qph.cf.quoracdn.net/main-qimg-90d9a2ceb96f836e0b724027c2aba723\" alt=\"Kinect Sensor\" width=\"676\" height=\"666\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8079,
      "author_name": "solorzano",
      "author_url": "",
      "post_date": "01/12/2012 21:49:21",
      "content": "<p>[quote=Ben Hamner;8077]</p>\r\n<p>[quote=Jose H. Solorzano;8055]</p>\r\n<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>\r\n<p>[/quote]</p>\r\n<p>Hi Jose - the IR camera is slightly offset from the RGB camera in the Kinect</p>\r\n<p>[/quote]</p>\r\n<p>Thanks, Ben, that's what I thought. There's probably no official/standard conversion, is there? It has to depend on the depth of the objects and other factors.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8171,
      "author_name": "ehouse1",
      "author_url": "",
      "post_date": "01/19/2012 19:47:56",
      "content": "<p>Still haven't heard back from you if we are allowed to add additional components to work alongside the Kinect sensor as a hardware excelerator?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8178,
      "author_name": "mmolek",
      "author_url": "",
      "post_date": "01/20/2012 04:44:18",
      "content": "<p>[quote=Jose H. Solorzano;8079]</p>\r\n<p>[quote=Ben Hamner;8077]</p>\r\n<p>[quote=Jose H. Solorzano;8055]</p>\r\n<p>It appears that Depth images are displaced slightly to the left relative to corresponding RGB images. Is this simply an artifact of how Kinect works?</p>\r\n<p>[/quote]</p>\r\n<p>Hi Jose - the IR camera is slightly offset from the RGB camera in the Kinect</p>\r\n<p>[/quote]</p>\r\n<p>Thanks, Ben, that's what I thought. There's probably no official/standard conversion, is there? It has to depend on the depth of the objects and other factors.</p>\r\n<p>[/quote]</p>\r\n<p>&nbsp;</p>\r\n<p>There is a conversion built into the kinect sdk, but since we're using pre-recorded video, and not a live kinect, the kinect sdk is useless.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8188,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "01/20/2012 20:10:25",
      "content": "<p>[quote=Ehouse;8171]</p>\r\n<p>Still haven't heard back from you if we are allowed to add additional components to work alongside the Kinect sensor as a hardware excelerator?</p>\r\n<p>[/quote]</p>\r\n<p>You can use anything you want (hardware or software) as long as you have the legal right to use it. If you are one of the winners and interested in licensing your methods to Microsoft, novelty/originality will play an important role in the decision and you\r\n will need to use only components that Micosoft can license. But to win the small prizes (1st place $5000, 2nd place $3000, 3rd place $2000), there are no restrictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8581,
      "author_name": "ehouse1",
      "author_url": "",
      "post_date": "02/15/2012 20:30:48",
      "content": "<p>I have a question regarding the rules for training. <br>\r\nThe system I am using to train, loads up the gesture video i want to train on. I train it on that gesture, then play the validation data to see if it recognizes it. If the system is having a hard time recognizing it, I will load up the original gesture video\r\n I trained on, and train it with further analysis on that original video. Is this against the rules of &quot;One shot learning&quot;.<br>\r\nI am still only using that one example video to train on the gesture, I just needed to fine tune the analysis I did in the training aspect of said gesture.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 8585,
      "author_name": "iguyon",
      "author_url": "",
      "post_date": "02/15/2012 22:31:59",
      "content": "<p>[quote=Ehouse;8581]</p>\r\n<p>I have a question regarding the rules for training. <br>\r\nThe system I am using to train, loads up the gesture video i want to train on. I train it on that gesture, then play the validation data to see if it recognizes it. If the system is having a hard time recognizing it, I will load up the original gesture video\r\n I trained on, and train it with further analysis on that original video. Is this against the rules of &quot;One shot learning&quot;.<br>\r\nI am still only using that one example video to train on the gesture, I just needed to fine tune the analysis I did in the training aspect of said gesture.</p>\r\n<p>[/quote]</p>\r\n<p>As long as you do not add any human-made information (like additional labels that were not provided), you can revisit the data in your procedure. By &quot;one-shot-learning&quot; we just mean that you have only one labeled example of each gesture in a given batch.\r\n It is also permitted to use the unlabeled examples as part of training (e.g. you can cluster the gestures if that helps you recognize them).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "7092": "",
    "7094": "",
    "7097": "",
    "7125": "",
    "7149": "",
    "7151": "",
    "7152": "",
    "7162": "",
    "7508": "",
    "7520": "",
    "7524": "",
    "7525": "",
    "7526": "",
    "7555": "",
    "7572": "",
    "7578": "",
    "7607": "",
    "8046": "",
    "8051": "",
    "8052": "",
    "8054": "",
    "8055": "",
    "8061": "",
    "8063": "",
    "8077": "",
    "8079": "",
    "8171": "",
    "8178": "",
    "8188": "",
    "8581": "",
    "8585": ""
  },
  "source": "meta"
}