{
  "id": 5040,
  "title": "Some hints",
  "url": "/competitions/multi-modal-gesture-recognition/discussion/5040",
  "author_name": "",
  "post_date": "2013-07-08T18:45:51.573Z",
  "votes": 2,
  "comment_count": 2,
  "views": 1292,
  "content": "<p>Hello all,<br>\r\n<br>\r\nThis competition looks extremely interesting but, due to the lack of time due to obvious summer commitments, I will not take place in it. However, I had a look at the data and got some ideas that I am sharing with you, in case someone would like to implement/develop\r\n them. <br>\r\n<br>\r\n<strong>First approach</strong></p>\r\n<ul>\r\n<li>A pure speech recognition approach. Data is clean enough to try a word spotter or a hidden Markov model-based (HMM) word recognition system. You have plenty of instances to decently train a HMM and there is some open source software that may save you programming\r\n the HMM by yourself (though a nice exercise). (<a target=\"_blank\" href=\"http://sourceforge.net/projects/hmm-asr-matlab/\">http://sourceforge.net/projects/hmm-asr-matlab/</a>)\r\n</li></ul>\r\n<p><strong>Second approach</strong></p>\r\n<ul>\r\n<li>You can mimic the idea I presented in a paper some years ago closely related with the topic of this competition:\r\n</li></ul>\r\n<p></p>\r\n<blockquote>\r\n<ul>\r\n<li>T. Butko, C. Canton-Ferrer, C. Segura, X. Giró, C. Nadeu, J. Hernando, J.R. Casas.\r\n<em>Acoustic Event Detection based on Feature-Level Fusion of Audio and Video Modalities</em>. EURASIP Journal on Advances in Signal Processing, Hindawi Publishing, January 2011. (<a href=\"http://www.cristiancanton.org/data/pubs/papers/2011-JASP-Canton.pdf\">http://www.cristiancanton.org/data/pubs/papers/2011-JASP-Canton.pdf</a>)\r\n</li></ul>\r\n</blockquote>\r\n<p></p>\r\n<blockquote>\r\n<ul>\r\n<li>C. Canton-Ferrer, T. Butko, C. Segura, X. Giró, C. Nadeu, J. Hernando, J.R. Casas.\r\n<em>Multimodal Acoustic Event Detection Towards Scene Understanding</em>. In Proc. of IEEE Workshop on Human Communicative Behavior Analysis within the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Miami (Florida, USA), June\r\n 2009. (<a href=\"http://www.cristiancanton.org/data/pubs/papers/2009-CVPR-Canton.pdf\">http://www.cristiancanton.org/data/pubs/papers/2009-CVPR-Canton.pdf</a>)\r\n</li></ul>\r\n</blockquote>\r\n<p>Maybe you can try some of the multimodal data fusion strategies presented there.<br>\r\n<br>\r\n<strong>Third approach</strong>. This is just the outline of the path I would follow if I would participate<br>\r\n</p>\r\n<ul>\r\n<li>Collect some features from the audio stream every T_a seconds (for audio features, have a look at the two previous papers or have look at Butko's PhD thesis (<a href=\"http://taras-butko.info/files/thesis/Thesis_v33.pdf\">http://taras-butko.info/files/thesis/Thesis_v33.pdf</a>)).\r\n</li><li>Collect some features from the provided skeleton every T_s seconds (e.g. position from the normalized center of the body, velocity and acceleration of every joint, etc.)\r\n</li><li>&nbsp;Create some informative features related to the body parts every T_i seconds (e.g. distance from every hand to the face, distance from every hand to each other, etc.)\r\n</li><li>Brew some features of your own every T_o seconds </li><li>Make sure they are well synchronized and aligned </li><li>For every action, get the average time. Let's image that for action A, we have an average of A_t seconds\r\n</li><li>Time warp all the features related to a given instance of action A to last A_t\r\n</li><li>For every action, train your favorite classification technique using the k features (per time instant) and the A_t samples (effectively generating A_t * k features )\r\n</li><li>Test:\r\n<ul>\r\n<li>Extract features </li><li>Find the most likely action at every time instant by sliding every trained classifier for every action over whole sequence. This will provide a sequence of actions (one action per time instant)\r\n</li><li>Group the sequence into clusters of repeated consecutive actions (or use a a heuristic approach)\r\n</li><li>Submit. </li></ul>\r\n</li></ul>\r\n<p><br>\r\nNot that this will put you high up in the leaderboard since this approach lacks many important bits and pieces, but I hope it gives some ideas to get started! Enjoy!</p>\r\n<p>Cristian</p>\r\n<p></p>\r\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "26810",
      "postDate": "07/08/2013 18:45:51",
      "content": "<p>Hello all,<br>\r\n<br>\r\nThis competition looks extremely interesting but, due to the lack of time due to obvious summer commitments, I will not take place in it. However, I had a look at the data and got some ideas that I am sharing with you, in case someone would like to implement/develop\r\n them. <br>\r\n<br>\r\n<strong>First approach</strong></p>\r\n<ul>\r\n<li>A pure speech recognition approach. Data is clean enough to try a word spotter or a hidden Markov model-based (HMM) word recognition system. You have plenty of instances to decently train a HMM and there is some open source software that may save you programming\r\n the HMM by yourself (though a nice exercise). (<a target=\"_blank\" href=\"http://sourceforge.net/projects/hmm-asr-matlab/\">http://sourceforge.net/projects/hmm-asr-matlab/</a>)\r\n</li></ul>\r\n<p><strong>Second approach</strong></p>\r\n<ul>\r\n<li>You can mimic the idea I presented in a paper some years ago closely related with the topic of this competition:\r\n</li></ul>\r\n<p></p>\r\n<blockquote>\r\n<ul>\r\n<li>T. Butko, C. Canton-Ferrer, C. Segura, X. Giró, C. Nadeu, J. Hernando, J.R. Casas.\r\n<em>Acoustic Event Detection based on Feature-Level Fusion of Audio and Video Modalities</em>. EURASIP Journal on Advances in Signal Processing, Hindawi Publishing, January 2011. (<a href=\"http://www.cristiancanton.org/data/pubs/papers/2011-JASP-Canton.pdf\">http://www.cristiancanton.org/data/pubs/papers/2011-JASP-Canton.pdf</a>)\r\n</li></ul>\r\n</blockquote>\r\n<p></p>\r\n<blockquote>\r\n<ul>\r\n<li>C. Canton-Ferrer, T. Butko, C. Segura, X. Giró, C. Nadeu, J. Hernando, J.R. Casas.\r\n<em>Multimodal Acoustic Event Detection Towards Scene Understanding</em>. In Proc. of IEEE Workshop on Human Communicative Behavior Analysis within the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Miami (Florida, USA), June\r\n 2009. (<a href=\"http://www.cristiancanton.org/data/pubs/papers/2009-CVPR-Canton.pdf\">http://www.cristiancanton.org/data/pubs/papers/2009-CVPR-Canton.pdf</a>)\r\n</li></ul>\r\n</blockquote>\r\n<p>Maybe you can try some of the multimodal data fusion strategies presented there.<br>\r\n<br>\r\n<strong>Third approach</strong>. This is just the outline of the path I would follow if I would participate<br>\r\n</p>\r\n<ul>\r\n<li>Collect some features from the audio stream every T_a seconds (for audio features, have a look at the two previous papers or have look at Butko's PhD thesis (<a href=\"http://taras-butko.info/files/thesis/Thesis_v33.pdf\">http://taras-butko.info/files/thesis/Thesis_v33.pdf</a>)).\r\n</li><li>Collect some features from the provided skeleton every T_s seconds (e.g. position from the normalized center of the body, velocity and acceleration of every joint, etc.)\r\n</li><li>&nbsp;Create some informative features related to the body parts every T_i seconds (e.g. distance from every hand to the face, distance from every hand to each other, etc.)\r\n</li><li>Brew some features of your own every T_o seconds </li><li>Make sure they are well synchronized and aligned </li><li>For every action, get the average time. Let's image that for action A, we have an average of A_t seconds\r\n</li><li>Time warp all the features related to a given instance of action A to last A_t\r\n</li><li>For every action, train your favorite classification technique using the k features (per time instant) and the A_t samples (effectively generating A_t * k features )\r\n</li><li>Test:\r\n<ul>\r\n<li>Extract features </li><li>Find the most likely action at every time instant by sliding every trained classifier for every action over whole sequence. This will provide a sequence of actions (one action per time instant)\r\n</li><li>Group the sequence into clusters of repeated consecutive actions (or use a a heuristic approach)\r\n</li><li>Submit. </li></ul>\r\n</li></ul>\r\n<p><br>\r\nNot that this will put you high up in the leaderboard since this approach lacks many important bits and pieces, but I hope it gives some ideas to get started! Enjoy!</p>\r\n<p>Cristian</p>\r\n<p></p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "26847",
      "postDate": "07/09/2013 10:35:24",
      "content": "<p>Dear Cristian,</p>\r\n<p><span id=\"x_x_result_box\" lang=\"en\"><span class=\"x_x_hps\">it is a pity</span><span class=\"x_x_hps\"> you can not</span>\r\n<span class=\"x_x_hps\">participate.</span> <span class=\"x_x_hps\">Thanks for the</span>\r\n<span class=\"x_x_hps\">information provided</span> <span class=\"x_x_hps\">to other participants</span><span>,</span> it\r\n<span class=\"x_x_hps\">is very useful</span><span>. Nevertheless, we suggest to combine audio information with other data cues (like your 3th approach). Due to the high penalization of extra added gestures, it is important to use all information to avoid false\r\n positives.</span></span></p>\r\n<p><span id=\"x_x_result_box\" lang=\"en\"><span>Xavier<br>\r\n</span></span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "26882",
      "postDate": "07/09/2013 15:49:59",
      "content": "<p>Absolutely! Actually, you can follow two paths when fusing information from several modalities: fusion at feature level and fusion at data level. The first, generating features from each modality and then fusing them using a combination of classifiers and,\r\n the second, generating multidimensional features and then classifying them. A good read is\r\n<a href=\"http://www.wiley.com/WileyCDA/WileyTitle/productCd-0471210781.html\">Kuncheva's book</a>.</p>\r\n<p>Enjoy!</p>\r\n<p>Cristian</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 26847,
      "author_name": "xbaro100272",
      "author_url": "",
      "post_date": "07/09/2013 10:35:24",
      "content": "<p>Dear Cristian,</p>\r\n<p><span id=\"x_x_result_box\" lang=\"en\"><span class=\"x_x_hps\">it is a pity</span><span class=\"x_x_hps\"> you can not</span>\r\n<span class=\"x_x_hps\">participate.</span> <span class=\"x_x_hps\">Thanks for the</span>\r\n<span class=\"x_x_hps\">information provided</span> <span class=\"x_x_hps\">to other participants</span><span>,</span> it\r\n<span class=\"x_x_hps\">is very useful</span><span>. Nevertheless, we suggest to combine audio information with other data cues (like your 3th approach). Due to the high penalization of extra added gestures, it is important to use all information to avoid false\r\n positives.</span></span></p>\r\n<p><span id=\"x_x_result_box\" lang=\"en\"><span>Xavier<br>\r\n</span></span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 26882,
      "author_name": "cristiancanton",
      "author_url": "",
      "post_date": "07/09/2013 15:49:59",
      "content": "<p>Absolutely! Actually, you can follow two paths when fusing information from several modalities: fusion at feature level and fusion at data level. The first, generating features from each modality and then fusing them using a combination of classifiers and,\r\n the second, generating multidimensional features and then classifying them. A good read is\r\n<a href=\"http://www.wiley.com/WileyCDA/WileyTitle/productCd-0471210781.html\">Kuncheva's book</a>.</p>\r\n<p>Enjoy!</p>\r\n<p>Cristian</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "26810": "",
    "26847": "",
    "26882": ""
  },
  "source": "meta"
}