{
  "id": 5541,
  "title": "Fact Sheets / Model Sharing",
  "url": "/competitions/multi-modal-gesture-recognition/discussion/5541",
  "author_name": "",
  "post_date": "2013-08-25T23:36:06.370Z",
  "votes": 1,
  "comment_count": 1,
  "views": 1419,
  "content": "<p>Since I don't know where to upload the fact sheet, and since many of us will share our models soon anyway, here's a description of my model. &nbsp;Fact sheet attached too as PDF.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>Finding Gestures:</p>\n<p>I search for ~2-second-long regions of high audio-energy to define periods of time that potentially contain a gesture. For the purpose of training, I create a 21st label to signify &quot;not a recognized gesture&quot;.</p>\n<p>&nbsp;</p>\n<p>Features:</p>\n<p>I use the joint positions and angles above the hips and a log-frequency-spaced spectrogram of the audio data as my features. I down-sample all data onto a 5 Hz grid and use ~2 seconds of data. I subtract the average 3d position of the left and right shoulders from each 3d joint position.</p>\n<p>&nbsp;</p>\n<p>Model:</p>\n<p>I train a random forest and a k-Nearest Neighbor model using these features and the labels. I average the posteriors from these models with equal weight. Finally, I have a simple heuristic (limit the number of gestures, no repeats) to convert these posteriors to a prediction for the sequence of gestures. &nbsp;I use python and scikit-learn throughout.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
  "messages": [
    {
      "id": "29457",
      "postDate": "08/25/2013 23:36:06",
      "content": "<p>Since I don't know where to upload the fact sheet, and since many of us will share our models soon anyway, here's a description of my model. &nbsp;Fact sheet attached too as PDF.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>Finding Gestures:</p>\n<p>I search for ~2-second-long regions of high audio-energy to define periods of time that potentially contain a gesture. For the purpose of training, I create a 21st label to signify &quot;not a recognized gesture&quot;.</p>\n<p>&nbsp;</p>\n<p>Features:</p>\n<p>I use the joint positions and angles above the hips and a log-frequency-spaced spectrogram of the audio data as my features. I down-sample all data onto a 5 Hz grid and use ~2 seconds of data. I subtract the average 3d position of the left and right shoulders from each 3d joint position.</p>\n<p>&nbsp;</p>\n<p>Model:</p>\n<p>I train a random forest and a k-Nearest Neighbor model using these features and the labels. I average the posteriors from these models with equal weight. Finally, I have a simple heuristic (limit the number of gestures, no repeats) to convert these posteriors to a prediction for the sequence of gestures. &nbsp;I use python and scikit-learn throughout.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "29571",
      "postDate": "08/27/2013 14:06:42",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 29571,
      "author_name": "xbaro100272",
      "author_url": "",
      "post_date": "08/27/2013 14:06:42",
      "content": "<p>Thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "29457": "",
    "29571": ""
  },
  "source": "meta"
}