{
  "id": 1591,
  "title": "Observations of the data",
  "url": "/competitions/GestureChallenge/discussion/1591",
  "author_name": "",
  "post_date": "2012-03-19T19:40:54.733Z",
  "votes": 2,
  "comment_count": 1,
  "views": 1439,
  "content": "<p>I have to withdraw from the competion, so I'm posting some observations and conclusions so that others can make use of them.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 1:</strong> The distance measurement is per channel, but offset.</p>\r\n<p>&nbsp;</p>\r\n<p>I took separate R, G, B histograms in the distance frames and compared the pixel counts. The data seem to indicate that the R, G, and B channels all measure the same distribution, but offset from each other. For example, the count of pixels which have a\r\n red = 10 is the same as the count of green = 12 and the same as blue = 8.</p>\r\n<p>&nbsp;</p>\r\n<p>I concluded from this that the three channels are measuring the same distribution but offset from each other, and that calculating the distance at a pixel by averaging the channels (R&#43;B&#43;G)/3 would effectively fuzz the data.</p>\r\n<p>&nbsp;</p>\r\n<p>To calculate the distance of a particular pixel, simply use the green channel value. This appears to have the best spread of values within the distribution.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 2:</strong> The distance values are not continuous</p>\r\n<p>&nbsp;</p>\r\n<p>The histograms also show gaps in the data counts - channel values which have a count of zero at periodic intervals.</p>\r\n<p>&nbsp;</p>\r\n<p>For example, a Red channel histogram is shown below. The values at positions 5, 12, 19, 26 and so on are zero. This indicates, for example, that there are no pixels in the distance image which have a Red component of 26. This is true for all distance frames\r\n in all videos, and since the green histogram is at an offset from the red data, there are corresponding zeroes in the green channel, and also in the blue channel.</p>\r\n<p>&nbsp;</p>\r\n<p>Calculating the *change* in value from one frame to the next without compensating for these zeroes will skew the results.</p>\r\n<p>&nbsp;</p>\r\n<p>For example, consider the MSE of the same pixel taken from two succeeding frames. One frame might show a distance of 11 at that location, and the next might show a 13. The MSE will appear to be (13-11)^2 = 4 when in reality the distances differ by 1 so the\r\n MSE should be (12-11)^2 = 1. The extra distance is an anomaly, caused by the hole in the digitization technique. Since no pixel can have a distance of 13, the new pixel has be the next higher value.</p>\r\n<p>&nbsp;</p>\r\n<p>When distances are remapped to avoid the zeroes, 212 separate distance values remain in a more-or-less continuous distribution.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 3:</strong> Black (R=G=B=0) means &quot;no information&quot;</p>\r\n<p>&nbsp;</p>\r\n<p>In the distance images, the value of &quot;black&quot; does not mean &quot;really close&quot;, but instead means &quot;no information&quot;. The Kinect apparently uses this value to report that the infrared signal has disappeared - hence, &quot;no information&quot;.</p>\r\n<p>&nbsp;</p>\r\n<p>For an example of this, look at devel05/K_26.avi and notice how the patterns change frame to frame. Any detection algorithm needs to account for this special value; for example, by relying on pixel area averages instead of absolute pixel counts.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 4:</strong> The &quot;no information&quot; value (Black) is low-pass filtered</p>\r\n<p>&nbsp;</p>\r\n<p>A histogram of the distance images plotted does not show a sharp vertical line at zero, but a smoothly sloping curve starting at around value 5 and rising up to meet the value at zero.</p>\r\n<p>&nbsp;</p>\r\n<p>This indicates that the black value has been put through a low-pass filter, with the result that some of the black values appear as small-numbered values close to black. Practically speaking, this means that values close to black (around 5 and less) should\r\n also be considered the &quot;no information&quot; value. How close depends on the specific video - I think this has to do with the specifics of the AVI encoding.</p>\r\n<p>&nbsp;</p>\r\n<p>Histogram of red channel in one of the distance frames:</p>\r\n<p>&nbsp;</p>\r\n<table border=\"0\" frame=\"VOID\" rules=\"NONE\" cellspacing=\"0\">\r\n<colgroup><col width=\"86\"></colgroup>\r\n<tbody>\r\n<tr>\r\n<td align=\"RIGHT\" width=\"86\" height=\"18\">1270</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">27</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">17</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">15</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">15</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">22</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">14</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">13</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">15</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">13</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">9</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">5</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">7</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">7</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">5</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">5</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">6</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n</tbody>\r\n</table>",
  "messages": [
    {
      "id": "9544",
      "postDate": "03/19/2012 19:40:54",
      "content": "<p>I have to withdraw from the competion, so I'm posting some observations and conclusions so that others can make use of them.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 1:</strong> The distance measurement is per channel, but offset.</p>\r\n<p>&nbsp;</p>\r\n<p>I took separate R, G, B histograms in the distance frames and compared the pixel counts. The data seem to indicate that the R, G, and B channels all measure the same distribution, but offset from each other. For example, the count of pixels which have a\r\n red = 10 is the same as the count of green = 12 and the same as blue = 8.</p>\r\n<p>&nbsp;</p>\r\n<p>I concluded from this that the three channels are measuring the same distribution but offset from each other, and that calculating the distance at a pixel by averaging the channels (R&#43;B&#43;G)/3 would effectively fuzz the data.</p>\r\n<p>&nbsp;</p>\r\n<p>To calculate the distance of a particular pixel, simply use the green channel value. This appears to have the best spread of values within the distribution.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 2:</strong> The distance values are not continuous</p>\r\n<p>&nbsp;</p>\r\n<p>The histograms also show gaps in the data counts - channel values which have a count of zero at periodic intervals.</p>\r\n<p>&nbsp;</p>\r\n<p>For example, a Red channel histogram is shown below. The values at positions 5, 12, 19, 26 and so on are zero. This indicates, for example, that there are no pixels in the distance image which have a Red component of 26. This is true for all distance frames\r\n in all videos, and since the green histogram is at an offset from the red data, there are corresponding zeroes in the green channel, and also in the blue channel.</p>\r\n<p>&nbsp;</p>\r\n<p>Calculating the *change* in value from one frame to the next without compensating for these zeroes will skew the results.</p>\r\n<p>&nbsp;</p>\r\n<p>For example, consider the MSE of the same pixel taken from two succeeding frames. One frame might show a distance of 11 at that location, and the next might show a 13. The MSE will appear to be (13-11)^2 = 4 when in reality the distances differ by 1 so the\r\n MSE should be (12-11)^2 = 1. The extra distance is an anomaly, caused by the hole in the digitization technique. Since no pixel can have a distance of 13, the new pixel has be the next higher value.</p>\r\n<p>&nbsp;</p>\r\n<p>When distances are remapped to avoid the zeroes, 212 separate distance values remain in a more-or-less continuous distribution.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 3:</strong> Black (R=G=B=0) means &quot;no information&quot;</p>\r\n<p>&nbsp;</p>\r\n<p>In the distance images, the value of &quot;black&quot; does not mean &quot;really close&quot;, but instead means &quot;no information&quot;. The Kinect apparently uses this value to report that the infrared signal has disappeared - hence, &quot;no information&quot;.</p>\r\n<p>&nbsp;</p>\r\n<p>For an example of this, look at devel05/K_26.avi and notice how the patterns change frame to frame. Any detection algorithm needs to account for this special value; for example, by relying on pixel area averages instead of absolute pixel counts.</p>\r\n<p>&nbsp;</p>\r\n<p><strong>Observation 4:</strong> The &quot;no information&quot; value (Black) is low-pass filtered</p>\r\n<p>&nbsp;</p>\r\n<p>A histogram of the distance images plotted does not show a sharp vertical line at zero, but a smoothly sloping curve starting at around value 5 and rising up to meet the value at zero.</p>\r\n<p>&nbsp;</p>\r\n<p>This indicates that the black value has been put through a low-pass filter, with the result that some of the black values appear as small-numbered values close to black. Practically speaking, this means that values close to black (around 5 and less) should\r\n also be considered the &quot;no information&quot; value. How close depends on the specific video - I think this has to do with the specifics of the AVI encoding.</p>\r\n<p>&nbsp;</p>\r\n<p>Histogram of red channel in one of the distance frames:</p>\r\n<p>&nbsp;</p>\r\n<table border=\"0\" frame=\"VOID\" rules=\"NONE\" cellspacing=\"0\">\r\n<colgroup><col width=\"86\"></colgroup>\r\n<tbody>\r\n<tr>\r\n<td align=\"RIGHT\" width=\"86\" height=\"18\">1270</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">27</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">17</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">15</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">15</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">22</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">14</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">13</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">15</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">13</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">9</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">5</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">7</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">2</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">7</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">5</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">0</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">5</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">4</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">6</td>\r\n</tr>\r\n<tr>\r\n<td align=\"RIGHT\" height=\"18\">3</td>\r\n</tr>\r\n</tbody>\r\n</table>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "9606",
      "postDate": "03/20/2012 17:18:32",
      "content": "<p>Observation: FFMPEG can turn the videos into images</p>\r\n<p>FFMPEG is free software and can be used to extract individual frames from the videos. The frames are usually easier to deal with than the raw video, and accessing the frames is generally much faster than rendering the video inline.</p>\r\n<p>The FFMPEG line to use is: &quot;ffmpeg -i $VideoPath $ImageDir/Something.%04d.bmp&quot;</p>\r\n<p>(Replace $VideoPath with the path to the video of interest, and $ImageDir with the destination directory on your system.)</p>\r\n<p>The -i argument identifies the input video file. The last part specifies the output image format (in this example, a BMP file) and the &quot;printf&quot; style feature at the end is recognized by FFMPEG as a file expression.</p>\r\n<p>This will generate image files &quot;Something.0001.bmp Something.0002.bmp Something.0003.bmp&quot; and so on in the specified $ImageDir. For example, the line: &quot;ffmpeg -i K<em>1.avi K</em>1/K<em>1.%04d.bmp&quot; will grab K</em>1.avi in the current directory, render it\r\n as frames, and store the frames as bitmap files K<em>1.0001.bmp and so on. The frames will be stored in a subdirectory names &quot;K</em>1&quot; in the current directory.</p>\r\n<p>Observation: Compressed data results in duplicated frames</p>\r\n<p>Using FFMPEG to render the frames, I noticed that frames are duplicated. For example, rendering devel01/K_1.avi results in 0001.bmp through 0004.bmp as the same image - identical in content. The first non-identical image is 0005.bmp. In similar manner, 0006.bmp\r\n is identical to 0005, and so on until 0009.bmp.</p>\r\n<p>I believe that this is an artifact of the compression used by the .AVI conversion. When compressing a video the system chooses a bit rate, for example 300K bits per second. If the video has a lot of change, then the compression will have to drop frames in\r\n order to achieve the specified compression.</p>\r\n<p>Since the amount of change is specific to the video, this means that the frame number of unique images will be different depending on which video is accessed. Devel01/K<em>1 renders 1 unique image in 5 frames, while devel01/K</em>10 renders one in two. There's\r\n apparently much less change frame by frame in video K_10, which allows the compression to render more unique frames.</p>\r\n<p>So if you render the video to individual frames, be on the lookout for identical frames. Or to put it more succinctly, the compression algorithm quantizes the actions in a temporal direction.</p>\r\n<p>Observation: The Depth image is offset from the color image, and a different size</p>\r\n<p>I've posted an image which overlays the Kinect and color images below. As can be seen from the image, the Kinect data and color data don't match - the Kinect image is a little to the left and above the color image.</p>\r\n<p>A system which attempts to correlate segments from one video stream to another needs to take this into account.</p>\r\n<p>Hmmm... Can't seem to get the image to load. I'll try to add it as an attachment</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 9606,
      "author_name": "rajstennajbarrabas",
      "author_url": "",
      "post_date": "03/20/2012 17:18:32",
      "content": "<p>Observation: FFMPEG can turn the videos into images</p>\r\n<p>FFMPEG is free software and can be used to extract individual frames from the videos. The frames are usually easier to deal with than the raw video, and accessing the frames is generally much faster than rendering the video inline.</p>\r\n<p>The FFMPEG line to use is: &quot;ffmpeg -i $VideoPath $ImageDir/Something.%04d.bmp&quot;</p>\r\n<p>(Replace $VideoPath with the path to the video of interest, and $ImageDir with the destination directory on your system.)</p>\r\n<p>The -i argument identifies the input video file. The last part specifies the output image format (in this example, a BMP file) and the &quot;printf&quot; style feature at the end is recognized by FFMPEG as a file expression.</p>\r\n<p>This will generate image files &quot;Something.0001.bmp Something.0002.bmp Something.0003.bmp&quot; and so on in the specified $ImageDir. For example, the line: &quot;ffmpeg -i K<em>1.avi K</em>1/K<em>1.%04d.bmp&quot; will grab K</em>1.avi in the current directory, render it\r\n as frames, and store the frames as bitmap files K<em>1.0001.bmp and so on. The frames will be stored in a subdirectory names &quot;K</em>1&quot; in the current directory.</p>\r\n<p>Observation: Compressed data results in duplicated frames</p>\r\n<p>Using FFMPEG to render the frames, I noticed that frames are duplicated. For example, rendering devel01/K_1.avi results in 0001.bmp through 0004.bmp as the same image - identical in content. The first non-identical image is 0005.bmp. In similar manner, 0006.bmp\r\n is identical to 0005, and so on until 0009.bmp.</p>\r\n<p>I believe that this is an artifact of the compression used by the .AVI conversion. When compressing a video the system chooses a bit rate, for example 300K bits per second. If the video has a lot of change, then the compression will have to drop frames in\r\n order to achieve the specified compression.</p>\r\n<p>Since the amount of change is specific to the video, this means that the frame number of unique images will be different depending on which video is accessed. Devel01/K<em>1 renders 1 unique image in 5 frames, while devel01/K</em>10 renders one in two. There's\r\n apparently much less change frame by frame in video K_10, which allows the compression to render more unique frames.</p>\r\n<p>So if you render the video to individual frames, be on the lookout for identical frames. Or to put it more succinctly, the compression algorithm quantizes the actions in a temporal direction.</p>\r\n<p>Observation: The Depth image is offset from the color image, and a different size</p>\r\n<p>I've posted an image which overlays the Kinect and color images below. As can be seen from the image, the Kinect data and color data don't match - the Kinect image is a little to the left and above the color image.</p>\r\n<p>A system which attempts to correlate segments from one video stream to another needs to take this into account.</p>\r\n<p>Hmmm... Can't seem to get the image to load. I'll try to add it as an attachment</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "9544": "",
    "9606": ""
  },
  "source": "meta"
}