{
  "id": 21209,
  "title": "Tutorial: visualizing difference between three pictures",
  "url": "/competitions/draper-satellite-image-chronology/discussion/21209",
  "author_name": "",
  "post_date": "2016-05-25T12:20:30.390Z",
  "votes": 7,
  "comment_count": 7,
  "views": 2571,
  "content": "<p>Hello,</p>\n\n<p>This is a tutorial to visualize the difference between a set of three different pictures. This technique generalizes to a 4 or more pictures (like 5 for Draper), although the interpretation becomes widely fuzzy (and difficult!). They can, however, be used for deep learning or any ML method you want.</p>\n\n<p>We will use the three last pictures of the set 79 (given in the train set). We are assuming you already aligned the pictures. Our hypothesis is that there are (mostly) disappearing containers throughout the 3 images.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4320/Montage_initial.jpg\" alt=\"Montage initial.jpg\" title></p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4321/Montage_aligned.jpg\" alt=\"Montage aligned.jpg\" title></p>\n\n<hr>\n\n<p>To visualize (meaningfully) the difference between three pictures, we need to follow the following steps:</p>\n\n<ul>\n<li>Convert the image to binary</li>\n<li>Optional: measure the area covered by the binary images (good when analyzing moving camps, deforestation, disappearing objects, etc.)</li>\n<li>Interpret the stack of binarized images to RGB</li>\n</ul>\n\n<p>The idea of visualizing the difference meaningfully is to minimize (to a certain extent) the human bias in perception of three consecutive pictures, given the knowledge of their order (or to help the human to perceive differences between pictures when too much information is provided).</p>\n\n<hr>\n\n<h2>I. Converting the image to binary</h2>\n\n<p>This step is probably the easiest (along with the binary stacking to RGB). You just need to use a binarization technique. Most of them incurs the following:</p>\n\n<ol>\n<li>Conversion to 8-bit image</li>\n<li>Conversion to binary using a preferred method with a black background</li>\n</ol>\n\n<p>Here are our three 8-bit images.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4322/Montage_aligned_8bit.jpg\" alt=\"Montage aligned 8bit.jpg\" title></p>\n\n<p>We convert the 8-bit images to binary per image (and not per stack) using black as background:</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\" alt=\"Montage aligned binary.jpg\" title></p>\n\n<p>Until now... nothing hard right?</p>\n\n<hr>\n\n<h2>(Optional) II. Measuring the area covered by the binary images</h2>\n\n<p>This step is a good way to quantify the missing difference between two pictures. However, it quantifies the following:</p>\n\n<ul>\n<li>Adds all the missing elements between two pictures</li>\n<li>Subtracts all the appearing elements between two pictures</li>\n</ul>\n\n<p>This is an issue you must take into account when you are analyzing a set of images! As you will compute the aggregated value (difference between the Add/Substract), you may end up with wrong interpretations using only that information.</p>\n\n<p>To measure the area covered by the binary images, a first step is to threshold temporarily the images. All the white pixels must turn colored. If not, adjust appropriately the sliders of the software you are using.</p>\n\n<p>After, you must specify the area you need to measure. We select the area on the right to the left, omitting the road (you can notice in the image 2, the road is black while in the image 1/3 it is white.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\" alt=\"Montage aligned binary.jpg\" title></p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4324/Stack_aligned_selection.gif\" alt=\"Stack aligned selection.gif\" title></p>\n\n<p>Now, we can compute the area. We get the following:</p>\n\n<pre><code>    Area\n1 2693107\n2 2704704\n3 2819013\n</code></pre>\n\n<p>Did we expect this ordering? Yes, because we hypothesized containers to disappear. Hence:</p>\n\n<ul>\n<li>If a container has disappeared, its area becomes fully white</li>\n<li>If a light colored container has appeared, its contour becomes black along with the shadows</li>\n<li>If a dark colored container has appeared, its area become fully black along with the shadows</li>\n</ul>\n\n<p>This confirms our hypothesis about the containers disappearing gradually from the image 1 to 2 to 3.</p>\n\n<p>This is all good already!</p>\n\n<hr>\n\n<h2>III. Interpret the stack of binarized images and RGB</h2>\n\n<p>This step is the hardest of perform, although it is trivial for a set of only three images. You must convert the three pictures to RGB channels:</p>\n\n<ul>\n<li>Image 1's whiteness become the Red channel (Red)</li>\n<li>Image 2's whiteness become the Green channel (Green)</li>\n<li>Image 3's whiteness become the Blue channel (Blue)</li>\n</ul>\n\n<p>Remember to remove the threshold filter if you did the previous step.</p>\n\n<p>You get the following picture:</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4325/RGB_stack.jpg\" alt=\"RGB stack.jpg\" title></p>\n\n<p>When composing three colors together, you need to remember how additive colors are:</p>\n\n<ul>\n<li>Red =&gt; Red</li>\n<li>Blue =&gt; Blue</li>\n<li>Green =&gt; Green</li>\n<li>Blue+Red =&gt; Magenta</li>\n<li>Red+Green =&gt; Yellow</li>\n<li>Green+Blue =&gt; Cyan</li>\n<li>Red+Green+Blue =&gt; White</li>\n<li>None =&gt; Black</li>\n</ul>\n\n<p>Visual image taken online:</p>\n\n<p><img src=\"http://hyperphysics.phy-astr.gsu.edu/hbase/vision/imgvis/addspotl.gif\" alt=\"Image Online\" title></p>\n\n<p>Note:</p>\n\n<p>If you assume each picture is a color channel (picture 1/2/3 's whiteness =&gt; R/G/B), you get the following:</p>\n\n<pre><code>Color    1 2 3\nRed      W B B\nGreen    B W B\nBlue     B B W\nCyan     B W W\nMagenta  W B W\nYellow   W W B\nWhite    W W W\nBlack    B B B\n</code></pre>\n\n<p>If we suppose white (W) is presence (of missing) and black (B) is absence (of missing), we have the following:</p>\n\n<pre><code>Color    Presence  Absence\nRed            1      2, 3\nGreen          2      1, 3\nBlue           3      1, 2\nCyan        1, 2         1\nMagenta     1, 3         2\nYellow      2, 3         3\nWhite    1, 2, 3\nBlack              1, 2, 3\n</code></pre>\n\n<p>Therefore, the pictures:</p>\n\n<ul>\n<li>(1) should have the most red (disappearing in (2) and (3) )</li>\n<li>(2) should have the most magenta (disappearing in (3) + the common elements with (1) )</li>\n<li>(3) should have the most black (the common elements of (1) and (2) with (3) )</li>\n</ul>\n\n<p>How is our RGB going? We just have to compute a count of unique values from a histogram (I selected a slightly different areas, so the values are different than part II.).</p>\n\n<pre><code>Selection count:   3789045\nGrey (0) count:     797752 (none)\nGrey (85) count:    267728\nGrey (170) count:   309932\nGrey (255) count:  2413633 (1, 2, 3)\nRed (255) count:   2662863 (1 =&gt; present in 1 only)\nGreen (255) count: 2675696 (2 =&gt; present in 2 only)\nBlue (255) count:  2789932 (3 =&gt; present in 3 only)\n</code></pre>\n\n<p>The grey values are not interesting for us (difficult to infer). The three last lines are the most interesting for us: if our hypothesis (stuff disappearing) is right, we should have Red &gt; Green &gt; Blue. This is the case, which also confirms our hypothesis.</p>\n\n<p>If you want to look at the five stacks... not recommending it!</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4326/Five_stacks.jpg\" alt=\"Five stacks.jpg\" title></p>\n\n<hr>\n\n<h2>IV. Deep learning method for differentiation</h2>\n\n<p>What you can also do is using deep learning to compute the most probable sequence of images. This is a transformation of ranking to classification problem.</p>\n\n<p>You would do the following:</p>\n\n<ul>\n<li>For each set, create all existing combinations of images with the differences computed; {1, 2, 3, 4, 5}, {1, 2, 3, 5, 4}, {1, 2, 4, 3, 5}, {1, 2, 4, 5, 3}... overall you end up with 120 pictures for 1 set.</li>\n<li>Transform the ranking issue to classification, assigning the label 1 for the correct sequence, and 0 for the incorrect (or vice-versa).</li>\n<li>Train a deep learning model on the whole training set using the network of your choice.</li>\n<li>Predict on the test set. Per set of 120 pictures, the highest probability picture is the most predicted right order by your deep learning network model.</li>\n</ul>\n\n<p>There is a possible way also by subsampling each set by 3 pictures: instead of using the set {1, 2, 3, 4, 5}, you would label {1, 2, 3}, {2, 3, 4}, and {3, 4, 5} as right. However, you would label {1, 2, 4}, {1, 2, 5}, {1, 3, 2}... as wrong (total: 60 pictures per set, with 3 right and 57 wrong). You would then train a model, and the prediction would be the first unconflicting subset of the highest predictions per set.</p>\n\n<hr>\n\n<p>Off-topic: Hmm...?</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4327/Kaggle_too_many_requests.JPG\" alt=\"enter image description here\" title></p>",
  "messages": [
    {
      "id": "121295",
      "postDate": "05/25/2016 12:20:30",
      "content": "<p>Hello,</p>\n\n<p>This is a tutorial to visualize the difference between a set of three different pictures. This technique generalizes to a 4 or more pictures (like 5 for Draper), although the interpretation becomes widely fuzzy (and difficult!). They can, however, be used for deep learning or any ML method you want.</p>\n\n<p>We will use the three last pictures of the set 79 (given in the train set). We are assuming you already aligned the pictures. Our hypothesis is that there are (mostly) disappearing containers throughout the 3 images.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4320/Montage_initial.jpg\" alt=\"Montage initial.jpg\" title></p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4321/Montage_aligned.jpg\" alt=\"Montage aligned.jpg\" title></p>\n\n<hr>\n\n<p>To visualize (meaningfully) the difference between three pictures, we need to follow the following steps:</p>\n\n<ul>\n<li>Convert the image to binary</li>\n<li>Optional: measure the area covered by the binary images (good when analyzing moving camps, deforestation, disappearing objects, etc.)</li>\n<li>Interpret the stack of binarized images to RGB</li>\n</ul>\n\n<p>The idea of visualizing the difference meaningfully is to minimize (to a certain extent) the human bias in perception of three consecutive pictures, given the knowledge of their order (or to help the human to perceive differences between pictures when too much information is provided).</p>\n\n<hr>\n\n<h2>I. Converting the image to binary</h2>\n\n<p>This step is probably the easiest (along with the binary stacking to RGB). You just need to use a binarization technique. Most of them incurs the following:</p>\n\n<ol>\n<li>Conversion to 8-bit image</li>\n<li>Conversion to binary using a preferred method with a black background</li>\n</ol>\n\n<p>Here are our three 8-bit images.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4322/Montage_aligned_8bit.jpg\" alt=\"Montage aligned 8bit.jpg\" title></p>\n\n<p>We convert the 8-bit images to binary per image (and not per stack) using black as background:</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\" alt=\"Montage aligned binary.jpg\" title></p>\n\n<p>Until now... nothing hard right?</p>\n\n<hr>\n\n<h2>(Optional) II. Measuring the area covered by the binary images</h2>\n\n<p>This step is a good way to quantify the missing difference between two pictures. However, it quantifies the following:</p>\n\n<ul>\n<li>Adds all the missing elements between two pictures</li>\n<li>Subtracts all the appearing elements between two pictures</li>\n</ul>\n\n<p>This is an issue you must take into account when you are analyzing a set of images! As you will compute the aggregated value (difference between the Add/Substract), you may end up with wrong interpretations using only that information.</p>\n\n<p>To measure the area covered by the binary images, a first step is to threshold temporarily the images. All the white pixels must turn colored. If not, adjust appropriately the sliders of the software you are using.</p>\n\n<p>After, you must specify the area you need to measure. We select the area on the right to the left, omitting the road (you can notice in the image 2, the road is black while in the image 1/3 it is white.</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\" alt=\"Montage aligned binary.jpg\" title></p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4324/Stack_aligned_selection.gif\" alt=\"Stack aligned selection.gif\" title></p>\n\n<p>Now, we can compute the area. We get the following:</p>\n\n<pre><code>    Area\n1 2693107\n2 2704704\n3 2819013\n</code></pre>\n\n<p>Did we expect this ordering? Yes, because we hypothesized containers to disappear. Hence:</p>\n\n<ul>\n<li>If a container has disappeared, its area becomes fully white</li>\n<li>If a light colored container has appeared, its contour becomes black along with the shadows</li>\n<li>If a dark colored container has appeared, its area become fully black along with the shadows</li>\n</ul>\n\n<p>This confirms our hypothesis about the containers disappearing gradually from the image 1 to 2 to 3.</p>\n\n<p>This is all good already!</p>\n\n<hr>\n\n<h2>III. Interpret the stack of binarized images and RGB</h2>\n\n<p>This step is the hardest of perform, although it is trivial for a set of only three images. You must convert the three pictures to RGB channels:</p>\n\n<ul>\n<li>Image 1's whiteness become the Red channel (Red)</li>\n<li>Image 2's whiteness become the Green channel (Green)</li>\n<li>Image 3's whiteness become the Blue channel (Blue)</li>\n</ul>\n\n<p>Remember to remove the threshold filter if you did the previous step.</p>\n\n<p>You get the following picture:</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4325/RGB_stack.jpg\" alt=\"RGB stack.jpg\" title></p>\n\n<p>When composing three colors together, you need to remember how additive colors are:</p>\n\n<ul>\n<li>Red =&gt; Red</li>\n<li>Blue =&gt; Blue</li>\n<li>Green =&gt; Green</li>\n<li>Blue+Red =&gt; Magenta</li>\n<li>Red+Green =&gt; Yellow</li>\n<li>Green+Blue =&gt; Cyan</li>\n<li>Red+Green+Blue =&gt; White</li>\n<li>None =&gt; Black</li>\n</ul>\n\n<p>Visual image taken online:</p>\n\n<p><img src=\"http://hyperphysics.phy-astr.gsu.edu/hbase/vision/imgvis/addspotl.gif\" alt=\"Image Online\" title></p>\n\n<p>Note:</p>\n\n<p>If you assume each picture is a color channel (picture 1/2/3 's whiteness =&gt; R/G/B), you get the following:</p>\n\n<pre><code>Color    1 2 3\nRed      W B B\nGreen    B W B\nBlue     B B W\nCyan     B W W\nMagenta  W B W\nYellow   W W B\nWhite    W W W\nBlack    B B B\n</code></pre>\n\n<p>If we suppose white (W) is presence (of missing) and black (B) is absence (of missing), we have the following:</p>\n\n<pre><code>Color    Presence  Absence\nRed            1      2, 3\nGreen          2      1, 3\nBlue           3      1, 2\nCyan        1, 2         1\nMagenta     1, 3         2\nYellow      2, 3         3\nWhite    1, 2, 3\nBlack              1, 2, 3\n</code></pre>\n\n<p>Therefore, the pictures:</p>\n\n<ul>\n<li>(1) should have the most red (disappearing in (2) and (3) )</li>\n<li>(2) should have the most magenta (disappearing in (3) + the common elements with (1) )</li>\n<li>(3) should have the most black (the common elements of (1) and (2) with (3) )</li>\n</ul>\n\n<p>How is our RGB going? We just have to compute a count of unique values from a histogram (I selected a slightly different areas, so the values are different than part II.).</p>\n\n<pre><code>Selection count:   3789045\nGrey (0) count:     797752 (none)\nGrey (85) count:    267728\nGrey (170) count:   309932\nGrey (255) count:  2413633 (1, 2, 3)\nRed (255) count:   2662863 (1 =&gt; present in 1 only)\nGreen (255) count: 2675696 (2 =&gt; present in 2 only)\nBlue (255) count:  2789932 (3 =&gt; present in 3 only)\n</code></pre>\n\n<p>The grey values are not interesting for us (difficult to infer). The three last lines are the most interesting for us: if our hypothesis (stuff disappearing) is right, we should have Red &gt; Green &gt; Blue. This is the case, which also confirms our hypothesis.</p>\n\n<p>If you want to look at the five stacks... not recommending it!</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4326/Five_stacks.jpg\" alt=\"Five stacks.jpg\" title></p>\n\n<hr>\n\n<h2>IV. Deep learning method for differentiation</h2>\n\n<p>What you can also do is using deep learning to compute the most probable sequence of images. This is a transformation of ranking to classification problem.</p>\n\n<p>You would do the following:</p>\n\n<ul>\n<li>For each set, create all existing combinations of images with the differences computed; {1, 2, 3, 4, 5}, {1, 2, 3, 5, 4}, {1, 2, 4, 3, 5}, {1, 2, 4, 5, 3}... overall you end up with 120 pictures for 1 set.</li>\n<li>Transform the ranking issue to classification, assigning the label 1 for the correct sequence, and 0 for the incorrect (or vice-versa).</li>\n<li>Train a deep learning model on the whole training set using the network of your choice.</li>\n<li>Predict on the test set. Per set of 120 pictures, the highest probability picture is the most predicted right order by your deep learning network model.</li>\n</ul>\n\n<p>There is a possible way also by subsampling each set by 3 pictures: instead of using the set {1, 2, 3, 4, 5}, you would label {1, 2, 3}, {2, 3, 4}, and {3, 4, 5} as right. However, you would label {1, 2, 4}, {1, 2, 5}, {1, 3, 2}... as wrong (total: 60 pictures per set, with 3 right and 57 wrong). You would then train a model, and the prediction would be the first unconflicting subset of the highest predictions per set.</p>\n\n<hr>\n\n<p>Off-topic: Hmm...?</p>\n\n<p><img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4327/Kaggle_too_many_requests.JPG\" alt=\"enter image description here\" title></p>",
      "rawMarkdown": "Hello,\r\n\r\nThis is a tutorial to visualize the difference between a set of three different pictures. This technique generalizes to a 4 or more pictures (like 5 for Draper), although the interpretation becomes widely fuzzy (and difficult!). They can, however, be used for deep learning or any ML method you want.\r\n\r\nWe will use the three last pictures of the set 79 (given in the train set). We are assuming you already aligned the pictures. Our hypothesis is that there are (mostly) disappearing containers throughout the 3 images.\r\n\r\n![Montage initial.jpg][1]\r\n\r\n![Montage aligned.jpg][2]\r\n\r\n----------\r\n\r\n\r\nTo visualize (meaningfully) the difference between three pictures, we need to follow the following steps:\r\n\r\n* Convert the image to binary\r\n* Optional: measure the area covered by the binary images (good when analyzing moving camps, deforestation, disappearing objects, etc.)\r\n* Interpret the stack of binarized images to RGB\r\n\r\nThe idea of visualizing the difference meaningfully is to minimize (to a certain extent) the human bias in perception of three consecutive pictures, given the knowledge of their order (or to help the human to perceive differences between pictures when too much information is provided).\r\n\r\n\r\n----------\r\n\r\n\r\n## I. Converting the image to binary\r\n\r\nThis step is probably the easiest (along with the binary stacking to RGB). You just need to use a binarization technique. Most of them incurs the following:\r\n\r\n 1. Conversion to 8-bit image\r\n 2. Conversion to binary using a preferred method with a black background\r\n\r\nHere are our three 8-bit images.\r\n\r\n![Montage aligned 8bit.jpg][3]\r\n\r\nWe convert the 8-bit images to binary per image (and not per stack) using black as background:\r\n\r\n![Montage aligned binary.jpg][4]\r\n\r\nUntil now... nothing hard right?\r\n\r\n\r\n----------\r\n\r\n\r\n## (Optional) II. Measuring the area covered by the binary images\r\n\r\nThis step is a good way to quantify the missing difference between two pictures. However, it quantifies the following:\r\n\r\n* Adds all the missing elements between two pictures\r\n* Subtracts all the appearing elements between two pictures\r\n\r\nThis is an issue you must take into account when you are analyzing a set of images! As you will compute the aggregated value (difference between the Add/Substract), you may end up with wrong interpretations using only that information.\r\n\r\nTo measure the area covered by the binary images, a first step is to threshold temporarily the images. All the white pixels must turn colored. If not, adjust appropriately the sliders of the software you are using.\r\n\r\nAfter, you must specify the area you need to measure. We select the area on the right to the left, omitting the road (you can notice in the image 2, the road is black while in the image 1/3 it is white.\r\n\r\n![Montage aligned binary.jpg][5]\r\n\r\n![Stack aligned selection.gif][6]\r\n\r\nNow, we can compute the area. We get the following:\r\n\r\n        Area\r\n    1 2693107\r\n    2 2704704\r\n    3 2819013\r\n\r\nDid we expect this ordering? Yes, because we hypothesized containers to disappear. Hence:\r\n\r\n* If a container has disappeared, its area becomes fully white\r\n* If a light colored container has appeared, its contour becomes black along with the shadows\r\n* If a dark colored container has appeared, its area become fully black along with the shadows\r\n\r\nThis confirms our hypothesis about the containers disappearing gradually from the image 1 to 2 to 3.\r\n\r\nThis is all good already!\r\n\r\n\r\n----------\r\n\r\n\r\n## III. Interpret the stack of binarized images and RGB\r\n\r\nThis step is the hardest of perform, although it is trivial for a set of only three images. You must convert the three pictures to RGB channels:\r\n\r\n* Image 1's whiteness become the Red channel (Red)\r\n* Image 2's whiteness become the Green channel (Green)\r\n* Image 3's whiteness become the Blue channel (Blue)\r\n\r\nRemember to remove the threshold filter if you did the previous step.\r\n\r\nYou get the following picture:\r\n\r\n![RGB stack.jpg][7]\r\n\r\nWhen composing three colors together, you need to remember how additive colors are:\r\n\r\n* Red => Red\r\n* Blue => Blue\r\n* Green => Green\r\n* Blue+Red => Magenta\r\n* Red+Green => Yellow\r\n* Green+Blue => Cyan\r\n* Red+Green+Blue => White\r\n* None => Black\r\n\r\nVisual image taken online:\r\n\r\n![Image Online][8]\r\n\r\nNote:\r\n\r\nIf you assume each picture is a color channel (picture 1/2/3 's whiteness => R/G/B), you get the following:\r\n\r\n    Color    1 2 3\r\n    Red      W B B\r\n    Green    B W B\r\n    Blue     B B W\r\n    Cyan     B W W\r\n    Magenta  W B W\r\n    Yellow   W W B\r\n    White    W W W\r\n    Black    B B B\r\n\r\nIf we suppose white (W) is presence (of missing) and black (B) is absence (of missing), we have the following:\r\n\r\n    Color    Presence  Absence\r\n    Red            1      2, 3\r\n    Green          2      1, 3\r\n    Blue           3      1, 2\r\n    Cyan        1, 2         1\r\n    Magenta     1, 3         2\r\n    Yellow      2, 3         3\r\n    White    1, 2, 3\r\n    Black              1, 2, 3\r\n\r\nTherefore, the pictures:\r\n\r\n* (1) should have the most red (disappearing in (2) and (3) )\r\n* (2) should have the most magenta (disappearing in (3) + the common elements with (1) )\r\n* (3) should have the most black (the common elements of (1) and (2) with (3) )\r\n\r\nHow is our RGB going? We just have to compute a count of unique values from a histogram (I selected a slightly different areas, so the values are different than part II.).\r\n\r\n    Selection count:   3789045\r\n    Grey (0) count:     797752 (none)\r\n    Grey (85) count:    267728\r\n    Grey (170) count:   309932\r\n    Grey (255) count:  2413633 (1, 2, 3)\r\n    Red (255) count:   2662863 (1 => present in 1 only)\r\n    Green (255) count: 2675696 (2 => present in 2 only)\r\n    Blue (255) count:  2789932 (3 => present in 3 only)\r\n\r\nThe grey values are not interesting for us (difficult to infer). The three last lines are the most interesting for us: if our hypothesis (stuff disappearing) is right, we should have Red > Green > Blue. This is the case, which also confirms our hypothesis.\r\n\r\nIf you want to look at the five stacks... not recommending it!\r\n\r\n![Five stacks.jpg][9]\r\n\r\n----------\r\n\r\n## IV. Deep learning method for differentiation\r\n\r\nWhat you can also do is using deep learning to compute the most probable sequence of images. This is a transformation of ranking to classification problem.\r\n\r\nYou would do the following:\r\n\r\n* For each set, create all existing combinations of images with the differences computed; {1, 2, 3, 4, 5}, {1, 2, 3, 5, 4}, {1, 2, 4, 3, 5}, {1, 2, 4, 5, 3}... overall you end up with 120 pictures for 1 set.\r\n* Transform the ranking issue to classification, assigning the label 1 for the correct sequence, and 0 for the incorrect (or vice-versa).\r\n* Train a deep learning model on the whole training set using the network of your choice.\r\n* Predict on the test set. Per set of 120 pictures, the highest probability picture is the most predicted right order by your deep learning network model.\r\n\r\nThere is a possible way also by subsampling each set by 3 pictures: instead of using the set {1, 2, 3, 4, 5}, you would label {1, 2, 3}, {2, 3, 4}, and {3, 4, 5} as right. However, you would label {1, 2, 4}, {1, 2, 5}, {1, 3, 2}... as wrong (total: 60 pictures per set, with 3 right and 57 wrong). You would then train a model, and the prediction would be the first unconflicting subset of the highest predictions per set.\r\n\r\n\r\n----------\r\n\r\n\r\nOff-topic: Hmm...?\r\n\r\n![enter image description here][10]\r\n\r\n\r\n  [1]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4320/Montage_initial.jpg\r\n  [2]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4321/Montage_aligned.jpg\r\n  [3]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4322/Montage_aligned_8bit.jpg\r\n  [4]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\r\n  [5]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\r\n  [6]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4324/Stack_aligned_selection.gif\r\n  [7]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4325/RGB_stack.jpg\r\n  [8]: http://hyperphysics.phy-astr.gsu.edu/hbase/vision/imgvis/addspotl.gif\r\n  [9]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4326/Five_stacks.jpg\r\n  [10]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4327/Kaggle_too_many_requests.JPG",
      "votes": null
    },
    {
      "id": "121379",
      "postDate": "05/26/2016 01:14:06",
      "content": "<p>Laurae,</p>\n\n<p>Thanks again for your all posts.\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.</p>",
      "rawMarkdown": "Laurae,\r\n\r\nThanks again for your all posts.\r\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.",
      "votes": null
    },
    {
      "id": "121381",
      "postDate": "05/26/2016 01:27:01",
      "content": "<p>[quote=eagle4;121379]</p>\n\n<p>Laurae,</p>\n\n<p>Thanks again for your all posts.\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.</p>\n\n<p>[/quote]</p>\n\n<p>Two steps from the original images:</p>\n\n<ul>\n<li>Registration</li>\n<li>Mathematical masking</li>\n</ul>\n\n<p>First, use a registration method. You have many algorithms available, like:</p>\n\n<ul>\n<li>KAZE/AKAZE</li>\n<li>SIFT</li>\n<li>Elastic alignment</li>\n<li>Unwarping</li>\n<li>Moving Least Squares</li>\n</ul>\n\n<p>If you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.</p>\n\n<p>Then <strong><em>(I guess it is the part you are interested in)</em></strong>, you use a typical mathematical filtering formula on pixels to &quot;crop&quot; properly using the numerical property of a complete black color:</p>\n\n<ul>\n<li>((((Image1 &lt; multiply &gt; Image2) &lt; multiply &gt; Image3) &lt; multiply &gt; Image4) &lt; multiply &gt; Image5) =&gt; ImageTemp (we use the latter as a mask)</li>\n<li>Image1 &lt; AND &gt; ImageTemp =&gt; Image1</li>\n<li>Image2 &lt; AND &gt; ImageTemp =&gt; Image2</li>\n<li>Image3 &lt; AND &gt; ImageTemp =&gt; Image3</li>\n<li>Image4 &lt; AND &gt; ImageTemp =&gt; Image4</li>\n<li>Image5 &lt; AND &gt; ImageTemp =&gt; Image5</li>\n</ul>\n\n<p>And you get the same black crops I have.</p>\n\n<p>In case something gets cropped out of nowhere, deviate all blacks by like 1/255 before aligning images, as one complete black has repercussions on all the others.</p>",
      "rawMarkdown": "[quote=eagle4;121379]\r\n\r\nLaurae,\r\n\r\nThanks again for your all posts.\r\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.\r\n\r\n[/quote]\r\n\r\nTwo steps from the original images:\r\n\r\n* Registration\r\n* Mathematical masking\r\n\r\nFirst, use a registration method. You have many algorithms available, like:\r\n\r\n* KAZE/AKAZE\r\n* SIFT\r\n* Elastic alignment\r\n* Unwarping\r\n* Moving Least Squares\r\n\r\nIf you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.\r\n\r\nThen ***(I guess it is the part you are interested in)***, you use a typical mathematical filtering formula on pixels to \"crop\" properly using the numerical property of a complete black color:\r\n\r\n* ((((Image1 < multiply > Image2) < multiply > Image3) < multiply > Image4) < multiply > Image5) => ImageTemp (we use the latter as a mask)\r\n* Image1 < AND > ImageTemp => Image1\r\n* Image2 < AND > ImageTemp => Image2\r\n* Image3 < AND > ImageTemp => Image3\r\n* Image4 < AND > ImageTemp => Image4\r\n* Image5 < AND > ImageTemp => Image5\r\n\r\nAnd you get the same black crops I have.\r\n\r\nIn case something gets cropped out of nowhere, deviate all blacks by like 1/255 before aligning images, as one complete black has repercussions on all the others.",
      "votes": null
    },
    {
      "id": "123302",
      "postDate": "06/10/2016 23:30:00",
      "content": "<p>@Laurae Nice Post</p>\n\n<p>Can we do something like this:</p>\n\n<p>Take a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 &amp; 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.</p>\n\n<p>Then same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.</p>\n\n<p>You can reduce errors by taking many rectangular portions and then comparing and voting.</p>",
      "rawMarkdown": "Laurae Nice Post\r\n\r\nCan we do something like this:\r\n\r\nTake a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 & 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.\r\n\r\nThen same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.\r\n\r\nYou can reduce errors by taking many rectangular portions and then comparing and voting.",
      "votes": null
    },
    {
      "id": "123756",
      "postDate": "06/13/2016 22:01:55",
      "content": "<p>[quote=Shahnawaz Akhtar;123302]</p>\n\n<p>@Laurae Nice Post</p>\n\n<p>Can we do something like this:</p>\n\n<p>Take a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 &amp; 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.</p>\n\n<p>Then same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.</p>\n\n<p>You can reduce errors by taking many rectangular portions and then comparing and voting.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, it is possible. However, you will not be able to sort out whether you are ordering them in the right way (1-2-3-4-5 or 5-4-3-2-1). If you are comparing with the &quot;nearest train set&quot;, this might be possible to order them the right way but your model will not be able to generalize to exotic new samples (such as on the ones with forests).</p>",
      "rawMarkdown": "[quote=Shahnawaz Akhtar;123302]\r\n\r\n@Laurae Nice Post\r\n\r\nCan we do something like this:\r\n\r\nTake a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 & 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.\r\n\r\nThen same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.\r\n\r\nYou can reduce errors by taking many rectangular portions and then comparing and voting.\r\n\r\n[/quote]\r\n\r\nYes, it is possible. However, you will not be able to sort out whether you are ordering them in the right way (1-2-3-4-5 or 5-4-3-2-1). If you are comparing with the \"nearest train set\", this might be possible to order them the right way but your model will not be able to generalize to exotic new samples (such as on the ones with forests).",
      "votes": null
    },
    {
      "id": "135414",
      "postDate": "09/13/2016 22:56:51",
      "content": "<p>[quote=Laurae;121381]</p>\n\n<p>Two steps from the original images:</p>\n\n<ul>\n<li>Registration</li>\n<li>Mathematical masking</li>\n</ul>\n\n<p>First, use a registration method. You have many algorithms available, like:</p>\n\n<ul>\n<li>KAZE/AKAZE</li>\n<li>SIFT</li>\n<li>Elastic alignment</li>\n<li>Unwarping</li>\n<li>Moving Least Squares</li>\n</ul>\n\n<p>If you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.</p>\n\n<p>[/quote]</p>\n\n<p>Can I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "[quote=Laurae;121381]\r\n\r\nTwo steps from the original images:\r\n\r\n* Registration\r\n* Mathematical masking\r\n\r\nFirst, use a registration method. You have many algorithms available, like:\r\n\r\n* KAZE/AKAZE\r\n* SIFT\r\n* Elastic alignment\r\n* Unwarping\r\n* Moving Least Squares\r\n\r\nIf you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.\r\n\r\n[/quote]\r\n\r\nCan I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?\r\n\r\nThanks!",
      "votes": null
    },
    {
      "id": "138861",
      "postDate": "10/11/2016 09:54:25",
      "content": "<p>[quote=mintaka;135414]</p>\n\n<p>Can I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?</p>\n\n<p>[/quote]</p>\n\n<p>Unwarping + Elastic Alignment. There are no performance metrics associated with these transformations as I used only one combination. You can try SIFT / KAZE/AKAZE if you want to go fast and well without &quot;tuning&quot; the parameters, as Unwarping + Elastic Alignment requires accurate tuning to not get unstable results. SIFT is included in ImageJ, AKAZE can be found in Python (there are scripts pre-cooked for that already in this competition if I remember).</p>",
      "rawMarkdown": "[quote=mintaka;135414]\r\n\r\nCan I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?\r\n\r\n[/quote]\r\n\r\nUnwarping + Elastic Alignment. There are no performance metrics associated with these transformations as I used only one combination. You can try SIFT / KAZE/AKAZE if you want to go fast and well without \"tuning\" the parameters, as Unwarping + Elastic Alignment requires accurate tuning to not get unstable results. SIFT is included in ImageJ, AKAZE can be found in Python (there are scripts pre-cooked for that already in this competition if I remember).",
      "votes": null
    },
    {
      "id": "2327689",
      "postDate": "07/03/2023 05:44:09",
      "content": "<p>Brightness &amp; contrast controlled Realtime RGB Histogram is here : <a href=\"https://youtu.be/k5rLn7VlAhI\" target=\"_blank\">https://youtu.be/k5rLn7VlAhI</a></p>",
      "rawMarkdown": "Brightness & contrast controlled Realtime RGB Histogram is here : https://youtu.be/k5rLn7VlAhI",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2327689,
      "author_name": "bemorekgg",
      "author_url": "",
      "post_date": "07/03/2023 05:44:09",
      "content": "<p>Brightness &amp; contrast controlled Realtime RGB Histogram is here : <a href=\"https://youtu.be/k5rLn7VlAhI\" target=\"_blank\">https://youtu.be/k5rLn7VlAhI</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121379,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "05/26/2016 01:14:06",
      "content": "<p>Laurae,</p>\n\n<p>Thanks again for your all posts.\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121381,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "05/26/2016 01:27:01",
      "content": "<p>[quote=eagle4;121379]</p>\n\n<p>Laurae,</p>\n\n<p>Thanks again for your all posts.\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.</p>\n\n<p>[/quote]</p>\n\n<p>Two steps from the original images:</p>\n\n<ul>\n<li>Registration</li>\n<li>Mathematical masking</li>\n</ul>\n\n<p>First, use a registration method. You have many algorithms available, like:</p>\n\n<ul>\n<li>KAZE/AKAZE</li>\n<li>SIFT</li>\n<li>Elastic alignment</li>\n<li>Unwarping</li>\n<li>Moving Least Squares</li>\n</ul>\n\n<p>If you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.</p>\n\n<p>Then <strong><em>(I guess it is the part you are interested in)</em></strong>, you use a typical mathematical filtering formula on pixels to &quot;crop&quot; properly using the numerical property of a complete black color:</p>\n\n<ul>\n<li>((((Image1 &lt; multiply &gt; Image2) &lt; multiply &gt; Image3) &lt; multiply &gt; Image4) &lt; multiply &gt; Image5) =&gt; ImageTemp (we use the latter as a mask)</li>\n<li>Image1 &lt; AND &gt; ImageTemp =&gt; Image1</li>\n<li>Image2 &lt; AND &gt; ImageTemp =&gt; Image2</li>\n<li>Image3 &lt; AND &gt; ImageTemp =&gt; Image3</li>\n<li>Image4 &lt; AND &gt; ImageTemp =&gt; Image4</li>\n<li>Image5 &lt; AND &gt; ImageTemp =&gt; Image5</li>\n</ul>\n\n<p>And you get the same black crops I have.</p>\n\n<p>In case something gets cropped out of nowhere, deviate all blacks by like 1/255 before aligning images, as one complete black has repercussions on all the others.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123302,
      "author_name": "shahnawazakhtar",
      "author_url": "",
      "post_date": "06/10/2016 23:30:00",
      "content": "<p>@Laurae Nice Post</p>\n\n<p>Can we do something like this:</p>\n\n<p>Take a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 &amp; 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.</p>\n\n<p>Then same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.</p>\n\n<p>You can reduce errors by taking many rectangular portions and then comparing and voting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123756,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "06/13/2016 22:01:55",
      "content": "<p>[quote=Shahnawaz Akhtar;123302]</p>\n\n<p>@Laurae Nice Post</p>\n\n<p>Can we do something like this:</p>\n\n<p>Take a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 &amp; 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.</p>\n\n<p>Then same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.</p>\n\n<p>You can reduce errors by taking many rectangular portions and then comparing and voting.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, it is possible. However, you will not be able to sort out whether you are ordering them in the right way (1-2-3-4-5 or 5-4-3-2-1). If you are comparing with the &quot;nearest train set&quot;, this might be possible to order them the right way but your model will not be able to generalize to exotic new samples (such as on the ones with forests).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 135414,
      "author_name": "mintaka",
      "author_url": "",
      "post_date": "09/13/2016 22:56:51",
      "content": "<p>[quote=Laurae;121381]</p>\n\n<p>Two steps from the original images:</p>\n\n<ul>\n<li>Registration</li>\n<li>Mathematical masking</li>\n</ul>\n\n<p>First, use a registration method. You have many algorithms available, like:</p>\n\n<ul>\n<li>KAZE/AKAZE</li>\n<li>SIFT</li>\n<li>Elastic alignment</li>\n<li>Unwarping</li>\n<li>Moving Least Squares</li>\n</ul>\n\n<p>If you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.</p>\n\n<p>[/quote]</p>\n\n<p>Can I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138861,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "10/11/2016 09:54:25",
      "content": "<p>[quote=mintaka;135414]</p>\n\n<p>Can I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?</p>\n\n<p>[/quote]</p>\n\n<p>Unwarping + Elastic Alignment. There are no performance metrics associated with these transformations as I used only one combination. You can try SIFT / KAZE/AKAZE if you want to go fast and well without &quot;tuning&quot; the parameters, as Unwarping + Elastic Alignment requires accurate tuning to not get unstable results. SIFT is included in ImageJ, AKAZE can be found in Python (there are scripts pre-cooked for that already in this competition if I remember).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121295": "Hello,\r\n\r\nThis is a tutorial to visualize the difference between a set of three different pictures. This technique generalizes to a 4 or more pictures (like 5 for Draper), although the interpretation becomes widely fuzzy (and difficult!). They can, however, be used for deep learning or any ML method you want.\r\n\r\nWe will use the three last pictures of the set 79 (given in the train set). We are assuming you already aligned the pictures. Our hypothesis is that there are (mostly) disappearing containers throughout the 3 images.\r\n\r\n![Montage initial.jpg][1]\r\n\r\n![Montage aligned.jpg][2]\r\n\r\n----------\r\n\r\n\r\nTo visualize (meaningfully) the difference between three pictures, we need to follow the following steps:\r\n\r\n* Convert the image to binary\r\n* Optional: measure the area covered by the binary images (good when analyzing moving camps, deforestation, disappearing objects, etc.)\r\n* Interpret the stack of binarized images to RGB\r\n\r\nThe idea of visualizing the difference meaningfully is to minimize (to a certain extent) the human bias in perception of three consecutive pictures, given the knowledge of their order (or to help the human to perceive differences between pictures when too much information is provided).\r\n\r\n\r\n----------\r\n\r\n\r\n## I. Converting the image to binary\r\n\r\nThis step is probably the easiest (along with the binary stacking to RGB). You just need to use a binarization technique. Most of them incurs the following:\r\n\r\n 1. Conversion to 8-bit image\r\n 2. Conversion to binary using a preferred method with a black background\r\n\r\nHere are our three 8-bit images.\r\n\r\n![Montage aligned 8bit.jpg][3]\r\n\r\nWe convert the 8-bit images to binary per image (and not per stack) using black as background:\r\n\r\n![Montage aligned binary.jpg][4]\r\n\r\nUntil now... nothing hard right?\r\n\r\n\r\n----------\r\n\r\n\r\n## (Optional) II. Measuring the area covered by the binary images\r\n\r\nThis step is a good way to quantify the missing difference between two pictures. However, it quantifies the following:\r\n\r\n* Adds all the missing elements between two pictures\r\n* Subtracts all the appearing elements between two pictures\r\n\r\nThis is an issue you must take into account when you are analyzing a set of images! As you will compute the aggregated value (difference between the Add/Substract), you may end up with wrong interpretations using only that information.\r\n\r\nTo measure the area covered by the binary images, a first step is to threshold temporarily the images. All the white pixels must turn colored. If not, adjust appropriately the sliders of the software you are using.\r\n\r\nAfter, you must specify the area you need to measure. We select the area on the right to the left, omitting the road (you can notice in the image 2, the road is black while in the image 1/3 it is white.\r\n\r\n![Montage aligned binary.jpg][5]\r\n\r\n![Stack aligned selection.gif][6]\r\n\r\nNow, we can compute the area. We get the following:\r\n\r\n        Area\r\n    1 2693107\r\n    2 2704704\r\n    3 2819013\r\n\r\nDid we expect this ordering? Yes, because we hypothesized containers to disappear. Hence:\r\n\r\n* If a container has disappeared, its area becomes fully white\r\n* If a light colored container has appeared, its contour becomes black along with the shadows\r\n* If a dark colored container has appeared, its area become fully black along with the shadows\r\n\r\nThis confirms our hypothesis about the containers disappearing gradually from the image 1 to 2 to 3.\r\n\r\nThis is all good already!\r\n\r\n\r\n----------\r\n\r\n\r\n## III. Interpret the stack of binarized images and RGB\r\n\r\nThis step is the hardest of perform, although it is trivial for a set of only three images. You must convert the three pictures to RGB channels:\r\n\r\n* Image 1's whiteness become the Red channel (Red)\r\n* Image 2's whiteness become the Green channel (Green)\r\n* Image 3's whiteness become the Blue channel (Blue)\r\n\r\nRemember to remove the threshold filter if you did the previous step.\r\n\r\nYou get the following picture:\r\n\r\n![RGB stack.jpg][7]\r\n\r\nWhen composing three colors together, you need to remember how additive colors are:\r\n\r\n* Red => Red\r\n* Blue => Blue\r\n* Green => Green\r\n* Blue+Red => Magenta\r\n* Red+Green => Yellow\r\n* Green+Blue => Cyan\r\n* Red+Green+Blue => White\r\n* None => Black\r\n\r\nVisual image taken online:\r\n\r\n![Image Online][8]\r\n\r\nNote:\r\n\r\nIf you assume each picture is a color channel (picture 1/2/3 's whiteness => R/G/B), you get the following:\r\n\r\n    Color    1 2 3\r\n    Red      W B B\r\n    Green    B W B\r\n    Blue     B B W\r\n    Cyan     B W W\r\n    Magenta  W B W\r\n    Yellow   W W B\r\n    White    W W W\r\n    Black    B B B\r\n\r\nIf we suppose white (W) is presence (of missing) and black (B) is absence (of missing), we have the following:\r\n\r\n    Color    Presence  Absence\r\n    Red            1      2, 3\r\n    Green          2      1, 3\r\n    Blue           3      1, 2\r\n    Cyan        1, 2         1\r\n    Magenta     1, 3         2\r\n    Yellow      2, 3         3\r\n    White    1, 2, 3\r\n    Black              1, 2, 3\r\n\r\nTherefore, the pictures:\r\n\r\n* (1) should have the most red (disappearing in (2) and (3) )\r\n* (2) should have the most magenta (disappearing in (3) + the common elements with (1) )\r\n* (3) should have the most black (the common elements of (1) and (2) with (3) )\r\n\r\nHow is our RGB going? We just have to compute a count of unique values from a histogram (I selected a slightly different areas, so the values are different than part II.).\r\n\r\n    Selection count:   3789045\r\n    Grey (0) count:     797752 (none)\r\n    Grey (85) count:    267728\r\n    Grey (170) count:   309932\r\n    Grey (255) count:  2413633 (1, 2, 3)\r\n    Red (255) count:   2662863 (1 => present in 1 only)\r\n    Green (255) count: 2675696 (2 => present in 2 only)\r\n    Blue (255) count:  2789932 (3 => present in 3 only)\r\n\r\nThe grey values are not interesting for us (difficult to infer). The three last lines are the most interesting for us: if our hypothesis (stuff disappearing) is right, we should have Red > Green > Blue. This is the case, which also confirms our hypothesis.\r\n\r\nIf you want to look at the five stacks... not recommending it!\r\n\r\n![Five stacks.jpg][9]\r\n\r\n----------\r\n\r\n## IV. Deep learning method for differentiation\r\n\r\nWhat you can also do is using deep learning to compute the most probable sequence of images. This is a transformation of ranking to classification problem.\r\n\r\nYou would do the following:\r\n\r\n* For each set, create all existing combinations of images with the differences computed; {1, 2, 3, 4, 5}, {1, 2, 3, 5, 4}, {1, 2, 4, 3, 5}, {1, 2, 4, 5, 3}... overall you end up with 120 pictures for 1 set.\r\n* Transform the ranking issue to classification, assigning the label 1 for the correct sequence, and 0 for the incorrect (or vice-versa).\r\n* Train a deep learning model on the whole training set using the network of your choice.\r\n* Predict on the test set. Per set of 120 pictures, the highest probability picture is the most predicted right order by your deep learning network model.\r\n\r\nThere is a possible way also by subsampling each set by 3 pictures: instead of using the set {1, 2, 3, 4, 5}, you would label {1, 2, 3}, {2, 3, 4}, and {3, 4, 5} as right. However, you would label {1, 2, 4}, {1, 2, 5}, {1, 3, 2}... as wrong (total: 60 pictures per set, with 3 right and 57 wrong). You would then train a model, and the prediction would be the first unconflicting subset of the highest predictions per set.\r\n\r\n\r\n----------\r\n\r\n\r\nOff-topic: Hmm...?\r\n\r\n![enter image description here][10]\r\n\r\n\r\n  [1]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4320/Montage_initial.jpg\r\n  [2]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4321/Montage_aligned.jpg\r\n  [3]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4322/Montage_aligned_8bit.jpg\r\n  [4]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\r\n  [5]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4323/Montage_aligned_binary.jpg\r\n  [6]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4324/Stack_aligned_selection.gif\r\n  [7]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4325/RGB_stack.jpg\r\n  [8]: http://hyperphysics.phy-astr.gsu.edu/hbase/vision/imgvis/addspotl.gif\r\n  [9]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4326/Five_stacks.jpg\r\n  [10]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4327/Kaggle_too_many_requests.JPG",
    "121379": "Laurae,\r\n\r\nThanks again for your all posts.\r\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.",
    "121381": "[quote=eagle4;121379]\r\n\r\nLaurae,\r\n\r\nThanks again for your all posts.\r\nSorry for my ignorance but how did you obtain the Montage_aligned.jpg ? I am having a hard time with this.\r\n\r\n[/quote]\r\n\r\nTwo steps from the original images:\r\n\r\n* Registration\r\n* Mathematical masking\r\n\r\nFirst, use a registration method. You have many algorithms available, like:\r\n\r\n* KAZE/AKAZE\r\n* SIFT\r\n* Elastic alignment\r\n* Unwarping\r\n* Moving Least Squares\r\n\r\nIf you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.\r\n\r\nThen ***(I guess it is the part you are interested in)***, you use a typical mathematical filtering formula on pixels to \"crop\" properly using the numerical property of a complete black color:\r\n\r\n* ((((Image1 < multiply > Image2) < multiply > Image3) < multiply > Image4) < multiply > Image5) => ImageTemp (we use the latter as a mask)\r\n* Image1 < AND > ImageTemp => Image1\r\n* Image2 < AND > ImageTemp => Image2\r\n* Image3 < AND > ImageTemp => Image3\r\n* Image4 < AND > ImageTemp => Image4\r\n* Image5 < AND > ImageTemp => Image5\r\n\r\nAnd you get the same black crops I have.\r\n\r\nIn case something gets cropped out of nowhere, deviate all blacks by like 1/255 before aligning images, as one complete black has repercussions on all the others.",
    "123302": "Laurae Nice Post\r\n\r\nCan we do something like this:\r\n\r\nTake a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 & 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.\r\n\r\nThen same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.\r\n\r\nYou can reduce errors by taking many rectangular portions and then comparing and voting.",
    "123756": "[quote=Shahnawaz Akhtar;123302]\r\n\r\n@Laurae Nice Post\r\n\r\nCan we do something like this:\r\n\r\nTake a small rectangular portion from image 1 of train set. Find that section in image 2,3,4 & 5. Note down transformations from original position that section goes through during each subsequent transformation by finding transformation of four corners. Then we can form a transformation vector say [tran1, trans2, trans3, trans4], where each translation vector element can be saved as quaternion.\r\n\r\nThen same procedure can be repeated in test image but this time since we can choose 5 images in 120 ways we need to make 120 transformation vectors for each set. Then we can classify based on nearest neighbour search. Since one of 120 transformation vector will be similar to some vector in train set, We can know the sequence that image was taken in because it will be same as sequence of nearest train set.\r\n\r\nYou can reduce errors by taking many rectangular portions and then comparing and voting.\r\n\r\n[/quote]\r\n\r\nYes, it is possible. However, you will not be able to sort out whether you are ordering them in the right way (1-2-3-4-5 or 5-4-3-2-1). If you are comparing with the \"nearest train set\", this might be possible to order them the right way but your model will not be able to generalize to exotic new samples (such as on the ones with forests).",
    "135414": "[quote=Laurae;121381]\r\n\r\nTwo steps from the original images:\r\n\r\n* Registration\r\n* Mathematical masking\r\n\r\nFirst, use a registration method. You have many algorithms available, like:\r\n\r\n* KAZE/AKAZE\r\n* SIFT\r\n* Elastic alignment\r\n* Unwarping\r\n* Moving Least Squares\r\n\r\nIf you are using ImageJ, it does not take long to learn how to register (align) images. You should get black parts in images which do not have matchings at least one image.\r\n\r\n[/quote]\r\n\r\nCan I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?\r\n\r\nThanks!",
    "138861": "[quote=mintaka;135414]\r\n\r\nCan I know which method(s) you were using in your experiments? And if you've tried some, what were their performances?\r\n\r\n[/quote]\r\n\r\nUnwarping + Elastic Alignment. There are no performance metrics associated with these transformations as I used only one combination. You can try SIFT / KAZE/AKAZE if you want to go fast and well without \"tuning\" the parameters, as Unwarping + Elastic Alignment requires accurate tuning to not get unstable results. SIFT is included in ImageJ, AKAZE can be found in Python (there are scripts pre-cooked for that already in this competition if I remember).",
    "2327689": "Brightness & contrast controlled Realtime RGB Histogram is here : https://youtu.be/k5rLn7VlAhI"
  },
  "source": "meta"
}