{
  "id": 21936,
  "title": "1st place - How to win the competition if you know nothing about image processing",
  "url": "/competitions/draper-satellite-image-chronology/discussion/21936",
  "author_name": "Vladimir Tomecek",
  "post_date": "2016-06-28T19:51:03.490000",
  "votes": 22,
  "comment_count": 10,
  "views": 2710,
  "content": "<p>For this competition I trained the neural network (the one inside my skull). </p>\n\n<p>It took about 2 weeks of staring at the images until the magic happened.</p>\n\n<p>Here are the main ideas.</p>\n\n<h2>High level algorithm:</h2>\n\n<ol>\n<li><p>Split the data by location</p></li>\n<li><p>For each location match together the pictures that were taken in the same day</p></li>\n<li><p>Solve each location using any clues from any picture set within that location</p></li>\n</ol>\n\n<h2>Ad 1:</h2>\n\n<p>We know that the pictures were taken in the southern California and on the pictures we see ports, airports, landfill, golf field, Point Loma Stadium ...</p>\n\n<p>So using Google Maps it is relatively easy to find all the locations:</p>\n\n<p>Train set: Carson (LA County), Port of Long Beach (LA County)</p>\n\n<p>Test set: San Diego (near downtown), Point Loma (W peninsula of SD), La Jolla (N of SD), Sycamore Landfill (E of San Diego/Santee), Sycamore Estates (N of landfill)</p>\n\n<p>It is not necessary to point the exact locations of the set, discovering the locations that were taken within one flight is sufficient. (I think that Sycamore Landfill + Sycamore Estates were taken within one flight, similarly San Diego + Point Loma and Carson + Long Beach, so there might be only 4 locations if we merge these together).</p>\n\n<h2>Ad 2:</h2>\n\n<p>The idea is that the plane flew into the area, took all the pictures within a couple of minutes and then flew away.\nThus if the pictures were taken in the morning, all the shadows should point to the left, if it was evening then they should point to the right, and during the noon to the top.\nIf it was cloudy during the flight over that location we should see only weak shadows (there was one such a day over the whole San Diego area, that was very easy to spot).\nIf we look at the map with marked locations, we see that the plane followed the straight lines when it took the pictures, so the neighboring images within that line should have very very similar orientation and zoom and in many cases also some overlap.</p>\n\n<p>Overlaps and shadows were most important features in this step (shadows were used mainly in cities because they very easier to evaluate by hand + they could bind the areas that didn't overlap; overlaps were the most reliable, but it took a little bit longer to match the pictures by hand using overlap than using shadows, they were mainly used in forests when the shadows very hardly visible, and when the two shadows seemed very similar).</p>\n\n<p>In some cases we could also use secondary features like presence/absense of cars at some location (indicator of the weekend), presense of clouds, presense of puddles after a rain. In the train set the zoom was an important feature.</p>\n\n<h2>Ad 3:</h2>\n\n<p>In this step we could use any clue from any set from the location to solve the whole location.</p>\n\n<p>Building sites, tracks in the mud, piles of sand were the most important and reliable features.</p>\n\n<p>Presence/absence of the cars in from of hypermarkets (weekend), presence of the cars in front of church (Sunday), presence of puddles after a rain (Point Loma) were also very important.</p>\n\n<p>Parking patterns of personal cars were not very reliable, because people use them every day and they may have a reserved or preffered parking place and the presence of the car at the very same place might not be a strong indicator of the closeness of that days. However the parking patterns of the bigger trucks and boats were pretty reliable.</p>",
  "messages": [
    {
      "id": 125346,
      "postDate": "2016-06-28T19:51:03.490Z",
      "content": "<p>For this competition I trained the neural network (the one inside my skull). </p>\n\n<p>It took about 2 weeks of staring at the images until the magic happened.</p>\n\n<p>Here are the main ideas.</p>\n\n<h2>High level algorithm:</h2>\n\n<ol>\n<li><p>Split the data by location</p></li>\n<li><p>For each location match together the pictures that were taken in the same day</p></li>\n<li><p>Solve each location using any clues from any picture set within that location</p></li>\n</ol>\n\n<h2>Ad 1:</h2>\n\n<p>We know that the pictures were taken in the southern California and on the pictures we see ports, airports, landfill, golf field, Point Loma Stadium ...</p>\n\n<p>So using Google Maps it is relatively easy to find all the locations:</p>\n\n<p>Train set: Carson (LA County), Port of Long Beach (LA County)</p>\n\n<p>Test set: San Diego (near downtown), Point Loma (W peninsula of SD), La Jolla (N of SD), Sycamore Landfill (E of San Diego/Santee), Sycamore Estates (N of landfill)</p>\n\n<p>It is not necessary to point the exact locations of the set, discovering the locations that were taken within one flight is sufficient. (I think that Sycamore Landfill + Sycamore Estates were taken within one flight, similarly San Diego + Point Loma and Carson + Long Beach, so there might be only 4 locations if we merge these together).</p>\n\n<h2>Ad 2:</h2>\n\n<p>The idea is that the plane flew into the area, took all the pictures within a couple of minutes and then flew away.\nThus if the pictures were taken in the morning, all the shadows should point to the left, if it was evening then they should point to the right, and during the noon to the top.\nIf it was cloudy during the flight over that location we should see only weak shadows (there was one such a day over the whole San Diego area, that was very easy to spot).\nIf we look at the map with marked locations, we see that the plane followed the straight lines when it took the pictures, so the neighboring images within that line should have very very similar orientation and zoom and in many cases also some overlap.</p>\n\n<p>Overlaps and shadows were most important features in this step (shadows were used mainly in cities because they very easier to evaluate by hand + they could bind the areas that didn't overlap; overlaps were the most reliable, but it took a little bit longer to match the pictures by hand using overlap than using shadows, they were mainly used in forests when the shadows very hardly visible, and when the two shadows seemed very similar).</p>\n\n<p>In some cases we could also use secondary features like presence/absense of cars at some location (indicator of the weekend), presense of clouds, presense of puddles after a rain. In the train set the zoom was an important feature.</p>\n\n<h2>Ad 3:</h2>\n\n<p>In this step we could use any clue from any set from the location to solve the whole location.</p>\n\n<p>Building sites, tracks in the mud, piles of sand were the most important and reliable features.</p>\n\n<p>Presence/absence of the cars in from of hypermarkets (weekend), presence of the cars in front of church (Sunday), presence of puddles after a rain (Point Loma) were also very important.</p>\n\n<p>Parking patterns of personal cars were not very reliable, because people use them every day and they may have a reserved or preffered parking place and the presence of the car at the very same place might not be a strong indicator of the closeness of that days. However the parking patterns of the bigger trucks and boats were pretty reliable.</p>",
      "rawMarkdown": "For this competition I trained the neural network (the one inside my skull). \r\n\r\nIt took about 2 weeks of staring at the images until the magic happened.\r\n\r\nHere are the main ideas.\r\n\r\n\r\nHigh level algorithm:\r\n---------------------\r\n\r\n\r\n 1. Split the data by location\r\n\r\n 2. For each location match together the pictures that were taken in the same day\r\n\r\n 3. Solve each location using any clues from any picture set within that location\r\n\r\n\r\nAd 1:\r\n-----\r\n\r\nWe know that the pictures were taken in the southern California and on the pictures we see ports, airports, landfill, golf field, Point Loma Stadium ...\r\n\r\nSo using Google Maps it is relatively easy to find all the locations:\r\n\r\nTrain set: Carson (LA County), Port of Long Beach (LA County)\r\n\r\nTest set: San Diego (near downtown), Point Loma (W peninsula of SD), La Jolla (N of SD), Sycamore Landfill (E of San Diego/Santee), Sycamore Estates (N of landfill)\r\n\r\nIt is not necessary to point the exact locations of the set, discovering the locations that were taken within one flight is sufficient. (I think that Sycamore Landfill + Sycamore Estates were taken within one flight, similarly San Diego + Point Loma and Carson + Long Beach, so there might be only 4 locations if we merge these together).\r\n\r\n\r\nAd 2:\r\n-----\r\n\r\nThe idea is that the plane flew into the area, took all the pictures within a couple of minutes and then flew away.\r\nThus if the pictures were taken in the morning, all the shadows should point to the left, if it was evening then they should point to the right, and during the noon to the top.\r\nIf it was cloudy during the flight over that location we should see only weak shadows (there was one such a day over the whole San Diego area, that was very easy to spot).\r\nIf we look at the map with marked locations, we see that the plane followed the straight lines when it took the pictures, so the neighboring images within that line should have very very similar orientation and zoom and in many cases also some overlap.\r\n\r\nOverlaps and shadows were most important features in this step (shadows were used mainly in cities because they very easier to evaluate by hand + they could bind the areas that didn't overlap; overlaps were the most reliable, but it took a little bit longer to match the pictures by hand using overlap than using shadows, they were mainly used in forests when the shadows very hardly visible, and when the two shadows seemed very similar).\r\n\r\nIn some cases we could also use secondary features like presence/absense of cars at some location (indicator of the weekend), presense of clouds, presense of puddles after a rain. In the train set the zoom was an important feature.\r\n\r\n\r\nAd 3:\r\n-----\r\n\r\nIn this step we could use any clue from any set from the location to solve the whole location.\r\n\r\nBuilding sites, tracks in the mud, piles of sand were the most important and reliable features.\r\n\r\nPresence/absence of the cars in from of hypermarkets (weekend), presence of the cars in front of church (Sunday), presence of puddles after a rain (Point Loma) were also very important.\r\n\r\nParking patterns of personal cars were not very reliable, because people use them every day and they may have a reserved or preffered parking place and the presence of the car at the very same place might not be a strong indicator of the closeness of that days. However the parking patterns of the bigger trucks and boats were pretty reliable.\r\n",
      "votes": 22
    },
    {
      "id": 125351,
      "postDate": "2016-06-28T20:41:09.347Z",
      "content": "<p>I have spent a couple of days on this, and I have solved about 5-10% of the images, the easiest 5-10%. In the end I would have probably managed to get close to 100% accuracy, but it was such hard work and the probability of being the 10th with the perfect score was so high, that I decided to pass the competition. A few weeks later Vladimir was already on the 1st place. Excellent effort. Now looking at what it took to get on top, I am glad I stepped out.</p>\n\n<p>I still think that what I said 2 months ago holds true: <a href=\"https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20565/leak-in-the-dataset/117709#post117709\">https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20565/leak-in-the-dataset/117709#post117709</a></p>\n\n<blockquote>\n  <p>This competition has little, if anything, to do with machine learning. It feels like it's a recruiting challenge for NSA or NRO.</p>\n</blockquote>\n\n<p>But I am not sure Vladimir has the right passport to work for those guys ;)</p>\n\n<p>Congratulations!</p>",
      "rawMarkdown": "I have spent a couple of days on this, and I have solved about 5-10% of the images, the easiest 5-10%. In the end I would have probably managed to get close to 100% accuracy, but it was such hard work and the probability of being the 10th with the perfect score was so high, that I decided to pass the competition. A few weeks later Vladimir was already on the 1st place. Excellent effort. Now looking at what it took to get on top, I am glad I stepped out.\r\n\r\nI still think that what I said 2 months ago holds true: https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20565/leak-in-the-dataset/117709#post117709\r\n\r\n> This competition has little, if anything, to do with machine learning. It feels like it's a recruiting challenge for NSA or NRO.\r\n\r\nBut I am not sure Vladimir has the right passport to work for those guys ;)\r\n\r\nCongratulations!\r\n\r\n\r\n",
      "votes": 3
    },
    {
      "id": 125428,
      "postDate": "2016-06-29T11:44:48.050Z",
      "content": "<p>[quote=the1owl;125397]</p>\n\n<p>Can you send me a copy of your Neural Network :) ?</p>\n\n<p>[/quote]</p>\n\n<p>I wouldn't send a copy of my brain to anyone!</p>\n\n<p>But this competition clearly showed that humans are superior to machines for this kind of task - for now.</p>\n\n<p>[quote=the1owl;125397]</p>\n\n<p>How do we know they were all taken in southern California (before solutions were posted)? </p>\n\n<p>[/quote]</p>\n\n<p>As Rich already said, it was in the data description.</p>\n\n<p>We even knew how the landfill looks like:</p>\n\n<p><a href=\"http://www.draper.com/news/cracking-code-satellite-data\">http://www.draper.com/news/cracking-code-satellite-data</a></p>\n\n<p>[quote=the1owl;125397]</p>\n\n<p>Did you use any automation or just browse through Google maps to find locations? </p>\n\n<p>[/quote]</p>\n\n<p>For the Step 1 I took advantage of how human brain works (associative memory). I first browsed throught all images in test set and tried to describe every image (e.g. pentagon building, route layout that resembles trident, rounded rectangles etc.) and kept those descriptions in the LibreOffice sheet.\nThen when I discovered a new location I just wandered around on Google Maps and the associations started to emerge. I then looked in the LibreOffice sheet for the reference so I don't have to browse through all image sets over and over again.</p>\n\n<p>The easiest locations to find was Port of Long Beach and Port of San Diego - I just followed the coastline. Carson was near the Long Beach and Point Loma near San Diego (plus it has also this eye-catching Point Loma Stadium).\nThe others were a little bit harder, but when you look carefully you see that all sets in test set have similar light conditions (e.g. one day with practically no shadow). So it was clear that they should be somewhere near San Diego. La Jolla was detectable by its golf field and the landfill is such a big object that it should catch your eyes sooner or later. The last one (Sycamore Estates) was the hardest, I discovered it by following the power lines from landfill area.</p>\n\n<p>After my first submission (score ~0.95) I used stitching to the background and visual inspection, but it turned out that I had it correct on the first attempt, I just made a mistake for Step3 in San Diego and Point Loma.</p>\n\n<p>I don't know if it comply with the Google rules, because they gave me 24h ban for auto-downloading their maps and after that I had to increase the sleep period significantly in order to not receive it again (or switch to Bing Maps).</p>\n\n<p>Thankfully this is only optional step. Moreover it didn't work well over the landfill and other locations that looks differently on aerial pictures than on satellite images of Google. It also had problems with the cloudy pictures.</p>\n\n<p>Alternative to this is intra-set stitching like you did here:</p>\n\n<p><a href=\"https://www.kaggle.com/the1owl/draper-satellite-image-chronology/stitch-and-predict\">https://www.kaggle.com/the1owl/draper-satellite-image-chronology/stitch-and-predict</a></p>\n\n<p>and then measuring relative rotation, zoom and shift;\nthe problem is that there are some sets that do not follow the common pattern, so you need to do outlier detection and then pass these sets to human for the fix. Plus there is still problem with cloudy pictures.</p>\n\n<p>Or you can do inter-set stitching, but this time you need to define the neighbors (otherwise it would be O(n^2)), but not all neighbors overlaps, some only have the same rotation/zoom, but shows very little or no overlap.</p>\n\n<p>So the best approach would be to do combination of intra and inter-set stitching with predefined neighbor relationships backed-up with the outlier detection so that human can fix possible problems and this is so difficult that given the size of the dataset you would prefer to do it by hand.</p>\n\n<p>By hand it wasn't that hard, once I trained &quot;the neural network&quot; it went pretty fast. There was one day with practically no shadow, one day with the shadow to the left, one with the shadow to the right and this held true for practically every image set. The last two days could be usually separated in the sets that had an asphalt road (when you looked carefully one shadow was usually weaker or was oriented a little bit more to the left than the other shadow). And for the other sets I matched them to the already resolved sets by other features like overlap/orientation or in one case (landfill) I had to use &quot;presence of the cars&quot; feature to distinguish between two days.</p>",
      "rawMarkdown": "[quote=the1owl;125397]\r\n\r\nCan you send me a copy of your Neural Network :) ?\r\n\r\n[/quote]\r\n\r\nI wouldn't send a copy of my brain to anyone!\r\n\r\nBut this competition clearly showed that humans are superior to machines for this kind of task - for now.\r\n\r\n[quote=the1owl;125397]\r\n\r\nHow do we know they were all taken in southern California (before solutions were posted)? \r\n\r\n[/quote]\r\n\r\nAs Rich already said, it was in the data description.\r\n\r\nWe even knew how the landfill looks like:\r\n\r\nhttp://www.draper.com/news/cracking-code-satellite-data\r\n\r\n[quote=the1owl;125397]\r\n\r\nDid you use any automation or just browse through Google maps to find locations? \r\n\r\n[/quote]\r\n\r\nFor the Step 1 I took advantage of how human brain works (associative memory). I first browsed throught all images in test set and tried to describe every image (e.g. pentagon building, route layout that resembles trident, rounded rectangles etc.) and kept those descriptions in the LibreOffice sheet.\r\nThen when I discovered a new location I just wandered around on Google Maps and the associations started to emerge. I then looked in the LibreOffice sheet for the reference so I don't have to browse through all image sets over and over again.\r\n\r\nThe easiest locations to find was Port of Long Beach and Port of San Diego - I just followed the coastline. Carson was near the Long Beach and Point Loma near San Diego (plus it has also this eye-catching Point Loma Stadium).\r\nThe others were a little bit harder, but when you look carefully you see that all sets in test set have similar light conditions (e.g. one day with practically no shadow). So it was clear that they should be somewhere near San Diego. La Jolla was detectable by its golf field and the landfill is such a big object that it should catch your eyes sooner or later. The last one (Sycamore Estates) was the hardest, I discovered it by following the power lines from landfill area.\r\n\r\n\r\nAfter my first submission (score ~0.95) I used stitching to the background and visual inspection, but it turned out that I had it correct on the first attempt, I just made a mistake for Step3 in San Diego and Point Loma.\r\n\r\nI don't know if it comply with the Google rules, because they gave me 24h ban for auto-downloading their maps and after that I had to increase the sleep period significantly in order to not receive it again (or switch to Bing Maps).\r\n\r\nThankfully this is only optional step. Moreover it didn't work well over the landfill and other locations that looks differently on aerial pictures than on satellite images of Google. It also had problems with the cloudy pictures.\r\n\r\nAlternative to this is intra-set stitching like you did here:\r\n\r\nhttps://www.kaggle.com/the1owl/draper-satellite-image-chronology/stitch-and-predict\r\n\r\nand then measuring relative rotation, zoom and shift;\r\nthe problem is that there are some sets that do not follow the common pattern, so you need to do outlier detection and then pass these sets to human for the fix. Plus there is still problem with cloudy pictures.\r\n\r\nOr you can do inter-set stitching, but this time you need to define the neighbors (otherwise it would be O(n^2)), but not all neighbors overlaps, some only have the same rotation/zoom, but shows very little or no overlap.\r\n\r\nSo the best approach would be to do combination of intra and inter-set stitching with predefined neighbor relationships backed-up with the outlier detection so that human can fix possible problems and this is so difficult that given the size of the dataset you would prefer to do it by hand.\r\n\r\nBy hand it wasn't that hard, once I trained \"the neural network\" it went pretty fast. There was one day with practically no shadow, one day with the shadow to the left, one with the shadow to the right and this held true for practically every image set. The last two days could be usually separated in the sets that had an asphalt road (when you looked carefully one shadow was usually weaker or was oriented a little bit more to the left than the other shadow). And for the other sets I matched them to the already resolved sets by other features like overlap/orientation or in one case (landfill) I had to use \"presence of the cars\" feature to distinguish between two days.\r\n",
      "votes": 1
    },
    {
      "id": 125491,
      "postDate": "2016-06-29T22:52:44.697Z",
      "content": "<p>[quote=Dmitry Grishin;125485]</p>\n\n<p>Vladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?</p>\n\n<p>[/quote]</p>\n\n<p>Stitching to the background (background = satellite images from Google Maps).</p>\n\n<p><strong>Location 1:</strong>\nPoint Loma (dayA = 5th day after solving) + San Diego (dayC = 5th day)</p>\n\n<p><strong>Location 2:</strong>\nLandfill + Sycamore Estates (both dayA = 2nd day)</p>\n\n<p><strong>Location 3 (LA):</strong>\nPort + Carson (3rd day)</p>\n\n<p><strong>Location 4:</strong>\nLa Jolla (dayA = 1st day, dayB = 2nd day; see the post above)</p>\n\n<p>Landfill didn't stitch properly because the landfill on the aerial images is little bit different than the landfill on the satellite images. Maybe if I used that image from Draper's web as a background it would be a better fit.</p>\n\n<p>Port have a lot of water around + the arrangement of the containers change every day, so the stitching is epic fail. I chose dayC because it produced best results.</p>",
      "rawMarkdown": "[quote=Dmitry Grishin;125485]\r\n\r\nVladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?\r\n\r\n[/quote]\r\n\r\nStitching to the background (background = satellite images from Google Maps).\r\n\r\n**Location 1:**\r\nPoint Loma (dayA = 5th day after solving) + San Diego (dayC = 5th day)\r\n\r\n**Location 2:**\r\nLandfill + Sycamore Estates (both dayA = 2nd day)\r\n\r\n**Location 3 (LA):**\r\nPort + Carson (3rd day)\r\n\r\n**Location 4:**\r\nLa Jolla (dayA = 1st day, dayB = 2nd day; see the post above)\r\n\r\nLandfill didn't stitch properly because the landfill on the aerial images is little bit different than the landfill on the satellite images. Maybe if I used that image from Draper's web as a background it would be a better fit.\r\n\r\nPort have a lot of water around + the arrangement of the containers change every day, so the stitching is epic fail. I chose dayC because it produced best results.",
      "votes": 2
    },
    {
      "id": 125541,
      "postDate": "2016-06-30T07:53:21.280Z",
      "content": "<p>To follow the fine tradition of <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20535/it-s-time-to-play-the-lb-shake-up-prediction-game/117513#post117513\">Shake-up</a>, this one earned 0.185 -- third place.</p>",
      "rawMarkdown": "To follow the fine tradition of [Shake-up][1], this one earned 0.185 -- third place.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20535/it-s-time-to-play-the-lb-shake-up-prediction-game/117513#post117513"
    },
    {
      "id": 125485,
      "postDate": "2016-06-29T21:47:17.543Z",
      "content": "<p>Vladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?</p>",
      "rawMarkdown": "Vladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?"
    },
    {
      "id": 125478,
      "postDate": "2016-06-29T20:58:02.653Z",
      "content": "<p>4 locations! that's why test set had not 100% homogeneous rotations and zoom. That's what I did not think of, thank you Vladimir!</p>\n\n<p>BTW I thought you built airplane trajectory for each day on Google Maps to predict.</p>",
      "rawMarkdown": "4 locations! that's why test set had not 100% homogeneous rotations and zoom. That's what I did not think of, thank you Vladimir!\r\n\r\nBTW I thought you built airplane trajectory for each day on Google Maps to predict."
    },
    {
      "id": 125418,
      "postDate": "2016-06-29T10:20:12.817Z",
      "content": "<p>In response to the1owl:</p>\n\n<p>Under the description of the dataset, we are told:</p>\n\n<p>&quot;This dataset contains over one thousand high-resolution images of aerial photographs taken in southern California. The photographs were taken from a plane and are meant as a reasonable facsimile for satellite images.&quot;</p>\n\n<p>-Rich</p>",
      "rawMarkdown": "In response to the1owl:\r\n\r\nUnder the description of the dataset, we are told:\r\n\r\n\"This dataset contains over one thousand high-resolution images of aerial photographs taken in southern California. The photographs were taken from a plane and are meant as a reasonable facsimile for satellite images.\"\r\n\r\n-Rich"
    },
    {
      "id": 125408,
      "postDate": "2016-06-29T07:29:44.147Z",
      "content": "<p>Congrats Vladimir. Good job.</p>",
      "rawMarkdown": "Congrats Vladimir. Good job."
    },
    {
      "id": 125450,
      "postDate": "2016-06-29T15:24:23.663Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 125397,
      "postDate": "2016-06-29T04:51:18.107Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 125351,
      "author_name": "Radu Stoicescu",
      "author_url": "",
      "post_date": "2016-06-28T20:41:09.347000",
      "content": "<p>I have spent a couple of days on this, and I have solved about 5-10% of the images, the easiest 5-10%. In the end I would have probably managed to get close to 100% accuracy, but it was such hard work and the probability of being the 10th with the perfect score was so high, that I decided to pass the competition. A few weeks later Vladimir was already on the 1st place. Excellent effort. Now looking at what it took to get on top, I am glad I stepped out.</p>\n\n<p>I still think that what I said 2 months ago holds true: <a href=\"https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20565/leak-in-the-dataset/117709#post117709\">https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20565/leak-in-the-dataset/117709#post117709</a></p>\n\n<blockquote>\n  <p>This competition has little, if anything, to do with machine learning. It feels like it's a recruiting challenge for NSA or NRO.</p>\n</blockquote>\n\n<p>But I am not sure Vladimir has the right passport to work for those guys ;)</p>\n\n<p>Congratulations!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 125428,
      "author_name": "Vladimir Tomecek",
      "author_url": "",
      "post_date": "2016-06-29T11:44:48.050000",
      "content": "<p>[quote=the1owl;125397]</p>\n\n<p>Can you send me a copy of your Neural Network :) ?</p>\n\n<p>[/quote]</p>\n\n<p>I wouldn't send a copy of my brain to anyone!</p>\n\n<p>But this competition clearly showed that humans are superior to machines for this kind of task - for now.</p>\n\n<p>[quote=the1owl;125397]</p>\n\n<p>How do we know they were all taken in southern California (before solutions were posted)? </p>\n\n<p>[/quote]</p>\n\n<p>As Rich already said, it was in the data description.</p>\n\n<p>We even knew how the landfill looks like:</p>\n\n<p><a href=\"http://www.draper.com/news/cracking-code-satellite-data\">http://www.draper.com/news/cracking-code-satellite-data</a></p>\n\n<p>[quote=the1owl;125397]</p>\n\n<p>Did you use any automation or just browse through Google maps to find locations? </p>\n\n<p>[/quote]</p>\n\n<p>For the Step 1 I took advantage of how human brain works (associative memory). I first browsed throught all images in test set and tried to describe every image (e.g. pentagon building, route layout that resembles trident, rounded rectangles etc.) and kept those descriptions in the LibreOffice sheet.\nThen when I discovered a new location I just wandered around on Google Maps and the associations started to emerge. I then looked in the LibreOffice sheet for the reference so I don't have to browse through all image sets over and over again.</p>\n\n<p>The easiest locations to find was Port of Long Beach and Port of San Diego - I just followed the coastline. Carson was near the Long Beach and Point Loma near San Diego (plus it has also this eye-catching Point Loma Stadium).\nThe others were a little bit harder, but when you look carefully you see that all sets in test set have similar light conditions (e.g. one day with practically no shadow). So it was clear that they should be somewhere near San Diego. La Jolla was detectable by its golf field and the landfill is such a big object that it should catch your eyes sooner or later. The last one (Sycamore Estates) was the hardest, I discovered it by following the power lines from landfill area.</p>\n\n<p>After my first submission (score ~0.95) I used stitching to the background and visual inspection, but it turned out that I had it correct on the first attempt, I just made a mistake for Step3 in San Diego and Point Loma.</p>\n\n<p>I don't know if it comply with the Google rules, because they gave me 24h ban for auto-downloading their maps and after that I had to increase the sleep period significantly in order to not receive it again (or switch to Bing Maps).</p>\n\n<p>Thankfully this is only optional step. Moreover it didn't work well over the landfill and other locations that looks differently on aerial pictures than on satellite images of Google. It also had problems with the cloudy pictures.</p>\n\n<p>Alternative to this is intra-set stitching like you did here:</p>\n\n<p><a href=\"https://www.kaggle.com/the1owl/draper-satellite-image-chronology/stitch-and-predict\">https://www.kaggle.com/the1owl/draper-satellite-image-chronology/stitch-and-predict</a></p>\n\n<p>and then measuring relative rotation, zoom and shift;\nthe problem is that there are some sets that do not follow the common pattern, so you need to do outlier detection and then pass these sets to human for the fix. Plus there is still problem with cloudy pictures.</p>\n\n<p>Or you can do inter-set stitching, but this time you need to define the neighbors (otherwise it would be O(n^2)), but not all neighbors overlaps, some only have the same rotation/zoom, but shows very little or no overlap.</p>\n\n<p>So the best approach would be to do combination of intra and inter-set stitching with predefined neighbor relationships backed-up with the outlier detection so that human can fix possible problems and this is so difficult that given the size of the dataset you would prefer to do it by hand.</p>\n\n<p>By hand it wasn't that hard, once I trained &quot;the neural network&quot; it went pretty fast. There was one day with practically no shadow, one day with the shadow to the left, one with the shadow to the right and this held true for practically every image set. The last two days could be usually separated in the sets that had an asphalt road (when you looked carefully one shadow was usually weaker or was oriented a little bit more to the left than the other shadow). And for the other sets I matched them to the already resolved sets by other features like overlap/orientation or in one case (landfill) I had to use &quot;presence of the cars&quot; feature to distinguish between two days.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 125491,
      "author_name": "Vladimir Tomecek",
      "author_url": "",
      "post_date": "2016-06-29T22:52:44.697000",
      "content": "<p>[quote=Dmitry Grishin;125485]</p>\n\n<p>Vladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?</p>\n\n<p>[/quote]</p>\n\n<p>Stitching to the background (background = satellite images from Google Maps).</p>\n\n<p><strong>Location 1:</strong>\nPoint Loma (dayA = 5th day after solving) + San Diego (dayC = 5th day)</p>\n\n<p><strong>Location 2:</strong>\nLandfill + Sycamore Estates (both dayA = 2nd day)</p>\n\n<p><strong>Location 3 (LA):</strong>\nPort + Carson (3rd day)</p>\n\n<p><strong>Location 4:</strong>\nLa Jolla (dayA = 1st day, dayB = 2nd day; see the post above)</p>\n\n<p>Landfill didn't stitch properly because the landfill on the aerial images is little bit different than the landfill on the satellite images. Maybe if I used that image from Draper's web as a background it would be a better fit.</p>\n\n<p>Port have a lot of water around + the arrangement of the containers change every day, so the stitching is epic fail. I chose dayC because it produced best results.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 125541,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2016-06-30T07:53:21.280000",
      "content": "<p>To follow the fine tradition of <a href=\"https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20535/it-s-time-to-play-the-lb-shake-up-prediction-game/117513#post117513\">Shake-up</a>, this one earned 0.185 -- third place.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125485,
      "author_name": "Dmitry Grishin",
      "author_url": "",
      "post_date": "2016-06-29T21:47:17.543000",
      "content": "<p>Vladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125478,
      "author_name": "Dmitry Grishin",
      "author_url": "",
      "post_date": "2016-06-29T20:58:02.653000",
      "content": "<p>4 locations! that's why test set had not 100% homogeneous rotations and zoom. That's what I did not think of, thank you Vladimir!</p>\n\n<p>BTW I thought you built airplane trajectory for each day on Google Maps to predict.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125418,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2016-06-29T10:20:12.817000",
      "content": "<p>In response to the1owl:</p>\n\n<p>Under the description of the dataset, we are told:</p>\n\n<p>&quot;This dataset contains over one thousand high-resolution images of aerial photographs taken in southern California. The photographs were taken from a plane and are meant as a reasonable facsimile for satellite images.&quot;</p>\n\n<p>-Rich</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125408,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2016-06-29T07:29:44.147000",
      "content": "<p>Congrats Vladimir. Good job.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125450,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-29T15:24:23.663000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125397,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-29T04:51:18.107000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "125346": "For this competition I trained the neural network (the one inside my skull). \r\n\r\nIt took about 2 weeks of staring at the images until the magic happened.\r\n\r\nHere are the main ideas.\r\n\r\n\r\nHigh level algorithm:\r\n---------------------\r\n\r\n\r\n 1. Split the data by location\r\n\r\n 2. For each location match together the pictures that were taken in the same day\r\n\r\n 3. Solve each location using any clues from any picture set within that location\r\n\r\n\r\nAd 1:\r\n-----\r\n\r\nWe know that the pictures were taken in the southern California and on the pictures we see ports, airports, landfill, golf field, Point Loma Stadium ...\r\n\r\nSo using Google Maps it is relatively easy to find all the locations:\r\n\r\nTrain set: Carson (LA County), Port of Long Beach (LA County)\r\n\r\nTest set: San Diego (near downtown), Point Loma (W peninsula of SD), La Jolla (N of SD), Sycamore Landfill (E of San Diego/Santee), Sycamore Estates (N of landfill)\r\n\r\nIt is not necessary to point the exact locations of the set, discovering the locations that were taken within one flight is sufficient. (I think that Sycamore Landfill + Sycamore Estates were taken within one flight, similarly San Diego + Point Loma and Carson + Long Beach, so there might be only 4 locations if we merge these together).\r\n\r\n\r\nAd 2:\r\n-----\r\n\r\nThe idea is that the plane flew into the area, took all the pictures within a couple of minutes and then flew away.\r\nThus if the pictures were taken in the morning, all the shadows should point to the left, if it was evening then they should point to the right, and during the noon to the top.\r\nIf it was cloudy during the flight over that location we should see only weak shadows (there was one such a day over the whole San Diego area, that was very easy to spot).\r\nIf we look at the map with marked locations, we see that the plane followed the straight lines when it took the pictures, so the neighboring images within that line should have very very similar orientation and zoom and in many cases also some overlap.\r\n\r\nOverlaps and shadows were most important features in this step (shadows were used mainly in cities because they very easier to evaluate by hand + they could bind the areas that didn't overlap; overlaps were the most reliable, but it took a little bit longer to match the pictures by hand using overlap than using shadows, they were mainly used in forests when the shadows very hardly visible, and when the two shadows seemed very similar).\r\n\r\nIn some cases we could also use secondary features like presence/absense of cars at some location (indicator of the weekend), presense of clouds, presense of puddles after a rain. In the train set the zoom was an important feature.\r\n\r\n\r\nAd 3:\r\n-----\r\n\r\nIn this step we could use any clue from any set from the location to solve the whole location.\r\n\r\nBuilding sites, tracks in the mud, piles of sand were the most important and reliable features.\r\n\r\nPresence/absence of the cars in from of hypermarkets (weekend), presence of the cars in front of church (Sunday), presence of puddles after a rain (Point Loma) were also very important.\r\n\r\nParking patterns of personal cars were not very reliable, because people use them every day and they may have a reserved or preffered parking place and the presence of the car at the very same place might not be a strong indicator of the closeness of that days. However the parking patterns of the bigger trucks and boats were pretty reliable.\r\n",
    "125351": "I have spent a couple of days on this, and I have solved about 5-10% of the images, the easiest 5-10%. In the end I would have probably managed to get close to 100% accuracy, but it was such hard work and the probability of being the 10th with the perfect score was so high, that I decided to pass the competition. A few weeks later Vladimir was already on the 1st place. Excellent effort. Now looking at what it took to get on top, I am glad I stepped out.\r\n\r\nI still think that what I said 2 months ago holds true: https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20565/leak-in-the-dataset/117709#post117709\r\n\r\n> This competition has little, if anything, to do with machine learning. It feels like it's a recruiting challenge for NSA or NRO.\r\n\r\nBut I am not sure Vladimir has the right passport to work for those guys ;)\r\n\r\nCongratulations!\r\n\r\n\r\n",
    "125428": "[quote=the1owl;125397]\r\n\r\nCan you send me a copy of your Neural Network :) ?\r\n\r\n[/quote]\r\n\r\nI wouldn't send a copy of my brain to anyone!\r\n\r\nBut this competition clearly showed that humans are superior to machines for this kind of task - for now.\r\n\r\n[quote=the1owl;125397]\r\n\r\nHow do we know they were all taken in southern California (before solutions were posted)? \r\n\r\n[/quote]\r\n\r\nAs Rich already said, it was in the data description.\r\n\r\nWe even knew how the landfill looks like:\r\n\r\nhttp://www.draper.com/news/cracking-code-satellite-data\r\n\r\n[quote=the1owl;125397]\r\n\r\nDid you use any automation or just browse through Google maps to find locations? \r\n\r\n[/quote]\r\n\r\nFor the Step 1 I took advantage of how human brain works (associative memory). I first browsed throught all images in test set and tried to describe every image (e.g. pentagon building, route layout that resembles trident, rounded rectangles etc.) and kept those descriptions in the LibreOffice sheet.\r\nThen when I discovered a new location I just wandered around on Google Maps and the associations started to emerge. I then looked in the LibreOffice sheet for the reference so I don't have to browse through all image sets over and over again.\r\n\r\nThe easiest locations to find was Port of Long Beach and Port of San Diego - I just followed the coastline. Carson was near the Long Beach and Point Loma near San Diego (plus it has also this eye-catching Point Loma Stadium).\r\nThe others were a little bit harder, but when you look carefully you see that all sets in test set have similar light conditions (e.g. one day with practically no shadow). So it was clear that they should be somewhere near San Diego. La Jolla was detectable by its golf field and the landfill is such a big object that it should catch your eyes sooner or later. The last one (Sycamore Estates) was the hardest, I discovered it by following the power lines from landfill area.\r\n\r\n\r\nAfter my first submission (score ~0.95) I used stitching to the background and visual inspection, but it turned out that I had it correct on the first attempt, I just made a mistake for Step3 in San Diego and Point Loma.\r\n\r\nI don't know if it comply with the Google rules, because they gave me 24h ban for auto-downloading their maps and after that I had to increase the sleep period significantly in order to not receive it again (or switch to Bing Maps).\r\n\r\nThankfully this is only optional step. Moreover it didn't work well over the landfill and other locations that looks differently on aerial pictures than on satellite images of Google. It also had problems with the cloudy pictures.\r\n\r\nAlternative to this is intra-set stitching like you did here:\r\n\r\nhttps://www.kaggle.com/the1owl/draper-satellite-image-chronology/stitch-and-predict\r\n\r\nand then measuring relative rotation, zoom and shift;\r\nthe problem is that there are some sets that do not follow the common pattern, so you need to do outlier detection and then pass these sets to human for the fix. Plus there is still problem with cloudy pictures.\r\n\r\nOr you can do inter-set stitching, but this time you need to define the neighbors (otherwise it would be O(n^2)), but not all neighbors overlaps, some only have the same rotation/zoom, but shows very little or no overlap.\r\n\r\nSo the best approach would be to do combination of intra and inter-set stitching with predefined neighbor relationships backed-up with the outlier detection so that human can fix possible problems and this is so difficult that given the size of the dataset you would prefer to do it by hand.\r\n\r\nBy hand it wasn't that hard, once I trained \"the neural network\" it went pretty fast. There was one day with practically no shadow, one day with the shadow to the left, one with the shadow to the right and this held true for practically every image set. The last two days could be usually separated in the sets that had an asphalt road (when you looked carefully one shadow was usually weaker or was oriented a little bit more to the left than the other shadow). And for the other sets I matched them to the already resolved sets by other features like overlap/orientation or in one case (landfill) I had to use \"presence of the cars\" feature to distinguish between two days.\r\n",
    "125491": "[quote=Dmitry Grishin;125485]\r\n\r\nVladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?\r\n\r\n[/quote]\r\n\r\nStitching to the background (background = satellite images from Google Maps).\r\n\r\n**Location 1:**\r\nPoint Loma (dayA = 5th day after solving) + San Diego (dayC = 5th day)\r\n\r\n**Location 2:**\r\nLandfill + Sycamore Estates (both dayA = 2nd day)\r\n\r\n**Location 3 (LA):**\r\nPort + Carson (3rd day)\r\n\r\n**Location 4:**\r\nLa Jolla (dayA = 1st day, dayB = 2nd day; see the post above)\r\n\r\nLandfill didn't stitch properly because the landfill on the aerial images is little bit different than the landfill on the satellite images. Maybe if I used that image from Draper's web as a background it would be a better fit.\r\n\r\nPort have a lot of water around + the arrangement of the containers change every day, so the stitching is epic fail. I chose dayC because it produced best results.",
    "125541": "To follow the fine tradition of [Shake-up][1], this one earned 0.185 -- third place.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/santander-customer-satisfaction/forums/t/20535/it-s-time-to-play-the-lb-shake-up-prediction-game/117513#post117513",
    "125485": "Vladimir, can you share day A compiled images for all 4 locations please? Do they all have same rotation within one day or mostly at least?",
    "125478": "4 locations! that's why test set had not 100% homogeneous rotations and zoom. That's what I did not think of, thank you Vladimir!\r\n\r\nBTW I thought you built airplane trajectory for each day on Google Maps to predict.",
    "125418": "In response to the1owl:\r\n\r\nUnder the description of the dataset, we are told:\r\n\r\n\"This dataset contains over one thousand high-resolution images of aerial photographs taken in southern California. The photographs were taken from a plane and are meant as a reasonable facsimile for satellite images.\"\r\n\r\n-Rich",
    "125408": "Congrats Vladimir. Good job.",
    "125450": "",
    "125397": ""
  }
}