{
  "id": 21052,
  "title": "Rules question: correlating overlapping same-time commercial imagery",
  "url": "/competitions/draper-satellite-image-chronology/discussion/21052",
  "author_name": "",
  "post_date": "2016-05-18T21:15:01.423Z",
  "votes": 2,
  "comment_count": 9,
  "views": 1480,
  "content": "<p>For approximately 252 test sets, overlapping commercial imagery exists that identifies the exact date that at least one photo from the set was taken.  The external imagery was captured within a few hours of the Draper commissioned set.  </p>\n\n<p>Example:</p>\n\n<p><img src=\"http://i.imgur.com/T1VFPJz.png\" alt=\"Example of overlap.\" title></p>\n\n<p>I know it has been clarified a few times that external imagery and reverse-engineering locations are allowed, but this seems a little different since it provides a flawless method for identifying which day many photos were taken.  It also makes other methods, like tide levels and solar noon shadow corroboration, much more straightforward to analyze.  It's still &quot;generalizable&quot; since it's probably impossible to capture aerial images from anywhere in the world that don't overlap with some other platform the same day.  </p>\n\n<p>Will final submissions be allowed if the output partially relied on correlating with such commercial imagery?</p>\n\n<p>I won't post the commercial source and the corresponding absolute date range the test photos were taken in unless the admins want those provided publicly.</p>",
  "messages": [
    {
      "id": "120526",
      "postDate": "05/18/2016 21:15:01",
      "content": "<p>For approximately 252 test sets, overlapping commercial imagery exists that identifies the exact date that at least one photo from the set was taken.  The external imagery was captured within a few hours of the Draper commissioned set.  </p>\n\n<p>Example:</p>\n\n<p><img src=\"http://i.imgur.com/T1VFPJz.png\" alt=\"Example of overlap.\" title></p>\n\n<p>I know it has been clarified a few times that external imagery and reverse-engineering locations are allowed, but this seems a little different since it provides a flawless method for identifying which day many photos were taken.  It also makes other methods, like tide levels and solar noon shadow corroboration, much more straightforward to analyze.  It's still &quot;generalizable&quot; since it's probably impossible to capture aerial images from anywhere in the world that don't overlap with some other platform the same day.  </p>\n\n<p>Will final submissions be allowed if the output partially relied on correlating with such commercial imagery?</p>\n\n<p>I won't post the commercial source and the corresponding absolute date range the test photos were taken in unless the admins want those provided publicly.</p>",
      "rawMarkdown": "For approximately 252 test sets, overlapping commercial imagery exists that identifies the exact date that at least one photo from the set was taken.  The external imagery was captured within a few hours of the Draper commissioned set.  \r\n\r\nExample:\r\n\r\n ![Example of overlap.][1]\r\n\r\nI know it has been clarified a few times that external imagery and reverse-engineering locations are allowed, but this seems a little different since it provides a flawless method for identifying which day many photos were taken.  It also makes other methods, like tide levels and solar noon shadow corroboration, much more straightforward to analyze.  It's still \"generalizable\" since it's probably impossible to capture aerial images from anywhere in the world that don't overlap with some other platform the same day.  \r\n\r\nWill final submissions be allowed if the output partially relied on correlating with such commercial imagery?\r\n\r\nI won't post the commercial source and the corresponding absolute date range the test photos were taken in unless the admins want those provided publicly.\r\n\r\n  [1]: http://i.imgur.com/T1VFPJz.png",
      "votes": null
    },
    {
      "id": "120540",
      "postDate": "05/18/2016 23:18:46",
      "content": "<p>Doesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...</p>",
      "rawMarkdown": "Doesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...",
      "votes": null
    },
    {
      "id": "120560",
      "postDate": "05/19/2016 03:36:19",
      "content": "<p>[quote=kes367;120540]</p>\n\n<p>Doesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...</p>\n\n<p>[/quote]\nNot exactly 0.96.  It's only one image in the set of five.  And it's also freely available commercial imagery.  </p>\n\n<p>Though I share your concern that it incentivizes a non-innovative race to the top.</p>",
      "rawMarkdown": "[quote=kes367;120540]\r\n\r\nDoesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...\r\n\r\n[/quote]\r\nNot exactly 0.96.  It's only one image in the set of five.  And it's also freely available commercial imagery.  \r\n\r\nThough I share your concern that it incentivizes a non-innovative race to the top.",
      "votes": null
    },
    {
      "id": "120562",
      "postDate": "05/19/2016 03:39:50",
      "content": "<p>Right, but if someone were able to find this, it isn't a huge stretch that they could purchase commercial satellite imagery for that date range around that time for a given location.</p>",
      "rawMarkdown": "Right, but if someone were able to find this, it isn't a huge stretch that they could purchase commercial satellite imagery for that date range around that time for a given location.",
      "votes": null
    },
    {
      "id": "120646",
      "postDate": "05/19/2016 18:43:43",
      "content": "<p>Thanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.</p>",
      "rawMarkdown": "Thanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.",
      "votes": null
    },
    {
      "id": "120657",
      "postDate": "05/19/2016 19:11:44",
      "content": "<p>[quote=William Cukierski;120646]</p>\n\n<p>Thanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.</p>\n\n<p>[/quote]\nThanks for the clarification!</p>",
      "rawMarkdown": "[quote=William Cukierski;120646]\r\n\r\nThanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.\r\n\r\n[/quote]\r\nThanks for the clarification!",
      "votes": null
    },
    {
      "id": "120682",
      "postDate": "05/19/2016 22:33:25",
      "content": "<p>I'm still not 100% clear on this.</p>\n\n<p>I haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is &quot;freely available&quot; (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?</p>\n\n<p>The second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not <em>directly</em> give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.</p>\n\n<p>So my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are &quot;freely available&quot; at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for <em>every</em> image, giving the answer up directly)</p>\n\n<p>I can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)</p>",
      "rawMarkdown": "I'm still not 100% clear on this.\r\n\r\nI haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is \"freely available\" (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?\r\n\r\nThe second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not _directly_ give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.\r\n\r\nSo my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are \"freely available\" at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for _every_ image, giving the answer up directly)\r\n\r\nI can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)",
      "votes": null
    },
    {
      "id": "120783",
      "postDate": "05/20/2016 14:44:15",
      "content": "<p>[quote=Jordie Fulton;120682]</p>\n\n<p>I'm still not 100% clear on this.</p>\n\n<p>I haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is &quot;freely available&quot; (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?</p>\n\n<p>The second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not <em>directly</em> give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.</p>\n\n<p>So my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are &quot;freely available&quot; at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for <em>every</em> image, giving the answer up directly)</p>\n\n<p>I can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)</p>\n\n<p>[/quote]\nThese are valid questions.  </p>\n\n<p>You're right that when I say free I mean free as in beer not free as in speech.  I certainly didn't pay for them; I just sifted through a lot of different aerial photo sites until I found some that overlapped.  You will probably always be able to do this in the United States, as evidenced by the introductory paragraph to this challenge which cites that we should have an image of everywhere every day by 2017.  But it might not generalize to, say, Antarctica or the middle of Africa or parts of Russia.  </p>\n\n<p>However, it's difficult to stay within the spirit of the contest when it's very unclear what the spirit of the contest actually is.  There is no real world application for putting satellite images in chronological order.  Most analysts (full disclosure: I'm an imagery analyst, not a ML programmer) won't even bother to look at an image if it doesn't contain sensor data that allows automatic orthorectification.  It's just bad practice.  Imagine if in the &quot;Ultrasound Nerve Segmentation&quot; contest we were only allowed to look at Polaroid photos of the subjects.  No doctor in their right mind would use that to find nerves, and there isn't a use case for it because far better methods exist.  It's interesting as a puzzle of sorts (which is why I'm here) but it's not practical.</p>\n\n<p>I also agree that it would be trivial to hand tag with the external data and come up with an in-image justification.  Alternatively, you might hand tag them and check the answers against the external data.  It seems that the latter technique should be allowed based on what William said, because:</p>\n\n<ul>\n<li>No one paid for imagery.</li>\n<li>The answers aren't given away in a direct manner.  The goal of the competition is to organize the images relative to each other and not to absolute dates on the calendar.  If only one image in the set is included, you would know little about where it fits in the set.</li>\n<li>It's within the spirit of the competition to be able to say both &quot;this photo was taken a few hours before the test photo&quot; and &quot;this test photo was taken the day after this other test photo.&quot;  Those are both temporal questions.</li>\n<li>There is a difference between the <em>approach</em> being open source and all the data utilized in the approach being open source.  We've already agreed that the data used can be external/copywritten, since Google Earth is allowed.</li>\n<li>It's difficult to know exactly what would help the host here.  The best practice approach to this problem (from my view) would be to analyze the <em>context</em> of the photographs well before the <em>content</em> of the photographs.  The date a photograph was taken is a context question, so outside information is the best way to answer it.  Using exclusively the contents of the image to answer the question would be a disservice to the host, since it would not be as consistently accurate.  If they had instead hosted a competition that asked &quot;how many vehicles are in this photograph?&quot; or &quot;what is the health of this forest?&quot; restricting outside information would make much more sense.</li>\n</ul>",
      "rawMarkdown": "[quote=Jordie Fulton;120682]\r\n\r\n\r\nI'm still not 100% clear on this.\r\n\r\nI haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is \"freely available\" (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?\r\n\r\nThe second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not _directly_ give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.\r\n\r\nSo my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are \"freely available\" at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for _every_ image, giving the answer up directly)\r\n\r\nI can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)\r\n\r\n\r\n[/quote]\r\nThese are valid questions.  \r\n\r\nYou're right that when I say free I mean free as in beer not free as in speech.  I certainly didn't pay for them; I just sifted through a lot of different aerial photo sites until I found some that overlapped.  You will probably always be able to do this in the United States, as evidenced by the introductory paragraph to this challenge which cites that we should have an image of everywhere every day by 2017.  But it might not generalize to, say, Antarctica or the middle of Africa or parts of Russia.  \r\n\r\nHowever, it's difficult to stay within the spirit of the contest when it's very unclear what the spirit of the contest actually is.  There is no real world application for putting satellite images in chronological order.  Most analysts (full disclosure: I'm an imagery analyst, not a ML programmer) won't even bother to look at an image if it doesn't contain sensor data that allows automatic orthorectification.  It's just bad practice.  Imagine if in the \"Ultrasound Nerve Segmentation\" contest we were only allowed to look at Polaroid photos of the subjects.  No doctor in their right mind would use that to find nerves, and there isn't a use case for it because far better methods exist.  It's interesting as a puzzle of sorts (which is why I'm here) but it's not practical.\r\n\r\nI also agree that it would be trivial to hand tag with the external data and come up with an in-image justification.  Alternatively, you might hand tag them and check the answers against the external data.  It seems that the latter technique should be allowed based on what William said, because:\r\n\r\n - No one paid for imagery.\r\n - The answers aren't given away in a direct manner.  The goal of the competition is to organize the images relative to each other and not to absolute dates on the calendar.  If only one image in the set is included, you would know little about where it fits in the set.\r\n - It's within the spirit of the competition to be able to say both \"this photo was taken a few hours before the test photo\" and \"this test photo was taken the day after this other test photo.\"  Those are both temporal questions.\r\n - There is a difference between the *approach* being open source and all the data utilized in the approach being open source.  We've already agreed that the data used can be external/copywritten, since Google Earth is allowed.\r\n - It's difficult to know exactly what would help the host here.  The best practice approach to this problem (from my view) would be to analyze the *context* of the photographs well before the *content* of the photographs.  The date a photograph was taken is a context question, so outside information is the best way to answer it.  Using exclusively the contents of the image to answer the question would be a disservice to the host, since it would not be as consistently accurate.  If they had instead hosted a competition that asked \"how many vehicles are in this photograph?\" or \"what is the health of this forest?\" restricting outside information would make much more sense.",
      "votes": null
    },
    {
      "id": "120817",
      "postDate": "05/20/2016 18:17:24",
      "content": "<p>All of the above mirrors my thinking -- I'm not particularly suggesting that one way or the other is more correct, but think it would be helpful if admin could address this specific issue with a clear yes/no on what is allowed/desirable for Draper. </p>\n\n<p>It is not clear to me either just what kind of solutions they are after, but it does seem as though the difficulty of the problem was overestimated. I find it particularly unlikely that &quot;align images with external data of known date&quot; is what they were thinking here, but the rules as stated so far do not seem to disallow it. :-)</p>\n\n<p>I am also enjoying the puzzle, and have been able to automate certain aspects, but more along the lines of an &quot;analyst's toolkit&quot; than machine learning. Curiously, I believe I can deduce specific dates for many of these particular images <em>without</em> using external imagery -- not the point of the contest as you say, but it is helpful to anchor time frames for the different sets.</p>",
      "rawMarkdown": "All of the above mirrors my thinking -- I'm not particularly suggesting that one way or the other is more correct, but think it would be helpful if admin could address this specific issue with a clear yes/no on what is allowed/desirable for Draper. \r\n\r\nIt is not clear to me either just what kind of solutions they are after, but it does seem as though the difficulty of the problem was overestimated. I find it particularly unlikely that \"align images with external data of known date\" is what they were thinking here, but the rules as stated so far do not seem to disallow it. :-)\r\n\r\nI am also enjoying the puzzle, and have been able to automate certain aspects, but more along the lines of an \"analyst's toolkit\" than machine learning. Curiously, I believe I can deduce specific dates for many of these particular images _without_ using external imagery -- not the point of the contest as you say, but it is helpful to anchor time frames for the different sets.",
      "votes": null
    },
    {
      "id": "120845",
      "postDate": "05/20/2016 22:03:26",
      "content": "<p>One should not use external images if their only redeeming contribution to the methodology is their timestamp. Going back to the spirit of the competition (namely, finding repeatable image features that distinguish images in time), such usage doesn't meet the aims of the host. The difficulty of enforcing this really depends on how much overlapping imagery is out there. We (Kaggle) were not aware of the extent of this external imagery at the time when we launched the competition and have, so far, only received reports of the one day's worth referenced earlier in this thread. No matter what, Draper will be reviewing the winners' methodology and have an eye out for suspicious post facto explanations.</p>\n\n<p>@Tyler re: &quot;Using exclusively the contents of the image to answer the question would be a disservice to the host&quot; Features based on image content is actually the goal here. It's not about the date/order these specific images were acquired, but rather in finding generalizable features within an image that give away the order (and beyond just the order, features that make other tasks possible). External data was permitted in hopes that people could construct more training data to improve the results.</p>",
      "rawMarkdown": "One should not use external images if their only redeeming contribution to the methodology is their timestamp. Going back to the spirit of the competition (namely, finding repeatable image features that distinguish images in time), such usage doesn't meet the aims of the host. The difficulty of enforcing this really depends on how much overlapping imagery is out there. We (Kaggle) were not aware of the extent of this external imagery at the time when we launched the competition and have, so far, only received reports of the one day's worth referenced earlier in this thread. No matter what, Draper will be reviewing the winners' methodology and have an eye out for suspicious post facto explanations.\r\n\r\n@Tyler re: \"Using exclusively the contents of the image to answer the question would be a disservice to the host\" Features based on image content is actually the goal here. It's not about the date/order these specific images were acquired, but rather in finding generalizable features within an image that give away the order (and beyond just the order, features that make other tasks possible). External data was permitted in hopes that people could construct more training data to improve the results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 120540,
      "author_name": "kes367",
      "author_url": "",
      "post_date": "05/18/2016 23:18:46",
      "content": "<p>Doesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120560,
      "author_name": "tylervigen",
      "author_url": "",
      "post_date": "05/19/2016 03:36:19",
      "content": "<p>[quote=kes367;120540]</p>\n\n<p>Doesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...</p>\n\n<p>[/quote]\nNot exactly 0.96.  It's only one image in the set of five.  And it's also freely available commercial imagery.  </p>\n\n<p>Though I share your concern that it incentivizes a non-innovative race to the top.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120562,
      "author_name": "kes367",
      "author_url": "",
      "post_date": "05/19/2016 03:39:50",
      "content": "<p>Right, but if someone were able to find this, it isn't a huge stretch that they could purchase commercial satellite imagery for that date range around that time for a given location.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120646,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "05/19/2016 18:43:43",
      "content": "<p>Thanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120657,
      "author_name": "tylervigen",
      "author_url": "",
      "post_date": "05/19/2016 19:11:44",
      "content": "<p>[quote=William Cukierski;120646]</p>\n\n<p>Thanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.</p>\n\n<p>[/quote]\nThanks for the clarification!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120682,
      "author_name": "jfulton",
      "author_url": "",
      "post_date": "05/19/2016 22:33:25",
      "content": "<p>I'm still not 100% clear on this.</p>\n\n<p>I haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is &quot;freely available&quot; (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?</p>\n\n<p>The second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not <em>directly</em> give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.</p>\n\n<p>So my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are &quot;freely available&quot; at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for <em>every</em> image, giving the answer up directly)</p>\n\n<p>I can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120783,
      "author_name": "tylervigen",
      "author_url": "",
      "post_date": "05/20/2016 14:44:15",
      "content": "<p>[quote=Jordie Fulton;120682]</p>\n\n<p>I'm still not 100% clear on this.</p>\n\n<p>I haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is &quot;freely available&quot; (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?</p>\n\n<p>The second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not <em>directly</em> give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.</p>\n\n<p>So my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are &quot;freely available&quot; at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for <em>every</em> image, giving the answer up directly)</p>\n\n<p>I can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)</p>\n\n<p>[/quote]\nThese are valid questions.  </p>\n\n<p>You're right that when I say free I mean free as in beer not free as in speech.  I certainly didn't pay for them; I just sifted through a lot of different aerial photo sites until I found some that overlapped.  You will probably always be able to do this in the United States, as evidenced by the introductory paragraph to this challenge which cites that we should have an image of everywhere every day by 2017.  But it might not generalize to, say, Antarctica or the middle of Africa or parts of Russia.  </p>\n\n<p>However, it's difficult to stay within the spirit of the contest when it's very unclear what the spirit of the contest actually is.  There is no real world application for putting satellite images in chronological order.  Most analysts (full disclosure: I'm an imagery analyst, not a ML programmer) won't even bother to look at an image if it doesn't contain sensor data that allows automatic orthorectification.  It's just bad practice.  Imagine if in the &quot;Ultrasound Nerve Segmentation&quot; contest we were only allowed to look at Polaroid photos of the subjects.  No doctor in their right mind would use that to find nerves, and there isn't a use case for it because far better methods exist.  It's interesting as a puzzle of sorts (which is why I'm here) but it's not practical.</p>\n\n<p>I also agree that it would be trivial to hand tag with the external data and come up with an in-image justification.  Alternatively, you might hand tag them and check the answers against the external data.  It seems that the latter technique should be allowed based on what William said, because:</p>\n\n<ul>\n<li>No one paid for imagery.</li>\n<li>The answers aren't given away in a direct manner.  The goal of the competition is to organize the images relative to each other and not to absolute dates on the calendar.  If only one image in the set is included, you would know little about where it fits in the set.</li>\n<li>It's within the spirit of the competition to be able to say both &quot;this photo was taken a few hours before the test photo&quot; and &quot;this test photo was taken the day after this other test photo.&quot;  Those are both temporal questions.</li>\n<li>There is a difference between the <em>approach</em> being open source and all the data utilized in the approach being open source.  We've already agreed that the data used can be external/copywritten, since Google Earth is allowed.</li>\n<li>It's difficult to know exactly what would help the host here.  The best practice approach to this problem (from my view) would be to analyze the <em>context</em> of the photographs well before the <em>content</em> of the photographs.  The date a photograph was taken is a context question, so outside information is the best way to answer it.  Using exclusively the contents of the image to answer the question would be a disservice to the host, since it would not be as consistently accurate.  If they had instead hosted a competition that asked &quot;how many vehicles are in this photograph?&quot; or &quot;what is the health of this forest?&quot; restricting outside information would make much more sense.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120817,
      "author_name": "jfulton",
      "author_url": "",
      "post_date": "05/20/2016 18:17:24",
      "content": "<p>All of the above mirrors my thinking -- I'm not particularly suggesting that one way or the other is more correct, but think it would be helpful if admin could address this specific issue with a clear yes/no on what is allowed/desirable for Draper. </p>\n\n<p>It is not clear to me either just what kind of solutions they are after, but it does seem as though the difficulty of the problem was overestimated. I find it particularly unlikely that &quot;align images with external data of known date&quot; is what they were thinking here, but the rules as stated so far do not seem to disallow it. :-)</p>\n\n<p>I am also enjoying the puzzle, and have been able to automate certain aspects, but more along the lines of an &quot;analyst's toolkit&quot; than machine learning. Curiously, I believe I can deduce specific dates for many of these particular images <em>without</em> using external imagery -- not the point of the contest as you say, but it is helpful to anchor time frames for the different sets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120845,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "05/20/2016 22:03:26",
      "content": "<p>One should not use external images if their only redeeming contribution to the methodology is their timestamp. Going back to the spirit of the competition (namely, finding repeatable image features that distinguish images in time), such usage doesn't meet the aims of the host. The difficulty of enforcing this really depends on how much overlapping imagery is out there. We (Kaggle) were not aware of the extent of this external imagery at the time when we launched the competition and have, so far, only received reports of the one day's worth referenced earlier in this thread. No matter what, Draper will be reviewing the winners' methodology and have an eye out for suspicious post facto explanations.</p>\n\n<p>@Tyler re: &quot;Using exclusively the contents of the image to answer the question would be a disservice to the host&quot; Features based on image content is actually the goal here. It's not about the date/order these specific images were acquired, but rather in finding generalizable features within an image that give away the order (and beyond just the order, features that make other tasks possible). External data was permitted in hopes that people could construct more training data to improve the results.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "120526": "For approximately 252 test sets, overlapping commercial imagery exists that identifies the exact date that at least one photo from the set was taken.  The external imagery was captured within a few hours of the Draper commissioned set.  \r\n\r\nExample:\r\n\r\n ![Example of overlap.][1]\r\n\r\nI know it has been clarified a few times that external imagery and reverse-engineering locations are allowed, but this seems a little different since it provides a flawless method for identifying which day many photos were taken.  It also makes other methods, like tide levels and solar noon shadow corroboration, much more straightforward to analyze.  It's still \"generalizable\" since it's probably impossible to capture aerial images from anywhere in the world that don't overlap with some other platform the same day.  \r\n\r\nWill final submissions be allowed if the output partially relied on correlating with such commercial imagery?\r\n\r\nI won't post the commercial source and the corresponding absolute date range the test photos were taken in unless the admins want those provided publicly.\r\n\r\n  [1]: http://i.imgur.com/T1VFPJz.png",
    "120540": "Doesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...",
    "120560": "[quote=kes367;120540]\r\n\r\nDoesn't seem like much of a competition if this is allowed...  Hypothetically, somebody could easily get 0.96 spearman in 1 submission if they did this...\r\n\r\n[/quote]\r\nNot exactly 0.96.  It's only one image in the set of five.  And it's also freely available commercial imagery.  \r\n\r\nThough I share your concern that it incentivizes a non-innovative race to the top.",
    "120562": "Right, but if someone were able to find this, it isn't a huge stretch that they could purchase commercial satellite imagery for that date range around that time for a given location.",
    "120646": "Thanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.",
    "120657": "[quote=William Cukierski;120646]\r\n\r\nThanks for bringing this to our attention. While I don't know the cost/use/license details of your source images, in general you should not use paid commercial imagery as part of your approach, nor images that give away the answer in a very direct manner. In addition to being against the spirit of the competition, the rules state that your approach should be open source, not trigger any royalties to third parties, and that you have the right to grant the license to what you use and produce. We understand the latter points can be tough to legally navigate, but the real concern is that such an approach does not help the host.\r\n\r\n[/quote]\r\nThanks for the clarification!",
    "120682": "I'm still not 100% clear on this.\r\n\r\nI haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is \"freely available\" (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?\r\n\r\nThe second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not _directly_ give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.\r\n\r\nSo my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are \"freely available\" at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for _every_ image, giving the answer up directly)\r\n\r\nI can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)",
    "120783": "[quote=Jordie Fulton;120682]\r\n\r\n\r\nI'm still not 100% clear on this.\r\n\r\nI haven't gone looking for the imagery in question as I agree that it is not in the spirit of the competition, but Tyler does say that this is \"freely available\" (I assume free as in beer), so I'm still not sure whether this type of data can be used or not?\r\n\r\nThe second part of William's response (that external images that give away the answer directly are not allowed) may or may not apply; certainly knowing the exact date of one image in a set does not _directly_ give away the order of the set, but looking at the data as a whole it makes handtagging or ML rather easier if an anchor exists -- given that the time-frames of the test sets are similar or the same, and the possibility of multiple external sources, this could amount to the same thing.\r\n\r\nSo my question I guess is: Assuming that dated images such as Tyler's example are available overlapping the test set, and that they are \"freely available\" at least in the sense of say Google Earth, is it acceptable to use dates derived in this way as a factor in automated or handtagging approaches? (assuming that you don't have exact matches for _every_ image, giving the answer up directly)\r\n\r\nI can't really see how this can be enforced in any case with a handtagging approach (this goes for whatever paid data sources may be out there also) as it should be very easy to construct a justification post facto from image content, given the correct order, in most cases; the poisoned information could then be removed from the workflow prior to submission. (not that anyone would do this of course)\r\n\r\n\r\n[/quote]\r\nThese are valid questions.  \r\n\r\nYou're right that when I say free I mean free as in beer not free as in speech.  I certainly didn't pay for them; I just sifted through a lot of different aerial photo sites until I found some that overlapped.  You will probably always be able to do this in the United States, as evidenced by the introductory paragraph to this challenge which cites that we should have an image of everywhere every day by 2017.  But it might not generalize to, say, Antarctica or the middle of Africa or parts of Russia.  \r\n\r\nHowever, it's difficult to stay within the spirit of the contest when it's very unclear what the spirit of the contest actually is.  There is no real world application for putting satellite images in chronological order.  Most analysts (full disclosure: I'm an imagery analyst, not a ML programmer) won't even bother to look at an image if it doesn't contain sensor data that allows automatic orthorectification.  It's just bad practice.  Imagine if in the \"Ultrasound Nerve Segmentation\" contest we were only allowed to look at Polaroid photos of the subjects.  No doctor in their right mind would use that to find nerves, and there isn't a use case for it because far better methods exist.  It's interesting as a puzzle of sorts (which is why I'm here) but it's not practical.\r\n\r\nI also agree that it would be trivial to hand tag with the external data and come up with an in-image justification.  Alternatively, you might hand tag them and check the answers against the external data.  It seems that the latter technique should be allowed based on what William said, because:\r\n\r\n - No one paid for imagery.\r\n - The answers aren't given away in a direct manner.  The goal of the competition is to organize the images relative to each other and not to absolute dates on the calendar.  If only one image in the set is included, you would know little about where it fits in the set.\r\n - It's within the spirit of the competition to be able to say both \"this photo was taken a few hours before the test photo\" and \"this test photo was taken the day after this other test photo.\"  Those are both temporal questions.\r\n - There is a difference between the *approach* being open source and all the data utilized in the approach being open source.  We've already agreed that the data used can be external/copywritten, since Google Earth is allowed.\r\n - It's difficult to know exactly what would help the host here.  The best practice approach to this problem (from my view) would be to analyze the *context* of the photographs well before the *content* of the photographs.  The date a photograph was taken is a context question, so outside information is the best way to answer it.  Using exclusively the contents of the image to answer the question would be a disservice to the host, since it would not be as consistently accurate.  If they had instead hosted a competition that asked \"how many vehicles are in this photograph?\" or \"what is the health of this forest?\" restricting outside information would make much more sense.",
    "120817": "All of the above mirrors my thinking -- I'm not particularly suggesting that one way or the other is more correct, but think it would be helpful if admin could address this specific issue with a clear yes/no on what is allowed/desirable for Draper. \r\n\r\nIt is not clear to me either just what kind of solutions they are after, but it does seem as though the difficulty of the problem was overestimated. I find it particularly unlikely that \"align images with external data of known date\" is what they were thinking here, but the rules as stated so far do not seem to disallow it. :-)\r\n\r\nI am also enjoying the puzzle, and have been able to automate certain aspects, but more along the lines of an \"analyst's toolkit\" than machine learning. Curiously, I believe I can deduce specific dates for many of these particular images _without_ using external imagery -- not the point of the contest as you say, but it is helpful to anchor time frames for the different sets.",
    "120845": "One should not use external images if their only redeeming contribution to the methodology is their timestamp. Going back to the spirit of the competition (namely, finding repeatable image features that distinguish images in time), such usage doesn't meet the aims of the host. The difficulty of enforcing this really depends on how much overlapping imagery is out there. We (Kaggle) were not aware of the extent of this external imagery at the time when we launched the competition and have, so far, only received reports of the one day's worth referenced earlier in this thread. No matter what, Draper will be reviewing the winners' methodology and have an eye out for suspicious post facto explanations.\r\n\r\n@Tyler re: \"Using exclusively the contents of the image to answer the question would be a disservice to the host\" Features based on image content is actually the goal here. It's not about the date/order these specific images were acquired, but rather in finding generalizable features within an image that give away the order (and beyond just the order, features that make other tasks possible). External data was permitted in hopes that people could construct more training data to improve the results."
  },
  "source": "meta"
}