{
  "id": 26506,
  "title": "Welcome! ",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/26506",
  "author_name": "Wendy Kan",
  "post_date": "2016-12-15T01:10:25.557000",
  "votes": 19,
  "comment_count": 83,
  "views": 4616,
  "content": "<p>On behalf of Kaggle and Dstl, I'd like to welcome you to Dstl Satellite Imagery Feature Detection competition. Satellite imagery and geometries are a new type of data for the Kaggle community, and we spent a good amount of time to make it work. We hope you enjoy the experience dealing with this type of data. </p>\n\n<p>We're happy to answer your questions here on the forum. Good luck and have fun!</p>",
  "messages": [
    {
      "id": 150394,
      "postDate": "2016-12-15T01:10:25.557Z",
      "content": "<p>On behalf of Kaggle and Dstl, I'd like to welcome you to Dstl Satellite Imagery Feature Detection competition. Satellite imagery and geometries are a new type of data for the Kaggle community, and we spent a good amount of time to make it work. We hope you enjoy the experience dealing with this type of data. </p>\n\n<p>We're happy to answer your questions here on the forum. Good luck and have fun!</p>",
      "rawMarkdown": "On behalf of Kaggle and Dstl, I'd like to welcome you to Dstl Satellite Imagery Feature Detection competition. Satellite imagery and geometries are a new type of data for the Kaggle community, and we spent a good amount of time to make it work. We hope you enjoy the experience dealing with this type of data. \r\n\r\nWe're happy to answer your questions here on the forum. Good luck and have fun!",
      "votes": 18
    },
    {
      "id": 159329,
      "postDate": "2017-02-01T20:47:27.213Z",
      "content": "<p>I can't submit file in new interface. The only thing I've got is: \"Your submission was unsuccessful.\"</p>",
      "rawMarkdown": "I can't submit file in new interface. The only thing I've got is: \"Your submission was unsuccessful.\"",
      "votes": 5,
      "replies": [
        {
          "id": 159359,
          "postDate": "2017-02-01T23:51:34.923Z",
          "content": "<p>same here. also, the counter of remaining submission is broken. on the \"Submit Predictions\" page, \"You have -3 submissions today.\"</p>",
          "rawMarkdown": "same here. also, the counter of remaining submission is broken. on the \"Submit Predictions\" page, \"You have -3 submissions today.\"",
          "votes": 3
        },
        {
          "id": 159360,
          "postDate": "2017-02-02T00:22:51.693Z",
          "content": "<p>@Wendy - couple of other bugs, other than submissions not getting through in this new interface:</p>\n\n<ol>\n<li><p>The daily quota counter decrements even though there is an error with submission (I just tried submitting today.  I suppose you can get negative submissions with this interface per response above.</p></li>\n<li><p>\"This leaderboard is calculated with approximately 1% of the test data.\nThe final results will be based on the other 99%, so the final standings may be different.\".  I suppose it is still 19% public?</p></li>\n</ol>\n\n<p>also, comparing the LB from half an hour back it seems some submissions went through, so it could either be a filesize or compression type (I am using .gz) change from the last interface?</p>",
          "rawMarkdown": "@Wendy - couple of other bugs, other than submissions not getting through in this new interface:\n\n1. The daily quota counter decrements even though there is an error with submission (I just tried submitting today.  I suppose you can get negative submissions with this interface per response above.\n\n2. \"This leaderboard is calculated with approximately 1% of the test data.\nThe final results will be based on the other 99%, so the final standings may be different.\".  I suppose it is still 19% public?\n\nalso, comparing the LB from half an hour back it seems some submissions went through, so it could either be a filesize or compression type (I am using .gz) change from the last interface?",
          "votes": 2
        },
        {
          "id": 159419,
          "postDate": "2017-02-02T07:01:49.193Z",
          "content": "<p>Some more problems: </p>\n\n<p>1) I often save useful data in filename. In new interface filename is not visible.</p>\n\n<p>2) I don't see how I can download my old submission.</p>",
          "rawMarkdown": "Some more problems: \n\n1) I often save useful data in filename. In new interface filename is not visible.\n\n2) I don't see how I can download my old submission.\n",
          "votes": 4
        },
        {
          "id": 159620,
          "postDate": "2017-02-03T02:59:55.503Z",
          "content": "<p>@ZFTurbo we're definitely fixing 1) and 2). Are you still having submission issues to this competition? I didn't see a failing submission from you.</p>",
          "rawMarkdown": "@ZFTurbo we're definitely fixing 1) and 2). Are you still having submission issues to this competition? I didn't see a failing submission from you."
        },
        {
          "id": 159651,
          "postDate": "2017-02-03T05:42:05.890Z",
          "content": "<p><strong>@Ben Hamner</strong>\nI still have a problem. I uploaded ZIP file with size around 200 MB. After it reached 100% I pressed \"Make submission\" it redirect me to Leaderboard without any notifications. When I go to \"My submissions\" page I don't see my latest submission. If I try to reupload the same ZIP-file Progressbar reach 100% instantly.</p>\n\n<p>I also see text errors on \"My submissions\" page:</p>\n\n<p>1) 90 submissions for ZFTurbo - As I remember I have less than 90 submissions in total.</p>\n\n<p>2) $100,000 · 0 teams · a month to go (a month to go until merger deadline)</p>",
          "rawMarkdown": "**@Ben Hamner**\nI still have a problem. I uploaded ZIP file with size around 200 MB. After it reached 100% I pressed \"Make submission\" it redirect me to Leaderboard without any notifications. When I go to \"My submissions\" page I don't see my latest submission. If I try to reupload the same ZIP-file Progressbar reach 100% instantly.\n\nI also see text errors on \"My submissions\" page:\n\n1) 90 submissions for ZFTurbo - As I remember I have less than 90 submissions in total.\n\n2) $100,000 · 0 teams · a month to go (a month to go until merger deadline)\n"
        }
      ]
    },
    {
      "id": 151324,
      "postDate": "2016-12-20T04:19:58.307Z",
      "content": "<p>woohoo, my first contribution to the Kaggle github repo :)</p>\n\n<p><a href=\"https://github.com/Kaggle/docker-python/pull/48\">https://github.com/Kaggle/docker-python/pull/48</a></p>",
      "rawMarkdown": "woohoo, my first contribution to the Kaggle github repo :)\r\n\r\nhttps://github.com/Kaggle/docker-python/pull/48",
      "votes": 6
    },
    {
      "id": 158643,
      "postDate": "2017-01-28T18:19:41.287Z",
      "content": "<p>@Wendy, from a business perspective, it looks that set up for this competition has some flaws.</p>\n\n<p>Only quarter of the participants are above the sample benchmark.... and it is not because problem is that hard, I mean it is hard to get a good result, but to get above benchmark is straightforward, any FCN will get your there.</p>\n\n<p>I would believe that DSTL wants to get a good solution, Kaggle wants to ensure that DSTL gets a good solution or at least decent interest from the Data Science community. </p>\n\n<p>I agree that Kaggle community is lazy and relaxed in a way that people do not want to invest their time on the engineering issues. But still, it is just easier to give up on this problem and switch to some other challenge.</p>\n\n<p>Boring Allstate competition attracted 3000+ participants, and this exciting problem has only 50 above benchmark and 150 below it. I do not know who is project owner for this problem, but it may be time to rethink a way that submission is performed to have satisfied client in 38 days.</p>\n\n<p>I really like an idea to do submission as polygons, as the client requires, transform it to mask and do the evaluation on the mask level.</p>\n\n<p>Or to provide a kernel with a function that will do mapping: mask =&gt; polygons in such a way that submission that is created with this function will not raise errors. </p>",
      "rawMarkdown": "@Wendy, from a business perspective, it looks that set up for this competition has some flaws.\r\n\r\nOnly quarter of the participants are above the sample benchmark.... and it is not because problem is that hard, I mean it is hard to get a good result, but to get above benchmark is straightforward, any FCN will get your there.\r\n\r\nI would believe that DSTL wants to get a good solution, Kaggle wants to ensure that DSTL gets a good solution or at least decent interest from the Data Science community. \r\n\r\nI agree that Kaggle community is lazy and relaxed in a way that people do not want to invest their time on the engineering issues. But still, it is just easier to give up on this problem and switch to some other challenge.\r\n\r\nBoring Allstate competition attracted 3000+ participants, and this exciting problem has only 50 above benchmark and 150 below it. I do not know who is project owner for this problem, but it may be time to rethink a way that submission is performed to have satisfied client in 38 days.\r\n\r\nI really like an idea to do submission as polygons, as the client requires, transform it to mask and do the evaluation on the mask level.\r\n\r\nOr to provide a kernel with a function that will do mapping: mask => polygons in such a way that submission that is created with this function will not raise errors. ",
      "votes": 4,
      "replies": [
        {
          "id": 159107,
          "postDate": "2017-01-31T20:30:12.843Z",
          "content": "<p>Hi Vladimir, </p>\n\n<p>Thanks for the suggestion. In fact, your suggestion of submission as polygons -&gt; transform into mask -&gt; evaluation was the original idea. However, after some experiments, I found that the performance is quite poor. The bottle neck was mainly in the generating of mask from polygons. When the geometry gets more complex, the repeated calls to polygon.contains(pixel) took a very long time since it had to be repeated for every pixel and was about 50-100x slower than calculating Jaccard directly on the polygons. Therefore, we re-wrote the metric to be purely vector-based. </p>\n\n<p>The metric execution time of this competition has been a challenge from the very beginning and we went through quite some iterations on designing this competition. Very different versions of metric and different implementations were experimented. We understand it's not going to be as traditional or trivial as some other Kaggle competitions, therefore not going to be as popular. But we do want to design these non-conventional competitions to expose our community to a completely different problem space. </p>",
          "rawMarkdown": "Hi Vladimir, \n\nThanks for the suggestion. In fact, your suggestion of submission as polygons -> transform into mask -> evaluation was the original idea. However, after some experiments, I found that the performance is quite poor. The bottle neck was mainly in the generating of mask from polygons. When the geometry gets more complex, the repeated calls to polygon.contains(pixel) took a very long time since it had to be repeated for every pixel and was about 50-100x slower than calculating Jaccard directly on the polygons. Therefore, we re-wrote the metric to be purely vector-based. \n\nThe metric execution time of this competition has been a challenge from the very beginning and we went through quite some iterations on designing this competition. Very different versions of metric and different implementations were experimented. We understand it's not going to be as traditional or trivial as some other Kaggle competitions, therefore not going to be as popular. But we do want to design these non-conventional competitions to expose our community to a completely different problem space. ",
          "votes": 6
        }
      ]
    },
    {
      "id": 160163,
      "postDate": "2017-02-06T12:02:35.513Z",
      "content": "<p>@wendy Kan Why can we not see the error message that we can use to get before. Debugging with the new interface without error message is very tough.</p>",
      "rawMarkdown": "@wendy Kan Why can we not see the error message that we can use to get before. Debugging with the new interface without error message is very tough.",
      "votes": 3,
      "replies": [
        {
          "id": 160256,
          "postDate": "2017-02-06T20:10:37.763Z",
          "content": "<p>@Attila,</p>\n\n<p>Very sorry about this - we are trying to fix this bug. In the mean time,  I'm going to email you with your submission errors as a temporary fix. </p>",
          "rawMarkdown": "@Attila,\n\nVery sorry about this - we are trying to fix this bug. In the mean time,  I'm going to email you with your submission errors as a temporary fix. ",
          "votes": 2
        },
        {
          "id": 160784,
          "postDate": "2017-02-09T13:46:51.277Z",
          "content": "<p>Hi Wendy,</p>\n\n<p>Not having the error message is a problem. I have a file that passes the tpex test tool (\"Good to go!\")  but is on error when submiting.\nI also have one submission that is still being processed since nearly 15 hours. Is this blocking further submissions ?</p>",
          "rawMarkdown": "Hi Wendy,\n\nNot having the error message is a problem. I have a file that passes the tpex test tool (\"Good to go!\")  but is on error when submiting.\nI also have one submission that is still being processed since nearly 15 hours. Is this blocking further submissions ?\n"
        },
        {
          "id": 160794,
          "postDate": "2017-02-09T14:38:11.733Z",
          "content": "<p>@WendyKan Thanks a lot for sending the error for each submission. </p>",
          "rawMarkdown": "@WendyKan Thanks a lot for sending the error for each submission. \n "
        },
        {
          "id": 161306,
          "postDate": "2017-02-13T04:53:21.533Z",
          "content": "<p>@Wendy, do you have an estimate about when this error is going to be fixed. It's really difficult to debug without any knowledge of what's going on during the evaluation.</p>",
          "rawMarkdown": "@Wendy, do you have an estimate about when this error is going to be fixed. It's really difficult to debug without any knowledge of what's going on during the evaluation."
        }
      ]
    },
    {
      "id": 159335,
      "postDate": "2017-02-01T21:34:03.027Z",
      "content": "<p>@Wendy,\nIs it possible to see an error message with the new interface? It is much easier to debug a code if you know the type of exception occurred.  It is especially important in this type of competition, where almost every participant experience submission problems. <br>\nUpdate: I got \"Evaluation Exception: Submission must have 4290 rows\" for one of my submissions. But submission file must have 1 header line and 4290 lines for each image =&gt; 4291 in total. What should I do to fix that problem? \nThanks in advance!</p>",
      "rawMarkdown": "@Wendy,\nIs it possible to see an error message with the new interface? It is much easier to debug a code if you know the type of exception occurred.  It is especially important in this type of competition, where almost every participant experience submission problems.  \nUpdate: I got \"Evaluation Exception: Submission must have 4290 rows\" for one of my submissions. But submission file must have 1 header line and 4290 lines for each image => 4291 in total. What should I do to fix that problem? \nThanks in advance!",
      "votes": 3
    },
    {
      "id": 158357,
      "postDate": "2017-01-27T02:05:26.383Z",
      "content": "<p>There's a much more probable explanation for timeouts.</p>",
      "rawMarkdown": "There's a much more probable explanation for timeouts.",
      "votes": 3
    },
    {
      "id": 157399,
      "postDate": "2017-01-20T19:52:34.690Z",
      "content": "<p>@ironbar,</p>\n\n<p>Yes, I tested the perfect submission prior to launching this competition and it finished evaluation under 6 mins, which is within the 8 minute time limit. </p>\n\n<p>I'm working with the engineering team to see if we can lift this limit for this competition. </p>",
      "rawMarkdown": "@ironbar,\n\nYes, I tested the perfect submission prior to launching this competition and it finished evaluation under 6 mins, which is within the 8 minute time limit. \n\nI'm working with the engineering team to see if we can lift this limit for this competition. ",
      "votes": 3,
      "replies": [
        {
          "id": 157463,
          "postDate": "2017-01-21T10:25:43.550Z",
          "content": "<p>@ Wendy, <br>\nHi thanks for your answer, <br>\nAs people is saying here, one alternative for computing the score is projecting the polygons to a raster and computing the score pixel-wise. <br>\n<a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/27673/non-noded-intersection-inconsistent-polygon-validity-results\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/27673/non-noded-intersection-inconsistent-polygon-validity-results</a></p>\n\n<p>I have tested already on the train set and it's much faster than making the polygon intersection, and the accuracy is very similar when using a raster of 4000x4000 or 8000x8000.  </p>\n\n<p>The speedup can be more than x10. <br>\nAlso that would eliminate the problem of Topology exception.</p>\n\n<p>I hope you and your team can come up to a good solution, because right now I have to downsize the submission from 96MB to 9MB to avoid the timeout error. And that hurts the score a lot. </p>",
          "rawMarkdown": "@ Wendy,  \r\nHi thanks for your answer,   \r\nAs people is saying here, one alternative for computing the score is projecting the polygons to a raster and computing the score pixel-wise.   \r\nhttps://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/27673/non-noded-intersection-inconsistent-polygon-validity-results\r\n\r\nI have tested already on the train set and it's much faster than making the polygon intersection, and the accuracy is very similar when using a raster of 4000x4000 or 8000x8000.  \r\n\r\nThe speedup can be more than x10.  \r\nAlso that would eliminate the problem of Topology exception.\r\n\r\nI hope you and your team can come up to a good solution, because right now I have to downsize the submission from 96MB to 9MB to avoid the timeout error. And that hurts the score a lot. \r\n\r\n",
          "votes": 3
        },
        {
          "id": 158156,
          "postDate": "2017-01-25T19:45:43.253Z",
          "content": "<p>update: We've upped the evaluation time limit to 10 mins (instead of 8) for this competition. </p>\n\n<p>I understand that pixel-based evaluation will be faster than geometry operations. However, we can't use a pixel based evaluation as it is part of the requirement from Dstl that the output of the algorithm has to be polygons. </p>",
          "rawMarkdown": "update: We've upped the evaluation time limit to 10 mins (instead of 8) for this competition. \n\nI understand that pixel-based evaluation will be faster than geometry operations. However, we can't use a pixel based evaluation as it is part of the requirement from Dstl that the output of the algorithm has to be polygons. ",
          "votes": 1
        },
        {
          "id": 158331,
          "postDate": "2017-01-26T22:10:26.367Z",
          "content": "<p>@Wendy</p>\n\n<blockquote>\n  <p>update: We've upped the evaluation time limit to 10 mins (instead of 8)</p>\n</blockquote>\n\n<p>Would that be real time or user time? Is it strongly affected by how busy your server is?</p>",
          "rawMarkdown": "@Wendy\n\n> update: We've upped the evaluation time limit to 10 mins (instead of 8)\n\nWould that be real time or user time? Is it strongly affected by how busy your server is?"
        },
        {
          "id": 158343,
          "postDate": "2017-01-26T23:18:22.410Z",
          "content": "<p>It's not affected by the busy-ness of server. The worker node starts its clock after we received your complete submission file (so your upload speed also doesn't matter). </p>",
          "rawMarkdown": "It's not affected by the busy-ness of server. The worker node starts its clock after we received your complete submission file (so your upload speed also doesn't matter). ",
          "votes": 1,
          "replies": [
            {
              "id": 158355,
              "postDate": "2017-01-27T01:02:46.217Z",
              "content": "<p>[quote=Wendy Kan;158343]</p>\n\n<p>It's not affected by the busy-ness of server. The worker node starts its clock after we received your complete submission file (so your upload speed also doesn't matter). </p>\n\n<p>[/quote]</p>\n\n<p>Wendy,</p>\n\n<p>I meant \"real time\" and \"user time\" in the Unix sense:</p>\n\n<p><a href=\"http://unix.stackexchange.com/questions/53302/why-would-the-real-time-be-much-higher-than-the-user-and-sys-times-combine\">http://unix.stackexchange.com/questions/53302/why-would-the-real-time-be-much-higher-than-the-user-and-sys-times-combine</a></p>\n\n<p>The former is what most people refer to as the \"clock\", but it's also very strongly affected by whether your worker node is processing other submissions simultaneously.</p>\n\n<p>That might explain why the users are timing out.</p>",
              "rawMarkdown": "[quote=Wendy Kan;158343]\r\n\r\nIt's not affected by the busy-ness of server. The worker node starts its clock after we received your complete submission file (so your upload speed also doesn't matter). \r\n\r\n[/quote]\r\n\r\nWendy,\r\n\r\nI meant \"real time\" and \"user time\" in the Unix sense:\r\n\r\nhttp://unix.stackexchange.com/questions/53302/why-would-the-real-time-be-much-higher-than-the-user-and-sys-times-combine\r\n\r\nThe former is what most people refer to as the \"clock\", but it's also very strongly affected by whether your worker node is processing other submissions simultaneously.\r\n\r\nThat might explain why the users are timing out.",
              "votes": -2
            }
          ]
        },
        {
          "id": 158514,
          "postDate": "2017-01-27T19:56:30.070Z",
          "content": "<p>It's real/clock time. However, the prediction competition scoring workers only do one request at a time, so they shouldn't be affected by noisy neighbors. </p>",
          "rawMarkdown": "It's real/clock time. However, the prediction competition scoring workers only do one request at a time, so they shouldn't be affected by noisy neighbors. "
        }
      ]
    },
    {
      "id": 157219,
      "postDate": "2017-01-19T18:28:14.903Z",
      "content": "<p>Hi Wendy, <br>\nI'm concerned about submission timeout. <br>\n¿Have you tried to submit the perfect submission?  ¿Does it have enough time to compute the score?</p>\n\n<p>I'm asking because from my experience:</p>\n\n<ul>\n<li>We are given polygons for 25 images for training</li>\n<li>The file of this polygons uses 50.3MB uncompressed (the precision of the points is 6)</li>\n<li>When we have to make the submission we are predicting for 429 images, so the file should be approximately 20 times bigger -&gt; 1000MB uncompressed</li>\n<li>In my case I can't submit files bigger than 36MB when using precision 6 (like in the training data). To get down to that size I have to simplify the prediction a lot and it hurts the score.</li>\n<li>If the files are bigger I get timeout error.</li>\n</ul>\n\n<p>So there's a big difference between the estimated size that a perfect submission will have and the size that until now I have been able to submit.</p>\n\n<p>Thanks\nIronbar</p>",
      "rawMarkdown": "Hi Wendy,   \r\nI'm concerned about submission timeout.   \r\n¿Have you tried to submit the perfect submission?  ¿Does it have enough time to compute the score?\r\n\r\nI'm asking because from my experience:\r\n\r\n* We are given polygons for 25 images for training\r\n* The file of this polygons uses 50.3MB uncompressed (the precision of the points is 6)\r\n* When we have to make the submission we are predicting for 429 images, so the file should be approximately 20 times bigger -> 1000MB uncompressed\r\n* In my case I can't submit files bigger than 36MB when using precision 6 (like in the training data). To get down to that size I have to simplify the prediction a lot and it hurts the score.\r\n* If the files are bigger I get timeout error.\r\n\r\nSo there's a big difference between the estimated size that a perfect submission will have and the size that until now I have been able to submit.\r\n\r\n\r\nThanks\r\nIronbar",
      "votes": 3
    },
    {
      "id": 150763,
      "postDate": "2016-12-16T16:17:44.447Z",
      "content": "<p>And to access images:</p>\n\n<pre><code>library(raster)\nraster_6044_4_4 &lt;- raster(\"./data/three_band/6040_4_4.tif\")\nplot(raster_6044_4_4)\n\nlibrary(rgdal)\ngdal_6044_4_4 &lt;- readGDAL(paste0(\"./data/three_band/\", '6040_4_4', \".tif\"))\nplot(gdal_6044_4_4)\n</code></pre>",
      "rawMarkdown": "And to access images:\r\n\r\n    library(raster)\r\n    raster_6044_4_4 <- raster(\"./data/three_band/6040_4_4.tif\")\r\n    plot(raster_6044_4_4)\r\n    \r\n    library(rgdal)\r\n    gdal_6044_4_4 <- readGDAL(paste0(\"./data/three_band/\", '6040_4_4', \".tif\"))\r\n    plot(gdal_6044_4_4)",
      "votes": 3
    },
    {
      "id": 150787,
      "postDate": "2016-12-16T18:21:17.540Z",
      "content": "<p>@ZFTurbo:</p>\n\n<p>Apologies! Our random train/test split picked something that has all empty multipolygons for class 7. They do exist in the test dataset. We will release another image that has class 7 shortly. </p>\n\n<p>EDIT: 6070_2_3 has been moved to train_wkt_v2.csv. It has non-empty multipolygon for class 7. Submission file remains the same while this image is ignored for scoring. </p>",
      "rawMarkdown": "@ZFTurbo:\r\n\r\nApologies! Our random train/test split picked something that has all empty multipolygons for class 7. They do exist in the test dataset. We will release another image that has class 7 shortly. \r\n\r\nEDIT: 6070_2_3 has been moved to train_wkt_v2.csv. It has non-empty multipolygon for class 7. Submission file remains the same while this image is ignored for scoring. ",
      "votes": 4
    },
    {
      "id": 150574,
      "postDate": "2016-12-15T20:42:39.053Z",
      "content": "<p>ClassType = 7 absent in train_wkt.csv. Does it exist in test?</p>",
      "rawMarkdown": "ClassType = 7 absent in train_wkt.csv. Does it exist in test?",
      "votes": 4
    },
    {
      "id": 150481,
      "postDate": "2016-12-15T11:41:32.023Z",
      "content": "<p>Can I ask what the Jaccard Index value will be if the predicted and actual MultipolygonWKT are both \"MULTIPOLYGON EMPTY\"? Thanks.</p>",
      "rawMarkdown": "Can I ask what the Jaccard Index value will be if the predicted and actual MultipolygonWKT are both \"MULTIPOLYGON EMPTY\"? Thanks.",
      "votes": 2,
      "replies": [
        {
          "id": 150587,
          "postDate": "2016-12-15T22:00:24.610Z",
          "content": "<p>@Yunfeng, </p>\n\n<p>The Jaccard index is not per-image, it's calculating a total TP/FP/FN over a bunch of images for a given class. So if they are both empty, then TP=FP=FN=0, so totalTP/totalFP/totalFN don't change for this image. </p>\n\n<p>[quote=Yunfeng Zhu;150481]</p>\n\n<p>Can I ask what the Jaccard Index value will be if the predicted and actual MultipolygonWKT are both \"MULTIPOLYGON EMPTY\"? Thanks.</p>\n\n<p>[/quote]</p>",
          "rawMarkdown": "@Yunfeng, \r\n\r\nThe Jaccard index is not per-image, it's calculating a total TP/FP/FN over a bunch of images for a given class. So if they are both empty, then TP=FP=FN=0, so totalTP/totalFP/totalFN don't change for this image. \r\n\r\n[quote=Yunfeng Zhu;150481]\r\n\r\nCan I ask what the Jaccard Index value will be if the predicted and actual MultipolygonWKT are both \"MULTIPOLYGON EMPTY\"? Thanks.\r\n\r\n[/quote]\r\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 166588,
      "postDate": "2017-03-10T08:14:59.950Z",
      "content": "<p>Hi @wendykan,\nnow that the competition is over are we allowed to submit ? There are ideas I haven't finish to explore and would like to check how well it performs against test dataset ? If yes, how long will it be open ?</p>",
      "rawMarkdown": "Hi @wendykan,\nnow that the competition is over are we allowed to submit ? There are ideas I haven't finish to explore and would like to check how well it performs against test dataset ? If yes, how long will it be open ?",
      "votes": 1,
      "replies": [
        {
          "id": 167462,
          "postDate": "2017-03-14T06:27:41.987Z",
          "content": "<p>Hi @Wendy,\nHaving no reply on the forum I sent you a direct message. Did you receive it ?</p>",
          "rawMarkdown": "Hi @Wendy,\nHaving no reply on the forum I sent you a direct message. Did you receive it ?"
        },
        {
          "id": 167600,
          "postDate": "2017-03-14T17:43:52.313Z",
          "content": "<p>I have been submitting to check ideas all week. Usually we have unlimited submissions once the competition is over,  but it looks like right now we only have 3 per day.  It would be nice if this limit was lifted. </p>",
          "rawMarkdown": "I have been submitting to check ideas all week. Usually we have unlimited submissions once the competition is over,  but it looks like right now we only have 3 per day.  It would be nice if this limit was lifted. "
        }
      ]
    },
    {
      "id": 158197,
      "postDate": "2017-01-26T02:47:14.503Z",
      "content": "<p>@Wendy </p>\n\n<p>Couldn't you round numbers in your wkt to remove too precise numbers just in case? Basically, the \"Non-Noded Intersection\" errors are unavoidable as long as intersections are calculated because of float number precision. But, I feel numbers like -0.007879000000000001 among rounded numbers are more likely to cause this problem. </p>\n\n<pre><code>6040_2_2,4,\"MULTIPOLYGON (((0.003025 -0.007879000000000001, 0.003074 -0.007931000000000001, 0.003123 -0.007996, 0.003182 \n</code></pre>",
      "rawMarkdown": "@Wendy \r\n\r\nCouldn't you round numbers in your wkt to remove too precise numbers just in case? Basically, the \"Non-Noded Intersection\" errors are unavoidable as long as intersections are calculated because of float number precision. But, I feel numbers like -0.007879000000000001 among rounded numbers are more likely to cause this problem. \r\n\r\n    6040_2_2,4,\"MULTIPOLYGON (((0.003025 -0.007879000000000001, 0.003074 -0.007931000000000001, 0.003123 -0.007996, 0.003182 \r\n\r\n",
      "votes": 1,
      "replies": [
        {
          "id": 158348,
          "postDate": "2017-01-26T23:27:41.703Z",
          "content": "<p>Interesting idea. Let me look into this. </p>",
          "rawMarkdown": "Interesting idea. Let me look into this. ",
          "votes": 1
        },
        {
          "id": 158548,
          "postDate": "2017-01-28T00:28:43.937Z",
          "content": "<p>Some updates to this: I reduced the rounding_precision in the geometry and did some testing, and it resulted in <strong>more</strong> non-noded intersection errors, even in the sample_submission file. </p>\n\n<p>I'll test more, but so far it doesn't seem like a good direction to go. </p>",
          "rawMarkdown": "Some updates to this: I reduced the rounding_precision in the geometry and did some testing, and it resulted in **more** non-noded intersection errors, even in the sample_submission file. \n\nI'll test more, but so far it doesn't seem like a good direction to go. ",
          "votes": 2,
          "replies": [
            {
              "id": 158561,
              "postDate": "2017-01-28T02:21:35.887Z",
              "content": "<p>[quote=Wendy Kan;158548]</p>\n\n<p>Some updates to this: I reduced the rounding_precision in the geometry and did some testing, and it resulted in <strong>more</strong> non-noded intersection errors, even in the sample_submission file. </p>\n\n<p>I'll test more, but so far it doesn't seem like a good direction to go. </p>\n\n<p>[/quote]</p>\n\n<p>I'm no expert, but I guess getting errors for sample_submission indicates geometries are broken. They were broken possibly because two non-intersected lines like (-1, 0) -&gt; (0, 0.99999999999) -&gt; (1, 0) and (1, 1) -&gt; (0, 1) -&gt; (1, 1) get touched to each other when rounded. I'm sorry, simply rounding numbers looks a wrong solution. Thank you, anyway.</p>",
              "rawMarkdown": "[quote=Wendy Kan;158548]\r\n\r\nSome updates to this: I reduced the rounding_precision in the geometry and did some testing, and it resulted in **more** non-noded intersection errors, even in the sample_submission file. \r\n\r\nI'll test more, but so far it doesn't seem like a good direction to go. \r\n\r\n[/quote]\r\n\r\nI'm no expert, but I guess getting errors for sample_submission indicates geometries are broken. They were broken possibly because two non-intersected lines like (-1, 0) -> (0, 0.99999999999) -> (1, 0) and (1, 1) -> (0, 1) -> (1, 1) get touched to each other when rounded. I'm sorry, simply rounding numbers looks a wrong solution. Thank you, anyway."
            }
          ]
        }
      ]
    },
    {
      "id": 157681,
      "postDate": "2017-01-23T03:31:29.410Z",
      "content": "<p>@Wendy, could you please clarify?</p>\n\n<p>Currently when we do submission is score calculated only for Public LB (19% of the test set images)\nor for the whole train set?</p>",
      "rawMarkdown": "@Wendy, could you please clarify?\r\n\r\nCurrently when we do submission is score calculated only for Public LB (19% of the test set images)\r\nor for the whole train set?",
      "votes": 1,
      "replies": [
        {
          "id": 158154,
          "postDate": "2017-01-25T19:37:30.353Z",
          "content": "<p>Your public LB score is based on the public split of the test dataset (19%). At the end of the competition you'll see your private LB scores that are based on the rest of the test set. </p>",
          "rawMarkdown": "Your public LB score is based on the public split of the test dataset (19%). At the end of the competition you'll see your private LB scores that are based on the rest of the test set. ",
          "votes": 1
        },
        {
          "id": 159464,
          "postDate": "2017-02-02T11:38:28.970Z",
          "content": "<p>@Wendy Kan, Do you know something about split between public and private? Is it estimated to be:\n19-81% (as has been given from the beginning of the competition) or 1-99% (as has been written in the current statement on leaderboard page)?</p>",
          "rawMarkdown": "@Wendy Kan, Do you know something about split between public and private? Is it estimated to be:\n19-81% (as has been given from the beginning of the competition) or 1-99% (as has been written in the current statement on leaderboard page)?"
        }
      ]
    },
    {
      "id": 153797,
      "postDate": "2017-01-03T12:19:23.187Z",
      "content": "<p>@Wendy Kan Hi, Kan. I have some problems in the submission. In the train_wkt_v4.csv, the results are listed with Multipolygon format, while in the sample_submission.csv, the results are listed with Polygon format, I was wondering, in the submission file, which format should we take? When I tried to submit the result with Multipolygon, I got some errors as following figure shows. Thank you very much! </p>",
      "rawMarkdown": "@Wendy Kan Hi, Kan. I have some problems in the submission. In the train_wkt_v4.csv, the results are listed with Multipolygon format, while in the sample_submission.csv, the results are listed with Polygon format, I was wondering, in the submission file, which format should we take? When I tried to submit the result with Multipolygon, I got some errors as following figure shows. Thank you very much! ",
      "votes": 1,
      "replies": [
        {
          "id": 153893,
          "postDate": "2017-01-03T23:06:06.067Z",
          "content": "<p>Hi @guangliang2016, both Polygon and MultiPolygon formats should work as long as they are valid. Seems like you were able to make some submissions without errors after this. </p>",
          "rawMarkdown": "Hi @guangliang2016, both Polygon and MultiPolygon formats should work as long as they are valid. Seems like you were able to make some submissions without errors after this. "
        }
      ]
    },
    {
      "id": 151261,
      "postDate": "2016-12-19T19:21:02.080Z",
      "content": "<p>@Gabriel Altay,</p>\n\n<p>Thanks for the suggestion! You're welcome to submit a pull request at <a href=\"https://github.com/Kaggle/docker-python\">https://github.com/Kaggle/docker-python</a> </p>",
      "rawMarkdown": "@Gabriel Altay,\r\n\r\nThanks for the suggestion! You're welcome to submit a pull request at https://github.com/Kaggle/docker-python \r\n\r\n",
      "votes": 1
    },
    {
      "id": 151238,
      "postDate": "2016-12-19T16:20:48.527Z",
      "content": "<p>@WendyKan would it be possible to make the following python packages available in the Kaggle kernel? </p>\n\n<ul>\n<li>tifffile (for reading multi-band tiff files) <a href=\"https://pypi.python.org/pypi/tifffile\">https://pypi.python.org/pypi/tifffile</a></li>\n<li>descartes (for dealing with polygons) <a href=\"https://pypi.python.org/pypi/descartes/1.0.2\">https://pypi.python.org/pypi/descartes/1.0.2</a></li>\n</ul>",
      "rawMarkdown": "@WendyKan would it be possible to make the following python packages available in the Kaggle kernel? \r\n\r\n - tifffile (for reading multi-band tiff files) https://pypi.python.org/pypi/tifffile\r\n - descartes (for dealing with polygons) https://pypi.python.org/pypi/descartes/1.0.2",
      "votes": 1
    },
    {
      "id": 150933,
      "postDate": "2016-12-17T17:45:49.353Z",
      "content": "<p>@shawn, Sorry about the discrepancy, something broke during the process of adding that one image to the training set. Please use train_wkt_v3.csv now. The train_geojson.zip will be updated on Monday. </p>",
      "rawMarkdown": "@shawn, Sorry about the discrepancy, something broke during the process of adding that one image to the training set. Please use train_wkt_v3.csv now. The train_geojson.zip will be updated on Monday. ",
      "votes": 1
    },
    {
      "id": 150503,
      "postDate": "2016-12-15T13:14:50.647Z",
      "content": "<p>Try this:</p>\n\n<p><a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial</a></p>",
      "rawMarkdown": "Try this:\r\n\r\nhttps://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial",
      "votes": 1
    },
    {
      "id": 152703,
      "postDate": "2016-12-27T23:52:13.927Z",
      "content": "<p>@RickTurner, I think your best bet would be to use a download manager that supports \"resume\" \n<a href=\"http://www.ghacks.net/2014/08/03/download-large-files/\">http://www.ghacks.net/2014/08/03/download-large-files/</a></p>",
      "rawMarkdown": "@RickTurner, I think your best bet would be to use a download manager that supports \"resume\" \r\nhttp://www.ghacks.net/2014/08/03/download-large-files/",
      "votes": 2
    },
    {
      "id": 150857,
      "postDate": "2016-12-17T03:10:32.263Z",
      "content": "<p>I just looked at the different train v2 formats and they do not have the same image id's in each.</p>\n\n<pre><code>diff -y in_wkt_csv in_geojson \n6010_4_2                              | 6010_1_2\n6010_4_4                            6010_4_4\n6040_1_0                            6040_1_0\n6040_1_3                            6040_1_3\n6040_2_2                            6040_2_2\n                                  &gt; 6040_4_4\n6060_2_3                            6060_2_3\n6070_2_3                            6070_2_3\n6090_2_0                            6090_2_0\n6100_1_3                            6100_1_3\n                                  &gt; 6100_2_2\n6100_2_3                            6100_2_3\n6110_1_2                            6110_1_2\n6110_3_1                            6110_3_1\n6110_4_0                              &lt;\n6120_2_0                            6120_2_0\n6120_2_2                            6120_2_2\n6140_1_2                            6140_1_2\n6140_3_1                              &lt;\n6150_2_3                              &lt;\n6160_2_1                            6160_2_1\n6170_0_4                            6170_0_4\n6170_2_4                            6170_2_4\n6170_4_1                            6170_4_1\n</code></pre>",
      "rawMarkdown": "I just looked at the different train v2 formats and they do not have the same image id's in each.\r\n\r\n   \r\n    diff -y in_wkt_csv in_geojson \r\n    6010_4_2\t\t\t\t\t\t      |\t6010_1_2\r\n    6010_4_4\t\t\t\t\t\t\t6010_4_4\r\n    6040_1_0\t\t\t\t\t\t\t6040_1_0\r\n    6040_1_3\t\t\t\t\t\t\t6040_1_3\r\n    6040_2_2\t\t\t\t\t\t\t6040_2_2\r\n    \t\t\t\t\t\t\t      >\t6040_4_4\r\n    6060_2_3\t\t\t\t\t\t\t6060_2_3\r\n    6070_2_3\t\t\t\t\t\t\t6070_2_3\r\n    6090_2_0\t\t\t\t\t\t\t6090_2_0\r\n    6100_1_3\t\t\t\t\t\t\t6100_1_3\r\n    \t\t\t\t\t\t\t      >\t6100_2_2\r\n    6100_2_3\t\t\t\t\t\t\t6100_2_3\r\n    6110_1_2\t\t\t\t\t\t\t6110_1_2\r\n    6110_3_1\t\t\t\t\t\t\t6110_3_1\r\n    6110_4_0\t\t\t\t\t\t      <\r\n    6120_2_0\t\t\t\t\t\t\t6120_2_0\r\n    6120_2_2\t\t\t\t\t\t\t6120_2_2\r\n    6140_1_2\t\t\t\t\t\t\t6140_1_2\r\n    6140_3_1\t\t\t\t\t\t      <\r\n    6150_2_3\t\t\t\t\t\t      <\r\n    6160_2_1\t\t\t\t\t\t\t6160_2_1\r\n    6170_0_4\t\t\t\t\t\t\t6170_0_4\r\n    6170_2_4\t\t\t\t\t\t\t6170_2_4\r\n    6170_4_1\t\t\t\t\t\t\t6170_4_1",
      "votes": 2
    },
    {
      "id": 150761,
      "postDate": "2016-12-16T16:02:06.290Z",
      "content": "<p>Try to start with <a href=\"https://github.com/ropensci/geojsonio\">https://github.com/ropensci/geojsonio</a> </p>\n\n<p>For example, to access to one element (6010 grid, element 4_4):</p>\n\n<pre><code>devtools::install_github(\"ropensci/geojsonio\")\nlibrary(\"geojsonio\")\n\ninstall.packages(\"rgdal\", type = \"source\")\ninstall.packages(\"rgeos\", type = \"source\")\nlibrary(\"rgdal\")\nlibrary(\"rgeos\")\nlibrary(ggplot2)\n\ngrid_6010_4_4 &lt;- geojson_read(\"./data/train_geojson/train_geojson/6010_4_4/Grid_6010.geojson\", method = local, what= 'sp')\n\nplot(grid_6010_4_4)\n\nggplot(grid_6010_4_4, aes(long, lat, group = group)) + geom_polygon()\n</code></pre>",
      "rawMarkdown": "Try to start with https://github.com/ropensci/geojsonio \r\n\r\nFor example, to access to one element (6010 grid, element 4_4):\r\n   \r\n\r\n    devtools::install_github(\"ropensci/geojsonio\")\r\n    library(\"geojsonio\")\r\n    \r\n    install.packages(\"rgdal\", type = \"source\")\r\n    install.packages(\"rgeos\", type = \"source\")\r\n    library(\"rgdal\")\r\n    library(\"rgeos\")\r\n    library(ggplot2)\r\n    \r\n    grid_6010_4_4 <- geojson_read(\"./data/train_geojson/train_geojson/6010_4_4/Grid_6010.geojson\", method = local, what= 'sp')\r\n    \r\n    plot(grid_6010_4_4)\r\n    \r\n    ggplot(grid_6010_4_4, aes(long, lat, group = group)) + geom_polygon()\r\n\r\n\r\n    \r\n",
      "votes": 2
    },
    {
      "id": 150575,
      "postDate": "2016-12-15T20:52:00.043Z",
      "content": "<p>Are these atmospherically corrected OR usually clear sky images? </p>",
      "rawMarkdown": "Are these atmospherically corrected OR usually clear sky images? ",
      "votes": 2
    },
    {
      "id": 152464,
      "postDate": "2016-12-26T13:53:54Z",
      "content": "<p>Wendy - question for you: I would like to enter this competition but am unable to download the datasets correctly due to their size - inevitably with multi-gigabyte files like these some glitch in the network links from my location to Kaggle causes the download to abend....   </p>\n\n<p>Even though I have fast fibre connections here, it took ~12 hours (and three tries) to get the 16-band data file downloaded, and at around 18 hours for the 3-band one it's proving impossible - this last time I got 13GB (87%) into the file before it hung..... frustrating in the extreme.</p>\n\n<p>So, any chance of getting these two huge zip files broken up into chunks of not more than a couple of GB each??</p>\n\n<p>And as a comment, if it is happening to me, it must be happening to others too, so can I suggest that Kaggle puts into place a 'rule' that causes large data files to be chunked up for all competitions...??</p>\n\n<p>Regards\nRick</p>",
      "rawMarkdown": "Wendy - question for you: I would like to enter this competition but am unable to download the datasets correctly due to their size - inevitably with multi-gigabyte files like these some glitch in the network links from my location to Kaggle causes the download to abend....   \r\n\r\nEven though I have fast fibre connections here, it took ~12 hours (and three tries) to get the 16-band data file downloaded, and at around 18 hours for the 3-band one it's proving impossible - this last time I got 13GB (87%) into the file before it hung..... frustrating in the extreme.\r\n\r\nSo, any chance of getting these two huge zip files broken up into chunks of not more than a couple of GB each??\r\n\r\nAnd as a comment, if it is happening to me, it must be happening to others too, so can I suggest that Kaggle puts into place a 'rule' that causes large data files to be chunked up for all competitions...??\r\n\r\nRegards\r\nRick"
    },
    {
      "id": 158345,
      "postDate": "2017-01-26T23:20:10.520Z",
      "content": "<p>@Wendy </p>\n\n<p>Is it possible that even if a submission file has broken polygons, it would go through, excluding lines with mistakes?</p>\n\n<p>Although it would be great to be able to have access to the error messages that occurred during the submission process.</p>",
      "rawMarkdown": "@Wendy \r\n\r\nIs it possible that even if a submission file has broken polygons, it would go through, excluding lines with mistakes?\r\n\r\nAlthough it would be great to be able to have access to the error messages that occurred during the submission process.",
      "votes": 1,
      "replies": [
        {
          "id": 158554,
          "postDate": "2017-01-28T01:12:02.017Z",
          "content": "<p>Do you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? </p>",
          "rawMarkdown": "Do you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? ",
          "replies": [
            {
              "id": 158568,
              "postDate": "2017-01-28T04:23:31.013Z",
              "content": "<p>[quote=Wendy Kan;158554]</p>\n\n<p>Do you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? </p>\n\n<p>[/quote]</p>\n\n<p>This wouldn't happen due to the evaluation formula, but in any case this idea of silently ignoring submission rows is a very bad idea as we'd be in the blind regarding which rows are hurting the score.</p>\n\n<p>The idea posted by Komaki in the other thread seems to be the only reasonable way of reliably evaluating polygon intersections like this (i.e. plotting polygons and calculating the formula over pixels not vectors). There are degenerate cases where even valid polygons can cause errors when calculating their intersection. The library used to calculate the intersection will generate an invalid polygon when doing the intersection,  and there's no guaranteed way to to make it work on all possible edge cases.</p>\n\n<p>@Wendy, I think the evaluation code used to calculate the score is an implementation detail. As far as the submission format is kept and scores are the same (up to rounding errors) the desire of DSTL to work with WKT in the competition is being respected.</p>",
              "rawMarkdown": "[quote=Wendy Kan;158554]\r\n\r\nDo you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? \r\n\r\n[/quote]\r\n\r\nThis wouldn't happen due to the evaluation formula, but in any case this idea of silently ignoring submission rows is a very bad idea as we'd be in the blind regarding which rows are hurting the score.\r\n\r\nThe idea posted by Komaki in the other thread seems to be the only reasonable way of reliably evaluating polygon intersections like this (i.e. plotting polygons and calculating the formula over pixels not vectors). There are degenerate cases where even valid polygons can cause errors when calculating their intersection. The library used to calculate the intersection will generate an invalid polygon when doing the intersection,  and there's no guaranteed way to to make it work on all possible edge cases.\r\n\r\n@Wendy, I think the evaluation code used to calculate the score is an implementation detail. As far as the submission format is kept and scores are the same (up to rounding errors) the desire of DSTL to work with WKT in the competition is being respected.\r\n\r\n\r\n\r\n",
              "votes": 2
            },
            {
              "id": 158575,
              "postDate": "2017-01-28T07:09:47.227Z",
              "content": "<p>[quote=Wendy Kan;158554]\nDo you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? \n[/quote]\n@Wendy Kan\nFor each couple of image/class where a problem is detected, the submitted multipolygon could be replaced by the polygon given in the sample file. I don't see how submitting broken polygons on purpose would give any advantage that way.\nA list of couples of image/class with broken polygons could be displayed as a warning.</p>\n\n<p>I suppose that modifying your scoring function this way wouldn't be too difficult. You could next consider further improvements if necessary.</p>\n\n<p>[quote=amaia;158568]\n...in any case this idea of silently ignoring submission rows is a very bad idea as we'd be in the blind regarding which rows are hurting the score.\n[/quote]\n@amaia\nIt doesn't have to be silently ignored. Rather than throwing an error, a warning could be displayed at the end of the submission with a list of all images where a problem was detected.  </p>",
              "rawMarkdown": "[quote=Wendy Kan;158554]\r\nDo you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? \r\n[/quote]\r\n@Wendy Kan\r\nFor each couple of image/class where a problem is detected, the submitted multipolygon could be replaced by the polygon given in the sample file. I don't see how submitting broken polygons on purpose would give any advantage that way.\r\nA list of couples of image/class with broken polygons could be displayed as a warning.\r\n\r\nI suppose that modifying your scoring function this way wouldn't be too difficult. You could next consider further improvements if necessary.\r\n\r\n[quote=amaia;158568]\r\n...in any case this idea of silently ignoring submission rows is a very bad idea as we'd be in the blind regarding which rows are hurting the score.\r\n[/quote]\r\n@amaia\r\nIt doesn't have to be silently ignored. Rather than throwing an error, a warning could be displayed at the end of the submission with a list of all images where a problem was detected.  \r\n",
              "votes": 1
            }
          ]
        },
        {
          "id": 160788,
          "postDate": "2017-02-09T14:16:41.127Z",
          "content": "<p>I completely agree with @amaia.\nPlotting expected and submitted polygons on \"big  enough\" layers (4000x4000 ?)  will avoid any geometry problem due to intersections even when submitting valid polygons.\nThis would allow to have more competitors and probably best solution at the end.</p>",
          "rawMarkdown": "I completely agree with @amaia.\nPlotting expected and submitted polygons on \"big  enough\" layers (4000x4000 ?)  will avoid any geometry problem due to intersections even when submitting valid polygons.\nThis would allow to have more competitors and probably best solution at the end.",
          "votes": 1
        }
      ]
    },
    {
      "id": 158264,
      "postDate": "2017-01-26T15:12:51.013Z",
      "content": "<p>@ Wendy <br>\nHi, <br>\nAny updates on the submission method or timeout problem?</p>\n\n<p>This days I have been making submissions of one class only to avoid the timeout error. If I combine the scores for all the classes I get an score of almost 0.44 on LB, instead because of simplifying I have a score of 0.37...</p>\n\n<p>I understand that the aim of the contest is to create the best model for segmenting satellite images, not to create a lossless method to compress wkt polygon files.</p>",
      "rawMarkdown": "@ Wendy   \r\nHi,   \r\nAny updates on the submission method or timeout problem?\r\n\r\nThis days I have been making submissions of one class only to avoid the timeout error. If I combine the scores for all the classes I get an score of almost 0.44 on LB, instead because of simplifying I have a score of 0.37...\r\n\r\nI understand that the aim of the contest is to create the best model for segmenting satellite images, not to create a lossless method to compress wkt polygon files.\r\n",
      "votes": 1,
      "replies": [
        {
          "id": 158315,
          "postDate": "2017-01-26T20:04:46.420Z",
          "content": "<p>Yes, update posted yesterday on increasing the timeout limit: <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/26506#158156\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/26506#158156</a></p>",
          "rawMarkdown": "Yes, update posted yesterday on increasing the timeout limit: https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/26506#158156",
          "votes": 1
        }
      ]
    },
    {
      "id": 159374,
      "postDate": "2017-02-02T01:29:42.197Z",
      "content": "<p>@Wendy - I also cannot get predictions through. I am pretty sure about my file format...</p>",
      "rawMarkdown": "@Wendy - I also cannot get predictions through. I am pretty sure about my file format...",
      "votes": -1
    },
    {
      "id": 159024,
      "postDate": "2017-01-31T09:52:50.530Z",
      "content": "<p>As people say above, projecting the wkt submission to an img and doing the evaluation on pixels will solve the two main problems: timeot errors and topology exception errors</p>",
      "rawMarkdown": "As people say above, projecting the wkt submission to an img and doing the evaluation on pixels will solve the two main problems: timeot errors and topology exception errors",
      "votes": -5
    },
    {
      "id": 163874,
      "postDate": "2017-02-26T16:03:15.040Z",
      "content": "<p>In the competition rules, it is said that</p>\n\n<pre><code>Pre-trained models are allowed in the competition, but need to be posted on the forum first.  \n</code></pre>\n\n<p>I'm going to use vgg 16 and vgg 19.<br>\n<a href=\"http://www.robots.ox.ac.uk/~vgg/\">http://www.robots.ox.ac.uk/~vgg/</a></p>",
      "rawMarkdown": "In the competition rules, it is said that\n\n    Pre-trained models are allowed in the competition, but need to be posted on the forum first.  \n\nI'm going to use vgg 16 and vgg 19.<br>\nhttp://www.robots.ox.ac.uk/~vgg/\n"
    },
    {
      "id": 163617,
      "postDate": "2017-02-24T19:38:13.760Z",
      "content": "<p>I'm using pretrained darknet model</p>\n\n<p><a href=\"https://pjreddie.com/darknet/yolo/\">https://pjreddie.com/darknet/yolo/</a></p>",
      "rawMarkdown": "I'm using pretrained darknet model\n\nhttps://pjreddie.com/darknet/yolo/"
    },
    {
      "id": 162211,
      "postDate": "2017-02-17T17:12:03.197Z",
      "content": "<p>I have made a submission and I got \"Error\" in my submissions page. How can I know the error?</p>",
      "rawMarkdown": "I have made a submission and I got \"Error\" in my submissions page. How can I know the error?"
    },
    {
      "id": 161642,
      "postDate": "2017-02-15T00:58:30.873Z",
      "content": "<p>Wendy, every time that I make a submission I get an Error. After clicking the Make Submission button I am redirected ot the Leaderboard, and I have no way of seeing the actual cause for the error. It has been very difficult for us to debug this issue, which has stopped us for the last 2 days. Please, can you help us with this issue?</p>",
      "rawMarkdown": "Wendy, every time that I make a submission I get an Error. After clicking the Make Submission button I am redirected ot the Leaderboard, and I have no way of seeing the actual cause for the error. It has been very difficult for us to debug this issue, which has stopped us for the last 2 days. Please, can you help us with this issue?"
    },
    {
      "id": 160810,
      "postDate": "2017-02-09T16:50:31.457Z",
      "content": "<p>Hi @wendy , until today kernel jupyter-notebooks contained python gdal module. However today \"import gdal\" fails. Did something changed in between? \nThanks</p>",
      "rawMarkdown": "Hi @wendy , until today kernel jupyter-notebooks contained python gdal module. However today \"import gdal\" fails. Did something changed in between? \nThanks"
    },
    {
      "id": 159293,
      "postDate": "2017-02-01T18:36:58.140Z",
      "content": "<p>@ Wendy, <br>\nI can't understand why painting the polygons is so slowly on your method. I have attached a python code for evaluating submissions. <br>\nIt takes 40s to evaluate the train set, and using intersections can take more than 30 minutes. </p>\n\n<p>I'm also attaching a submission for the train set so you can repeat the experiments.</p>",
      "rawMarkdown": "@ Wendy,   \r\nI can't understand why painting the polygons is so slowly on your method. I have attached a python code for evaluating submissions.   \r\nIt takes 40s to evaluate the train set, and using intersections can take more than 30 minutes. \r\n\r\nI'm also attaching a submission for the train set so you can repeat the experiments.",
      "replies": [
        {
          "id": 159387,
          "postDate": "2017-02-02T03:25:39.497Z",
          "content": "<p>Probably because I don't have access to cv2.fillpoly(). </p>",
          "rawMarkdown": "Probably because I don't have access to cv2.fillpoly(). "
        },
        {
          "id": 159845,
          "postDate": "2017-02-04T09:52:58.477Z",
          "content": "<p>Hi Wendy, <br>\nSorry to insist in the timeout problem. <br>\nYesterday I was able to submit a 13MB file at the cost of throwing away the tree class.  The score was 0.43 which shows is a good submission. <br>\nIncluding the tree class the size was of only 22MB. It is very simplified but still get timeout error. </p>\n\n<p>I think it's nonsense because I'm making a submission for a number of images that is x20 the train size, and I have to make a file that has a size almost equal to the train file (11MB).</p>\n\n<p>Please you have to solve this, I don't know if the solution is to double the time limit, change evaluating method to use cv2 or other.  But it's nonsense to train a good model if later I have to pass the prediction throught a very narrow bottleneck where a lot of information is lost.</p>\n\n<p>Thank you <br>\nironbar</p>",
          "rawMarkdown": "Hi Wendy,   \nSorry to insist in the timeout problem.  \nYesterday I was able to submit a 13MB file at the cost of throwing away the tree class.  The score was 0.43 which shows is a good submission.   \nIncluding the tree class the size was of only 22MB. It is very simplified but still get timeout error. \n\nI think it's nonsense because I'm making a submission for a number of images that is x20 the train size, and I have to make a file that has a size almost equal to the train file (11MB).\n\nPlease you have to solve this, I don't know if the solution is to double the time limit, change evaluating method to use cv2 or other.  But it's nonsense to train a good model if later I have to pass the prediction throught a very narrow bottleneck where a lot of information is lost.\n\nThank you   \nironbar",
          "votes": 2
        },
        {
          "id": 161779,
          "postDate": "2017-02-15T19:18:41.127Z",
          "content": "<p>@ Wendy, <br>\nThere are only 20 days until the end of the challenge. I still have timeout problems. <br>\nAt this point I don't think you are going to change evaluation method for the faster raster image method. But why you don't increase the timeout limit? I don't understand it. If it is a matter of computational power now I have to make many submissions that timeout until one succeeds, so I don't see the save. </p>\n\n<p>Thanks\nironbar</p>",
          "rawMarkdown": "@ Wendy,    \nThere are only 20 days until the end of the challenge. I still have timeout problems.   \nAt this point I don't think you are going to change evaluation method for the faster raster image method. But why you don't increase the timeout limit? I don't understand it. If it is a matter of computational power now I have to make many submissions that timeout until one succeeds, so I don't see the save. \n\nThanks\nironbar",
          "votes": 1
        }
      ]
    },
    {
      "id": 159183,
      "postDate": "2017-02-01T07:15:32.573Z",
      "content": "<p>Dear Wendy,</p>\n\n<p>Could you clarify if CUDA, cuBLAS and cuDNN are allowed as dependencies in this competition?</p>\n\n<p>In other words, if my code requires CUDA and can not be run without it, will I be disqualified?</p>\n\n<p>For context: <a href=\"https://www.kaggle.com/c/data-science-bowl-2017/forums/t/27820/matlab?forumMessageId=159167\">https://www.kaggle.com/c/data-science-bowl-2017/forums/t/27820/matlab?forumMessageId=159167</a> (DSB-2017 has the same license type)</p>\n\n<p>If this is left up to the discretion of the organizers (as seems to be the case), does this mean that one person may be disqualified for having CUDA as a requirement, while another might not be? That doesn't seem very sportsmanlike.</p>",
      "rawMarkdown": "Dear Wendy,\r\n\r\nCould you clarify if CUDA, cuBLAS and cuDNN are allowed as dependencies in this competition?\r\n\r\nIn other words, if my code requires CUDA and can not be run without it, will I be disqualified?\r\n\r\nFor context: https://www.kaggle.com/c/data-science-bowl-2017/forums/t/27820/matlab?forumMessageId=159167 (DSB-2017 has the same license type)\r\n\r\nIf this is left up to the discretion of the organizers (as seems to be the case), does this mean that one person may be disqualified for having CUDA as a requirement, while another might not be? That doesn't seem very sportsmanlike.",
      "replies": [
        {
          "id": 159388,
          "postDate": "2017-02-02T03:26:01.897Z",
          "content": "<p>I think you're safe. </p>",
          "rawMarkdown": "I think you're safe. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 158817,
      "postDate": "2017-01-30T09:04:55.603Z",
      "content": "<p>I wonder what the organizers tried to accomplish by allowing one but not the other: Some ML models are fully interconvertible with their training datasets. </p>",
      "rawMarkdown": "I wonder what the organizers tried to accomplish by allowing one but not the other: Some ML models are fully interconvertible with their training datasets. "
    },
    {
      "id": 158815,
      "postDate": "2017-01-30T08:37:39.297Z",
      "content": "<p>External data is not allowed.</p>\n\n<p>Pretrained yes if 1) you post on the forum and 2) the license allows.</p>",
      "rawMarkdown": "External data is not allowed.\r\n\r\nPretrained yes if 1) you post on the forum and 2) the license allows.\r\n"
    },
    {
      "id": 158802,
      "postDate": "2017-01-30T06:45:04.187Z",
      "content": "<p>Just to verify, are we allowed to use pre-trained networks and external data as long as they are published at the forum? </p>",
      "rawMarkdown": "Just to verify, are we allowed to use pre-trained networks and external data as long as they are published at the forum? "
    },
    {
      "id": 158157,
      "postDate": "2017-01-25T19:46:44.107Z",
      "content": "<p>This part is clear. </p>\n\n<p>I am more about, participants do not see the score on the Private LB, so from their perspective no need to perform calculations on the 81% of the data which may help to decrease errors that are related to the submission size and exceeding 8min limit.</p>",
      "rawMarkdown": "This part is clear. \r\n\r\nI am more about, participants do not see the score on the Private LB, so from their perspective no need to perform calculations on the 81% of the data which may help to decrease errors that are related to the submission size and exceeding 8min limit."
    },
    {
      "id": 154890,
      "postDate": "2017-01-08T16:01:43.340Z",
      "content": "<p>Hi @Wendy Kan , When I submit my results, I got an error that says \"Evaluation Exception: premature end of enumerator.\" </p>\n\n<p>Some results of our submission are in the submission file.</p>\n\n<p>Could you or could someone else help me find out how to solve this problem? </p>\n\n<p>Thank you very much！!！</p>",
      "rawMarkdown": "Hi @Wendy Kan , When I submit my results, I got an error that says \"Evaluation Exception: premature end of enumerator.\" \r\n\r\nSome results of our submission are in the submission file.\r\n\r\nCould you or could someone else help me find out how to solve this problem? \r\n\r\nThank you very much！!！",
      "replies": [
        {
          "id": 155172,
          "postDate": "2017-01-10T00:00:19.850Z",
          "content": "<p>@guangliang2016, just replied to your other thread. </p>",
          "rawMarkdown": "@guangliang2016, just replied to your other thread. "
        }
      ]
    },
    {
      "id": 151459,
      "postDate": "2016-12-20T22:01:54.997Z",
      "content": "<p>Looks like the PR was merged.  Note, I also added the geojson module so (soon) tifffile, descartes, and geojson python modules should be available in the kernels. </p>",
      "rawMarkdown": "Looks like the PR was merged.  Note, I also added the geojson module so (soon) tifffile, descartes, and geojson python modules should be available in the kernels. "
    },
    {
      "id": 150873,
      "postDate": "2016-12-17T07:19:58.290Z",
      "content": "<p>Wendy,</p>\n\n<p>Could ypu please inform us if we will can use train_geojson.zip (both to download and in kernels)? </p>\n\n<p>Thanks </p>",
      "rawMarkdown": "Wendy,\r\n\r\nCould ypu please inform us if we will can use train_geojson.zip (both to download and in kernels)? \r\n\r\nThanks "
    },
    {
      "id": 150772,
      "postDate": "2016-12-16T17:18:14.103Z",
      "content": "<p>Thanks that helps a lot. A long way to go :)</p>",
      "rawMarkdown": "Thanks that helps a lot. A long way to go :)"
    },
    {
      "id": 150756,
      "postDate": "2016-12-16T15:37:57.350Z",
      "content": "<p>How do we read the tiff files using R? The data processing tutorial only seems to cover Python and I'm having difficulty unzipping the tiffs in R:</p>\n\n<p><a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial</a></p>",
      "rawMarkdown": "How do we read the tiff files using R? The data processing tutorial only seems to cover Python and I'm having difficulty unzipping the tiffs in R:\r\n\r\nhttps://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial"
    },
    {
      "id": 150502,
      "postDate": "2016-12-15T13:11:46.893Z",
      "content": "<p>Hi! </p>\n\n<p>Thanks for this interesting competition!</p>\n\n<p>I think the link to data reading tutorial is dead (at least I get a 404 now) </p>\n\n<p><a href=\"https://www.kaggle.com/c/dstl-satellite-image-object-detection/details/data-processing-tutorial\">https://www.kaggle.com/c/dstl-satellite-image-object-detection/details/data-processing-tutorial</a></p>\n\n<p>Cristi</p>\n\n<p><strong>Update</strong></p>\n\n<p>The link is in the Data section. The Data Processing Tutorial link on the left menu area works.</p>",
      "rawMarkdown": "Hi! \r\n\r\nThanks for this interesting competition!\r\n\r\nI think the link to data reading tutorial is dead (at least I get a 404 now) \r\n\r\nhttps://www.kaggle.com/c/dstl-satellite-image-object-detection/details/data-processing-tutorial\r\n\r\nCristi\r\n\r\n**Update**\r\n\r\nThe link is in the Data section. The Data Processing Tutorial link on the left menu area works.\r\n"
    },
    {
      "id": 151327,
      "postDate": "2016-12-20T04:54:09.860Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 150951,
      "postDate": "2016-12-17T20:14:58.330Z",
      "content": "<p>Ok. Thank you</p>",
      "rawMarkdown": "Ok. Thank you"
    }
  ],
  "comments": [
    {
      "id": 159329,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2017-02-01T20:47:27.213000",
      "content": "<p>I can't submit file in new interface. The only thing I've got is: \"Your submission was unsuccessful.\"</p>",
      "votes": 5,
      "replies": [
        {
          "id": 159359,
          "author_name": "Kohei",
          "author_url": "",
          "post_date": "2017-02-01T23:51:34.923000",
          "content": "<p>same here. also, the counter of remaining submission is broken. on the \"Submit Predictions\" page, \"You have -3 submissions today.\"</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 159360,
          "author_name": "Kyle Lee",
          "author_url": "",
          "post_date": "2017-02-02T00:22:51.693000",
          "content": "<p>@Wendy - couple of other bugs, other than submissions not getting through in this new interface:</p>\n\n<ol>\n<li><p>The daily quota counter decrements even though there is an error with submission (I just tried submitting today.  I suppose you can get negative submissions with this interface per response above.</p></li>\n<li><p>\"This leaderboard is calculated with approximately 1% of the test data.\nThe final results will be based on the other 99%, so the final standings may be different.\".  I suppose it is still 19% public?</p></li>\n</ol>\n\n<p>also, comparing the LB from half an hour back it seems some submissions went through, so it could either be a filesize or compression type (I am using .gz) change from the last interface?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 159419,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2017-02-02T07:01:49.193000",
          "content": "<p>Some more problems: </p>\n\n<p>1) I often save useful data in filename. In new interface filename is not visible.</p>\n\n<p>2) I don't see how I can download my old submission.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 159620,
          "author_name": "Ben Hamner",
          "author_url": "",
          "post_date": "2017-02-03T02:59:55.503000",
          "content": "<p>@ZFTurbo we're definitely fixing 1) and 2). Are you still having submission issues to this competition? I didn't see a failing submission from you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 159651,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2017-02-03T05:42:05.890000",
          "content": "<p><strong>@Ben Hamner</strong>\nI still have a problem. I uploaded ZIP file with size around 200 MB. After it reached 100% I pressed \"Make submission\" it redirect me to Leaderboard without any notifications. When I go to \"My submissions\" page I don't see my latest submission. If I try to reupload the same ZIP-file Progressbar reach 100% instantly.</p>\n\n<p>I also see text errors on \"My submissions\" page:</p>\n\n<p>1) 90 submissions for ZFTurbo - As I remember I have less than 90 submissions in total.</p>\n\n<p>2) $100,000 · 0 teams · a month to go (a month to go until merger deadline)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 151324,
      "author_name": "Gabriel Altay",
      "author_url": "",
      "post_date": "2016-12-20T04:19:58.307000",
      "content": "<p>woohoo, my first contribution to the Kaggle github repo :)</p>\n\n<p><a href=\"https://github.com/Kaggle/docker-python/pull/48\">https://github.com/Kaggle/docker-python/pull/48</a></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 158643,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2017-01-28T18:19:41.287000",
      "content": "<p>@Wendy, from a business perspective, it looks that set up for this competition has some flaws.</p>\n\n<p>Only quarter of the participants are above the sample benchmark.... and it is not because problem is that hard, I mean it is hard to get a good result, but to get above benchmark is straightforward, any FCN will get your there.</p>\n\n<p>I would believe that DSTL wants to get a good solution, Kaggle wants to ensure that DSTL gets a good solution or at least decent interest from the Data Science community. </p>\n\n<p>I agree that Kaggle community is lazy and relaxed in a way that people do not want to invest their time on the engineering issues. But still, it is just easier to give up on this problem and switch to some other challenge.</p>\n\n<p>Boring Allstate competition attracted 3000+ participants, and this exciting problem has only 50 above benchmark and 150 below it. I do not know who is project owner for this problem, but it may be time to rethink a way that submission is performed to have satisfied client in 38 days.</p>\n\n<p>I really like an idea to do submission as polygons, as the client requires, transform it to mask and do the evaluation on the mask level.</p>\n\n<p>Or to provide a kernel with a function that will do mapping: mask =&gt; polygons in such a way that submission that is created with this function will not raise errors. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 159107,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-31T20:30:12.843000",
          "content": "<p>Hi Vladimir, </p>\n\n<p>Thanks for the suggestion. In fact, your suggestion of submission as polygons -&gt; transform into mask -&gt; evaluation was the original idea. However, after some experiments, I found that the performance is quite poor. The bottle neck was mainly in the generating of mask from polygons. When the geometry gets more complex, the repeated calls to polygon.contains(pixel) took a very long time since it had to be repeated for every pixel and was about 50-100x slower than calculating Jaccard directly on the polygons. Therefore, we re-wrote the metric to be purely vector-based. </p>\n\n<p>The metric execution time of this competition has been a challenge from the very beginning and we went through quite some iterations on designing this competition. Very different versions of metric and different implementations were experimented. We understand it's not going to be as traditional or trivial as some other Kaggle competitions, therefore not going to be as popular. But we do want to design these non-conventional competitions to expose our community to a completely different problem space. </p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 160163,
      "author_name": "Attila",
      "author_url": "",
      "post_date": "2017-02-06T12:02:35.513000",
      "content": "<p>@wendy Kan Why can we not see the error message that we can use to get before. Debugging with the new interface without error message is very tough.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 160256,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-02-06T20:10:37.763000",
          "content": "<p>@Attila,</p>\n\n<p>Very sorry about this - we are trying to fix this bug. In the mean time,  I'm going to email you with your submission errors as a temporary fix. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 160784,
          "author_name": "Cogitae _ Thomas Soumarmon",
          "author_url": "",
          "post_date": "2017-02-09T13:46:51.277000",
          "content": "<p>Hi Wendy,</p>\n\n<p>Not having the error message is a problem. I have a file that passes the tpex test tool (\"Good to go!\")  but is on error when submiting.\nI also have one submission that is still being processed since nearly 15 hours. Is this blocking further submissions ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 160794,
          "author_name": "Attila",
          "author_url": "",
          "post_date": "2017-02-09T14:38:11.733000",
          "content": "<p>@WendyKan Thanks a lot for sending the error for each submission. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 161306,
          "author_name": "Iñigo del Portillo",
          "author_url": "",
          "post_date": "2017-02-13T04:53:21.533000",
          "content": "<p>@Wendy, do you have an estimate about when this error is going to be fixed. It's really difficult to debug without any knowledge of what's going on during the evaluation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 159335,
      "author_name": "Alexey Lukyanov",
      "author_url": "",
      "post_date": "2017-02-01T21:34:03.027000",
      "content": "<p>@Wendy,\nIs it possible to see an error message with the new interface? It is much easier to debug a code if you know the type of exception occurred.  It is especially important in this type of competition, where almost every participant experience submission problems. <br>\nUpdate: I got \"Evaluation Exception: Submission must have 4290 rows\" for one of my submissions. But submission file must have 1 header line and 4290 lines for each image =&gt; 4291 in total. What should I do to fix that problem? \nThanks in advance!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 158357,
      "author_name": "amaia",
      "author_url": "",
      "post_date": "2017-01-27T02:05:26.383000",
      "content": "<p>There's a much more probable explanation for timeouts.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 157399,
      "author_name": "Wendy Kan",
      "author_url": "",
      "post_date": "2017-01-20T19:52:34.690000",
      "content": "<p>@ironbar,</p>\n\n<p>Yes, I tested the perfect submission prior to launching this competition and it finished evaluation under 6 mins, which is within the 8 minute time limit. </p>\n\n<p>I'm working with the engineering team to see if we can lift this limit for this competition. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 157463,
          "author_name": "Guillermo Barbadillo",
          "author_url": "",
          "post_date": "2017-01-21T10:25:43.550000",
          "content": "<p>@ Wendy, <br>\nHi thanks for your answer, <br>\nAs people is saying here, one alternative for computing the score is projecting the polygons to a raster and computing the score pixel-wise. <br>\n<a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/27673/non-noded-intersection-inconsistent-polygon-validity-results\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/forums/t/27673/non-noded-intersection-inconsistent-polygon-validity-results</a></p>\n\n<p>I have tested already on the train set and it's much faster than making the polygon intersection, and the accuracy is very similar when using a raster of 4000x4000 or 8000x8000.  </p>\n\n<p>The speedup can be more than x10. <br>\nAlso that would eliminate the problem of Topology exception.</p>\n\n<p>I hope you and your team can come up to a good solution, because right now I have to downsize the submission from 96MB to 9MB to avoid the timeout error. And that hurts the score a lot. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 158156,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-25T19:45:43.253000",
          "content": "<p>update: We've upped the evaluation time limit to 10 mins (instead of 8) for this competition. </p>\n\n<p>I understand that pixel-based evaluation will be faster than geometry operations. However, we can't use a pixel based evaluation as it is part of the requirement from Dstl that the output of the algorithm has to be polygons. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 158331,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2017-01-26T22:10:26.367000",
          "content": "<p>@Wendy</p>\n\n<blockquote>\n  <p>update: We've upped the evaluation time limit to 10 mins (instead of 8)</p>\n</blockquote>\n\n<p>Would that be real time or user time? Is it strongly affected by how busy your server is?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 158343,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-26T23:18:22.410000",
          "content": "<p>It's not affected by the busy-ness of server. The worker node starts its clock after we received your complete submission file (so your upload speed also doesn't matter). </p>",
          "votes": 1,
          "replies": [
            {
              "id": 158355,
              "author_name": "Oleg Trott",
              "author_url": "",
              "post_date": "2017-01-27T01:02:46.217000",
              "content": "<p>[quote=Wendy Kan;158343]</p>\n\n<p>It's not affected by the busy-ness of server. The worker node starts its clock after we received your complete submission file (so your upload speed also doesn't matter). </p>\n\n<p>[/quote]</p>\n\n<p>Wendy,</p>\n\n<p>I meant \"real time\" and \"user time\" in the Unix sense:</p>\n\n<p><a href=\"http://unix.stackexchange.com/questions/53302/why-would-the-real-time-be-much-higher-than-the-user-and-sys-times-combine\">http://unix.stackexchange.com/questions/53302/why-would-the-real-time-be-much-higher-than-the-user-and-sys-times-combine</a></p>\n\n<p>The former is what most people refer to as the \"clock\", but it's also very strongly affected by whether your worker node is processing other submissions simultaneously.</p>\n\n<p>That might explain why the users are timing out.</p>",
              "votes": -2,
              "replies": []
            }
          ]
        },
        {
          "id": 158514,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-27T19:56:30.070000",
          "content": "<p>It's real/clock time. However, the prediction competition scoring workers only do one request at a time, so they shouldn't be affected by noisy neighbors. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 157219,
      "author_name": "Guillermo Barbadillo",
      "author_url": "",
      "post_date": "2017-01-19T18:28:14.903000",
      "content": "<p>Hi Wendy, <br>\nI'm concerned about submission timeout. <br>\n¿Have you tried to submit the perfect submission?  ¿Does it have enough time to compute the score?</p>\n\n<p>I'm asking because from my experience:</p>\n\n<ul>\n<li>We are given polygons for 25 images for training</li>\n<li>The file of this polygons uses 50.3MB uncompressed (the precision of the points is 6)</li>\n<li>When we have to make the submission we are predicting for 429 images, so the file should be approximately 20 times bigger -&gt; 1000MB uncompressed</li>\n<li>In my case I can't submit files bigger than 36MB when using precision 6 (like in the training data). To get down to that size I have to simplify the prediction a lot and it hurts the score.</li>\n<li>If the files are bigger I get timeout error.</li>\n</ul>\n\n<p>So there's a big difference between the estimated size that a perfect submission will have and the size that until now I have been able to submit.</p>\n\n<p>Thanks\nIronbar</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 150763,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2016-12-16T16:17:44.447000",
      "content": "<p>And to access images:</p>\n\n<pre><code>library(raster)\nraster_6044_4_4 &lt;- raster(\"./data/three_band/6040_4_4.tif\")\nplot(raster_6044_4_4)\n\nlibrary(rgdal)\ngdal_6044_4_4 &lt;- readGDAL(paste0(\"./data/three_band/\", '6040_4_4', \".tif\"))\nplot(gdal_6044_4_4)\n</code></pre>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 150787,
      "author_name": "Wendy Kan",
      "author_url": "",
      "post_date": "2016-12-16T18:21:17.540000",
      "content": "<p>@ZFTurbo:</p>\n\n<p>Apologies! Our random train/test split picked something that has all empty multipolygons for class 7. They do exist in the test dataset. We will release another image that has class 7 shortly. </p>\n\n<p>EDIT: 6070_2_3 has been moved to train_wkt_v2.csv. It has non-empty multipolygon for class 7. Submission file remains the same while this image is ignored for scoring. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 150574,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2016-12-15T20:42:39.053000",
      "content": "<p>ClassType = 7 absent in train_wkt.csv. Does it exist in test?</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 150481,
      "author_name": "Yunfeng Zhu",
      "author_url": "",
      "post_date": "2016-12-15T11:41:32.023000",
      "content": "<p>Can I ask what the Jaccard Index value will be if the predicted and actual MultipolygonWKT are both \"MULTIPOLYGON EMPTY\"? Thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 150587,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2016-12-15T22:00:24.610000",
          "content": "<p>@Yunfeng, </p>\n\n<p>The Jaccard index is not per-image, it's calculating a total TP/FP/FN over a bunch of images for a given class. So if they are both empty, then TP=FP=FN=0, so totalTP/totalFP/totalFN don't change for this image. </p>\n\n<p>[quote=Yunfeng Zhu;150481]</p>\n\n<p>Can I ask what the Jaccard Index value will be if the predicted and actual MultipolygonWKT are both \"MULTIPOLYGON EMPTY\"? Thanks.</p>\n\n<p>[/quote]</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 166588,
      "author_name": "Cogitae _ Thomas Soumarmon",
      "author_url": "",
      "post_date": "2017-03-10T08:14:59.950000",
      "content": "<p>Hi @wendykan,\nnow that the competition is over are we allowed to submit ? There are ideas I haven't finish to explore and would like to check how well it performs against test dataset ? If yes, how long will it be open ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 167462,
          "author_name": "Cogitae _ Thomas Soumarmon",
          "author_url": "",
          "post_date": "2017-03-14T06:27:41.987000",
          "content": "<p>Hi @Wendy,\nHaving no reply on the forum I sent you a direct message. Did you receive it ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 167600,
          "author_name": "Devin Anzelmo",
          "author_url": "",
          "post_date": "2017-03-14T17:43:52.313000",
          "content": "<p>I have been submitting to check ideas all week. Usually we have unlimited submissions once the competition is over,  but it looks like right now we only have 3 per day.  It would be nice if this limit was lifted. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 158197,
      "author_name": "Komaki",
      "author_url": "",
      "post_date": "2017-01-26T02:47:14.503000",
      "content": "<p>@Wendy </p>\n\n<p>Couldn't you round numbers in your wkt to remove too precise numbers just in case? Basically, the \"Non-Noded Intersection\" errors are unavoidable as long as intersections are calculated because of float number precision. But, I feel numbers like -0.007879000000000001 among rounded numbers are more likely to cause this problem. </p>\n\n<pre><code>6040_2_2,4,\"MULTIPOLYGON (((0.003025 -0.007879000000000001, 0.003074 -0.007931000000000001, 0.003123 -0.007996, 0.003182 \n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 158348,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-26T23:27:41.703000",
          "content": "<p>Interesting idea. Let me look into this. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 158548,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-28T00:28:43.937000",
          "content": "<p>Some updates to this: I reduced the rounding_precision in the geometry and did some testing, and it resulted in <strong>more</strong> non-noded intersection errors, even in the sample_submission file. </p>\n\n<p>I'll test more, but so far it doesn't seem like a good direction to go. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 158561,
              "author_name": "Komaki",
              "author_url": "",
              "post_date": "2017-01-28T02:21:35.887000",
              "content": "<p>[quote=Wendy Kan;158548]</p>\n\n<p>Some updates to this: I reduced the rounding_precision in the geometry and did some testing, and it resulted in <strong>more</strong> non-noded intersection errors, even in the sample_submission file. </p>\n\n<p>I'll test more, but so far it doesn't seem like a good direction to go. </p>\n\n<p>[/quote]</p>\n\n<p>I'm no expert, but I guess getting errors for sample_submission indicates geometries are broken. They were broken possibly because two non-intersected lines like (-1, 0) -&gt; (0, 0.99999999999) -&gt; (1, 0) and (1, 1) -&gt; (0, 1) -&gt; (1, 1) get touched to each other when rounded. I'm sorry, simply rounding numbers looks a wrong solution. Thank you, anyway.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 157681,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2017-01-23T03:31:29.410000",
      "content": "<p>@Wendy, could you please clarify?</p>\n\n<p>Currently when we do submission is score calculated only for Public LB (19% of the test set images)\nor for the whole train set?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 158154,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-25T19:37:30.353000",
          "content": "<p>Your public LB score is based on the public split of the test dataset (19%). At the end of the competition you'll see your private LB scores that are based on the rest of the test set. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 159464,
          "author_name": "arnowaczynski",
          "author_url": "",
          "post_date": "2017-02-02T11:38:28.970000",
          "content": "<p>@Wendy Kan, Do you know something about split between public and private? Is it estimated to be:\n19-81% (as has been given from the beginning of the competition) or 1-99% (as has been written in the current statement on leaderboard page)?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 153797,
      "author_name": "guangliang2016",
      "author_url": "",
      "post_date": "2017-01-03T12:19:23.187000",
      "content": "<p>@Wendy Kan Hi, Kan. I have some problems in the submission. In the train_wkt_v4.csv, the results are listed with Multipolygon format, while in the sample_submission.csv, the results are listed with Polygon format, I was wondering, in the submission file, which format should we take? When I tried to submit the result with Multipolygon, I got some errors as following figure shows. Thank you very much! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 153893,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-03T23:06:06.067000",
          "content": "<p>Hi @guangliang2016, both Polygon and MultiPolygon formats should work as long as they are valid. Seems like you were able to make some submissions without errors after this. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 151261,
      "author_name": "Wendy Kan",
      "author_url": "",
      "post_date": "2016-12-19T19:21:02.080000",
      "content": "<p>@Gabriel Altay,</p>\n\n<p>Thanks for the suggestion! You're welcome to submit a pull request at <a href=\"https://github.com/Kaggle/docker-python\">https://github.com/Kaggle/docker-python</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 151238,
      "author_name": "Gabriel Altay",
      "author_url": "",
      "post_date": "2016-12-19T16:20:48.527000",
      "content": "<p>@WendyKan would it be possible to make the following python packages available in the Kaggle kernel? </p>\n\n<ul>\n<li>tifffile (for reading multi-band tiff files) <a href=\"https://pypi.python.org/pypi/tifffile\">https://pypi.python.org/pypi/tifffile</a></li>\n<li>descartes (for dealing with polygons) <a href=\"https://pypi.python.org/pypi/descartes/1.0.2\">https://pypi.python.org/pypi/descartes/1.0.2</a></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 150933,
      "author_name": "Wendy Kan",
      "author_url": "",
      "post_date": "2016-12-17T17:45:49.353000",
      "content": "<p>@shawn, Sorry about the discrepancy, something broke during the process of adding that one image to the training set. Please use train_wkt_v3.csv now. The train_geojson.zip will be updated on Monday. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 150503,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2016-12-15T13:14:50.647000",
      "content": "<p>Try this:</p>\n\n<p><a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 152703,
      "author_name": "Gabriel Altay",
      "author_url": "",
      "post_date": "2016-12-27T23:52:13.927000",
      "content": "<p>@RickTurner, I think your best bet would be to use a download manager that supports \"resume\" \n<a href=\"http://www.ghacks.net/2014/08/03/download-large-files/\">http://www.ghacks.net/2014/08/03/download-large-files/</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 150857,
      "author_name": "shawn",
      "author_url": "",
      "post_date": "2016-12-17T03:10:32.263000",
      "content": "<p>I just looked at the different train v2 formats and they do not have the same image id's in each.</p>\n\n<pre><code>diff -y in_wkt_csv in_geojson \n6010_4_2                              | 6010_1_2\n6010_4_4                            6010_4_4\n6040_1_0                            6040_1_0\n6040_1_3                            6040_1_3\n6040_2_2                            6040_2_2\n                                  &gt; 6040_4_4\n6060_2_3                            6060_2_3\n6070_2_3                            6070_2_3\n6090_2_0                            6090_2_0\n6100_1_3                            6100_1_3\n                                  &gt; 6100_2_2\n6100_2_3                            6100_2_3\n6110_1_2                            6110_1_2\n6110_3_1                            6110_3_1\n6110_4_0                              &lt;\n6120_2_0                            6120_2_0\n6120_2_2                            6120_2_2\n6140_1_2                            6140_1_2\n6140_3_1                              &lt;\n6150_2_3                              &lt;\n6160_2_1                            6160_2_1\n6170_0_4                            6170_0_4\n6170_2_4                            6170_2_4\n6170_4_1                            6170_4_1\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 150761,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2016-12-16T16:02:06.290000",
      "content": "<p>Try to start with <a href=\"https://github.com/ropensci/geojsonio\">https://github.com/ropensci/geojsonio</a> </p>\n\n<p>For example, to access to one element (6010 grid, element 4_4):</p>\n\n<pre><code>devtools::install_github(\"ropensci/geojsonio\")\nlibrary(\"geojsonio\")\n\ninstall.packages(\"rgdal\", type = \"source\")\ninstall.packages(\"rgeos\", type = \"source\")\nlibrary(\"rgdal\")\nlibrary(\"rgeos\")\nlibrary(ggplot2)\n\ngrid_6010_4_4 &lt;- geojson_read(\"./data/train_geojson/train_geojson/6010_4_4/Grid_6010.geojson\", method = local, what= 'sp')\n\nplot(grid_6010_4_4)\n\nggplot(grid_6010_4_4, aes(long, lat, group = group)) + geom_polygon()\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 150575,
      "author_name": "Manish Verma",
      "author_url": "",
      "post_date": "2016-12-15T20:52:00.043000",
      "content": "<p>Are these atmospherically corrected OR usually clear sky images? </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 152464,
      "author_name": "RickTurner",
      "author_url": "",
      "post_date": "2016-12-26T13:53:54",
      "content": "<p>Wendy - question for you: I would like to enter this competition but am unable to download the datasets correctly due to their size - inevitably with multi-gigabyte files like these some glitch in the network links from my location to Kaggle causes the download to abend....   </p>\n\n<p>Even though I have fast fibre connections here, it took ~12 hours (and three tries) to get the 16-band data file downloaded, and at around 18 hours for the 3-band one it's proving impossible - this last time I got 13GB (87%) into the file before it hung..... frustrating in the extreme.</p>\n\n<p>So, any chance of getting these two huge zip files broken up into chunks of not more than a couple of GB each??</p>\n\n<p>And as a comment, if it is happening to me, it must be happening to others too, so can I suggest that Kaggle puts into place a 'rule' that causes large data files to be chunked up for all competitions...??</p>\n\n<p>Regards\nRick</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 158345,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2017-01-26T23:20:10.520000",
      "content": "<p>@Wendy </p>\n\n<p>Is it possible that even if a submission file has broken polygons, it would go through, excluding lines with mistakes?</p>\n\n<p>Although it would be great to be able to have access to the error messages that occurred during the submission process.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 158554,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-28T01:12:02.017000",
          "content": "<p>Do you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 158568,
              "author_name": "amaia",
              "author_url": "",
              "post_date": "2017-01-28T04:23:31.013000",
              "content": "<p>[quote=Wendy Kan;158554]</p>\n\n<p>Do you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? </p>\n\n<p>[/quote]</p>\n\n<p>This wouldn't happen due to the evaluation formula, but in any case this idea of silently ignoring submission rows is a very bad idea as we'd be in the blind regarding which rows are hurting the score.</p>\n\n<p>The idea posted by Komaki in the other thread seems to be the only reasonable way of reliably evaluating polygon intersections like this (i.e. plotting polygons and calculating the formula over pixels not vectors). There are degenerate cases where even valid polygons can cause errors when calculating their intersection. The library used to calculate the intersection will generate an invalid polygon when doing the intersection,  and there's no guaranteed way to to make it work on all possible edge cases.</p>\n\n<p>@Wendy, I think the evaluation code used to calculate the score is an implementation detail. As far as the submission format is kept and scores are the same (up to rounding errors) the desire of DSTL to work with WKT in the competition is being respected.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 158575,
              "author_name": "Vincent L.",
              "author_url": "",
              "post_date": "2017-01-28T07:09:47.227000",
              "content": "<p>[quote=Wendy Kan;158554]\nDo you mean you would still get a score even with broken polygons? What if someone intentionally submits broken polygons so they're all ignored, then wins the competition with one very good polygon prediction? \n[/quote]\n@Wendy Kan\nFor each couple of image/class where a problem is detected, the submitted multipolygon could be replaced by the polygon given in the sample file. I don't see how submitting broken polygons on purpose would give any advantage that way.\nA list of couples of image/class with broken polygons could be displayed as a warning.</p>\n\n<p>I suppose that modifying your scoring function this way wouldn't be too difficult. You could next consider further improvements if necessary.</p>\n\n<p>[quote=amaia;158568]\n...in any case this idea of silently ignoring submission rows is a very bad idea as we'd be in the blind regarding which rows are hurting the score.\n[/quote]\n@amaia\nIt doesn't have to be silently ignored. Rather than throwing an error, a warning could be displayed at the end of the submission with a list of all images where a problem was detected.  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 160788,
          "author_name": "Cogitae _ Thomas Soumarmon",
          "author_url": "",
          "post_date": "2017-02-09T14:16:41.127000",
          "content": "<p>I completely agree with @amaia.\nPlotting expected and submitted polygons on \"big  enough\" layers (4000x4000 ?)  will avoid any geometry problem due to intersections even when submitting valid polygons.\nThis would allow to have more competitors and probably best solution at the end.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 158264,
      "author_name": "Guillermo Barbadillo",
      "author_url": "",
      "post_date": "2017-01-26T15:12:51.013000",
      "content": "<p>@ Wendy <br>\nHi, <br>\nAny updates on the submission method or timeout problem?</p>\n\n<p>This days I have been making submissions of one class only to avoid the timeout error. If I combine the scores for all the classes I get an score of almost 0.44 on LB, instead because of simplifying I have a score of 0.37...</p>\n\n<p>I understand that the aim of the contest is to create the best model for segmenting satellite images, not to create a lossless method to compress wkt polygon files.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 158315,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2017-01-26T20:04:46.420000",
          "content": "<p>Yes, update posted yesterday on increasing the timeout limit: <a href=\"https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/26506#158156\">https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/discussion/26506#158156</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 159374,
      "author_name": "J W Victor",
      "author_url": "",
      "post_date": "2017-02-02T01:29:42.197000",
      "content": "<p>@Wendy - I also cannot get predictions through. I am pretty sure about my file format...</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 159024,
      "author_name": "Guillermo Barbadillo",
      "author_url": "",
      "post_date": "2017-01-31T09:52:50.530000",
      "content": "<p>As people say above, projecting the wkt submission to an img and doing the evaluation on pixels will solve the two main problems: timeot errors and topology exception errors</p>",
      "votes": -5,
      "replies": []
    },
    {
      "id": 163874,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-02-26T16:03:15.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 163617,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-02-24T19:38:13.760000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 162211,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-02-17T17:12:03.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 161642,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-02-15T00:58:30.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 160810,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-02-09T16:50:31.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 159293,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-02-01T18:36:58.140000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 159387,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-02-02T03:25:39.497000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 159845,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-02-04T09:52:58.477000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 161779,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-02-15T19:18:41.127000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 159183,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-02-01T07:15:32.573000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 159388,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-02-02T03:26:01.897000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 158817,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-30T09:04:55.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 158815,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-30T08:37:39.297000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 158802,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-30T06:45:04.187000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 158157,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-25T19:46:44.107000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154890,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T16:01:43.340000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 155172,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-01-10T00:00:19.850000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 151459,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-20T22:01:54.997000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 150873,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-17T07:19:58.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 150772,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-16T17:18:14.103000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 150756,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-16T15:37:57.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 150502,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-15T13:11:46.893000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 151327,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-20T04:54:09.860000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 150951,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-17T20:14:58.330000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "150394": "On behalf of Kaggle and Dstl, I'd like to welcome you to Dstl Satellite Imagery Feature Detection competition. Satellite imagery and geometries are a new type of data for the Kaggle community, and we spent a good amount of time to make it work. We hope you enjoy the experience dealing with this type of data. \r\n\r\nWe're happy to answer your questions here on the forum. Good luck and have fun!",
    "159329": "I can't submit file in new interface. The only thing I've got is: \"Your submission was unsuccessful.\"",
    "151324": "woohoo, my first contribution to the Kaggle github repo :)\r\n\r\nhttps://github.com/Kaggle/docker-python/pull/48",
    "158643": "@Wendy, from a business perspective, it looks that set up for this competition has some flaws.\r\n\r\nOnly quarter of the participants are above the sample benchmark.... and it is not because problem is that hard, I mean it is hard to get a good result, but to get above benchmark is straightforward, any FCN will get your there.\r\n\r\nI would believe that DSTL wants to get a good solution, Kaggle wants to ensure that DSTL gets a good solution or at least decent interest from the Data Science community. \r\n\r\nI agree that Kaggle community is lazy and relaxed in a way that people do not want to invest their time on the engineering issues. But still, it is just easier to give up on this problem and switch to some other challenge.\r\n\r\nBoring Allstate competition attracted 3000+ participants, and this exciting problem has only 50 above benchmark and 150 below it. I do not know who is project owner for this problem, but it may be time to rethink a way that submission is performed to have satisfied client in 38 days.\r\n\r\nI really like an idea to do submission as polygons, as the client requires, transform it to mask and do the evaluation on the mask level.\r\n\r\nOr to provide a kernel with a function that will do mapping: mask => polygons in such a way that submission that is created with this function will not raise errors. ",
    "160163": "@wendy Kan Why can we not see the error message that we can use to get before. Debugging with the new interface without error message is very tough.",
    "159335": "@Wendy,\nIs it possible to see an error message with the new interface? It is much easier to debug a code if you know the type of exception occurred.  It is especially important in this type of competition, where almost every participant experience submission problems.  \nUpdate: I got \"Evaluation Exception: Submission must have 4290 rows\" for one of my submissions. But submission file must have 1 header line and 4290 lines for each image => 4291 in total. What should I do to fix that problem? \nThanks in advance!",
    "158357": "There's a much more probable explanation for timeouts.",
    "157399": "@ironbar,\n\nYes, I tested the perfect submission prior to launching this competition and it finished evaluation under 6 mins, which is within the 8 minute time limit. \n\nI'm working with the engineering team to see if we can lift this limit for this competition. ",
    "157219": "Hi Wendy,   \r\nI'm concerned about submission timeout.   \r\n¿Have you tried to submit the perfect submission?  ¿Does it have enough time to compute the score?\r\n\r\nI'm asking because from my experience:\r\n\r\n* We are given polygons for 25 images for training\r\n* The file of this polygons uses 50.3MB uncompressed (the precision of the points is 6)\r\n* When we have to make the submission we are predicting for 429 images, so the file should be approximately 20 times bigger -> 1000MB uncompressed\r\n* In my case I can't submit files bigger than 36MB when using precision 6 (like in the training data). To get down to that size I have to simplify the prediction a lot and it hurts the score.\r\n* If the files are bigger I get timeout error.\r\n\r\nSo there's a big difference between the estimated size that a perfect submission will have and the size that until now I have been able to submit.\r\n\r\n\r\nThanks\r\nIronbar",
    "150763": "And to access images:\r\n\r\n    library(raster)\r\n    raster_6044_4_4 <- raster(\"./data/three_band/6040_4_4.tif\")\r\n    plot(raster_6044_4_4)\r\n    \r\n    library(rgdal)\r\n    gdal_6044_4_4 <- readGDAL(paste0(\"./data/three_band/\", '6040_4_4', \".tif\"))\r\n    plot(gdal_6044_4_4)",
    "150787": "@ZFTurbo:\r\n\r\nApologies! Our random train/test split picked something that has all empty multipolygons for class 7. They do exist in the test dataset. We will release another image that has class 7 shortly. \r\n\r\nEDIT: 6070_2_3 has been moved to train_wkt_v2.csv. It has non-empty multipolygon for class 7. Submission file remains the same while this image is ignored for scoring. ",
    "150574": "ClassType = 7 absent in train_wkt.csv. Does it exist in test?",
    "150481": "Can I ask what the Jaccard Index value will be if the predicted and actual MultipolygonWKT are both \"MULTIPOLYGON EMPTY\"? Thanks.",
    "166588": "Hi @wendykan,\nnow that the competition is over are we allowed to submit ? There are ideas I haven't finish to explore and would like to check how well it performs against test dataset ? If yes, how long will it be open ?",
    "158197": "@Wendy \r\n\r\nCouldn't you round numbers in your wkt to remove too precise numbers just in case? Basically, the \"Non-Noded Intersection\" errors are unavoidable as long as intersections are calculated because of float number precision. But, I feel numbers like -0.007879000000000001 among rounded numbers are more likely to cause this problem. \r\n\r\n    6040_2_2,4,\"MULTIPOLYGON (((0.003025 -0.007879000000000001, 0.003074 -0.007931000000000001, 0.003123 -0.007996, 0.003182 \r\n\r\n",
    "157681": "@Wendy, could you please clarify?\r\n\r\nCurrently when we do submission is score calculated only for Public LB (19% of the test set images)\r\nor for the whole train set?",
    "153797": "@Wendy Kan Hi, Kan. I have some problems in the submission. In the train_wkt_v4.csv, the results are listed with Multipolygon format, while in the sample_submission.csv, the results are listed with Polygon format, I was wondering, in the submission file, which format should we take? When I tried to submit the result with Multipolygon, I got some errors as following figure shows. Thank you very much! ",
    "151261": "@Gabriel Altay,\r\n\r\nThanks for the suggestion! You're welcome to submit a pull request at https://github.com/Kaggle/docker-python \r\n\r\n",
    "151238": "@WendyKan would it be possible to make the following python packages available in the Kaggle kernel? \r\n\r\n - tifffile (for reading multi-band tiff files) https://pypi.python.org/pypi/tifffile\r\n - descartes (for dealing with polygons) https://pypi.python.org/pypi/descartes/1.0.2",
    "150933": "@shawn, Sorry about the discrepancy, something broke during the process of adding that one image to the training set. Please use train_wkt_v3.csv now. The train_geojson.zip will be updated on Monday. ",
    "150503": "Try this:\r\n\r\nhttps://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial",
    "152703": "@RickTurner, I think your best bet would be to use a download manager that supports \"resume\" \r\nhttp://www.ghacks.net/2014/08/03/download-large-files/",
    "150857": "I just looked at the different train v2 formats and they do not have the same image id's in each.\r\n\r\n   \r\n    diff -y in_wkt_csv in_geojson \r\n    6010_4_2\t\t\t\t\t\t      |\t6010_1_2\r\n    6010_4_4\t\t\t\t\t\t\t6010_4_4\r\n    6040_1_0\t\t\t\t\t\t\t6040_1_0\r\n    6040_1_3\t\t\t\t\t\t\t6040_1_3\r\n    6040_2_2\t\t\t\t\t\t\t6040_2_2\r\n    \t\t\t\t\t\t\t      >\t6040_4_4\r\n    6060_2_3\t\t\t\t\t\t\t6060_2_3\r\n    6070_2_3\t\t\t\t\t\t\t6070_2_3\r\n    6090_2_0\t\t\t\t\t\t\t6090_2_0\r\n    6100_1_3\t\t\t\t\t\t\t6100_1_3\r\n    \t\t\t\t\t\t\t      >\t6100_2_2\r\n    6100_2_3\t\t\t\t\t\t\t6100_2_3\r\n    6110_1_2\t\t\t\t\t\t\t6110_1_2\r\n    6110_3_1\t\t\t\t\t\t\t6110_3_1\r\n    6110_4_0\t\t\t\t\t\t      <\r\n    6120_2_0\t\t\t\t\t\t\t6120_2_0\r\n    6120_2_2\t\t\t\t\t\t\t6120_2_2\r\n    6140_1_2\t\t\t\t\t\t\t6140_1_2\r\n    6140_3_1\t\t\t\t\t\t      <\r\n    6150_2_3\t\t\t\t\t\t      <\r\n    6160_2_1\t\t\t\t\t\t\t6160_2_1\r\n    6170_0_4\t\t\t\t\t\t\t6170_0_4\r\n    6170_2_4\t\t\t\t\t\t\t6170_2_4\r\n    6170_4_1\t\t\t\t\t\t\t6170_4_1",
    "150761": "Try to start with https://github.com/ropensci/geojsonio \r\n\r\nFor example, to access to one element (6010 grid, element 4_4):\r\n   \r\n\r\n    devtools::install_github(\"ropensci/geojsonio\")\r\n    library(\"geojsonio\")\r\n    \r\n    install.packages(\"rgdal\", type = \"source\")\r\n    install.packages(\"rgeos\", type = \"source\")\r\n    library(\"rgdal\")\r\n    library(\"rgeos\")\r\n    library(ggplot2)\r\n    \r\n    grid_6010_4_4 <- geojson_read(\"./data/train_geojson/train_geojson/6010_4_4/Grid_6010.geojson\", method = local, what= 'sp')\r\n    \r\n    plot(grid_6010_4_4)\r\n    \r\n    ggplot(grid_6010_4_4, aes(long, lat, group = group)) + geom_polygon()\r\n\r\n\r\n    \r\n",
    "150575": "Are these atmospherically corrected OR usually clear sky images? ",
    "152464": "Wendy - question for you: I would like to enter this competition but am unable to download the datasets correctly due to their size - inevitably with multi-gigabyte files like these some glitch in the network links from my location to Kaggle causes the download to abend....   \r\n\r\nEven though I have fast fibre connections here, it took ~12 hours (and three tries) to get the 16-band data file downloaded, and at around 18 hours for the 3-band one it's proving impossible - this last time I got 13GB (87%) into the file before it hung..... frustrating in the extreme.\r\n\r\nSo, any chance of getting these two huge zip files broken up into chunks of not more than a couple of GB each??\r\n\r\nAnd as a comment, if it is happening to me, it must be happening to others too, so can I suggest that Kaggle puts into place a 'rule' that causes large data files to be chunked up for all competitions...??\r\n\r\nRegards\r\nRick",
    "158345": "@Wendy \r\n\r\nIs it possible that even if a submission file has broken polygons, it would go through, excluding lines with mistakes?\r\n\r\nAlthough it would be great to be able to have access to the error messages that occurred during the submission process.",
    "158264": "@ Wendy   \r\nHi,   \r\nAny updates on the submission method or timeout problem?\r\n\r\nThis days I have been making submissions of one class only to avoid the timeout error. If I combine the scores for all the classes I get an score of almost 0.44 on LB, instead because of simplifying I have a score of 0.37...\r\n\r\nI understand that the aim of the contest is to create the best model for segmenting satellite images, not to create a lossless method to compress wkt polygon files.\r\n",
    "159374": "@Wendy - I also cannot get predictions through. I am pretty sure about my file format...",
    "159024": "As people say above, projecting the wkt submission to an img and doing the evaluation on pixels will solve the two main problems: timeot errors and topology exception errors",
    "163874": "In the competition rules, it is said that\n\n    Pre-trained models are allowed in the competition, but need to be posted on the forum first.  \n\nI'm going to use vgg 16 and vgg 19.<br>\nhttp://www.robots.ox.ac.uk/~vgg/\n",
    "163617": "I'm using pretrained darknet model\n\nhttps://pjreddie.com/darknet/yolo/",
    "162211": "I have made a submission and I got \"Error\" in my submissions page. How can I know the error?",
    "161642": "Wendy, every time that I make a submission I get an Error. After clicking the Make Submission button I am redirected ot the Leaderboard, and I have no way of seeing the actual cause for the error. It has been very difficult for us to debug this issue, which has stopped us for the last 2 days. Please, can you help us with this issue?",
    "160810": "Hi @wendy , until today kernel jupyter-notebooks contained python gdal module. However today \"import gdal\" fails. Did something changed in between? \nThanks",
    "159293": "@ Wendy,   \r\nI can't understand why painting the polygons is so slowly on your method. I have attached a python code for evaluating submissions.   \r\nIt takes 40s to evaluate the train set, and using intersections can take more than 30 minutes. \r\n\r\nI'm also attaching a submission for the train set so you can repeat the experiments.",
    "159183": "Dear Wendy,\r\n\r\nCould you clarify if CUDA, cuBLAS and cuDNN are allowed as dependencies in this competition?\r\n\r\nIn other words, if my code requires CUDA and can not be run without it, will I be disqualified?\r\n\r\nFor context: https://www.kaggle.com/c/data-science-bowl-2017/forums/t/27820/matlab?forumMessageId=159167 (DSB-2017 has the same license type)\r\n\r\nIf this is left up to the discretion of the organizers (as seems to be the case), does this mean that one person may be disqualified for having CUDA as a requirement, while another might not be? That doesn't seem very sportsmanlike.",
    "158817": "I wonder what the organizers tried to accomplish by allowing one but not the other: Some ML models are fully interconvertible with their training datasets. ",
    "158815": "External data is not allowed.\r\n\r\nPretrained yes if 1) you post on the forum and 2) the license allows.\r\n",
    "158802": "Just to verify, are we allowed to use pre-trained networks and external data as long as they are published at the forum? ",
    "158157": "This part is clear. \r\n\r\nI am more about, participants do not see the score on the Private LB, so from their perspective no need to perform calculations on the 81% of the data which may help to decrease errors that are related to the submission size and exceeding 8min limit.",
    "154890": "Hi @Wendy Kan , When I submit my results, I got an error that says \"Evaluation Exception: premature end of enumerator.\" \r\n\r\nSome results of our submission are in the submission file.\r\n\r\nCould you or could someone else help me find out how to solve this problem? \r\n\r\nThank you very much！!！",
    "151459": "Looks like the PR was merged.  Note, I also added the geojson module so (soon) tifffile, descartes, and geojson python modules should be available in the kernels. ",
    "150873": "Wendy,\r\n\r\nCould ypu please inform us if we will can use train_geojson.zip (both to download and in kernels)? \r\n\r\nThanks ",
    "150772": "Thanks that helps a lot. A long way to go :)",
    "150756": "How do we read the tiff files using R? The data processing tutorial only seems to cover Python and I'm having difficulty unzipping the tiffs in R:\r\n\r\nhttps://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/details/data-processing-tutorial",
    "150502": "Hi! \r\n\r\nThanks for this interesting competition!\r\n\r\nI think the link to data reading tutorial is dead (at least I get a 404 now) \r\n\r\nhttps://www.kaggle.com/c/dstl-satellite-image-object-detection/details/data-processing-tutorial\r\n\r\nCristi\r\n\r\n**Update**\r\n\r\nThe link is in the Data section. The Data Processing Tutorial link on the left menu area works.\r\n",
    "151327": "",
    "150951": "Ok. Thank you"
  }
}