{
  "id": 30135,
  "title": "Private Leaderboard Final!",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/30135",
  "author_name": "",
  "post_date": "2017-03-15T19:29:51.604640200Z",
  "votes": 9,
  "comment_count": 18,
  "views": 0,
  "content": "<p>The private leaderboard has now been finalized and published.  Thank you for your patience, and congratulations to the winners!  This was a challenging competition with a unique custom metric and progressive area of study in detecting satellite imagery features.  Our community did not disappoint and clearly rose to the challenge.  Kudos to all of you for your impassioned participation and active sharing across forums.  We look forward to diving into the winners' solutions!</p>",
  "messages": [
    {
      "id": "167901",
      "postDate": "03/15/2017 19:29:51",
      "content": "<p>The private leaderboard has now been finalized and published.  Thank you for your patience, and congratulations to the winners!  This was a challenging competition with a unique custom metric and progressive area of study in detecting satellite imagery features.  Our community did not disappoint and clearly rose to the challenge.  Kudos to all of you for your impassioned participation and active sharing across forums.  We look forward to diving into the winners' solutions!</p>",
      "rawMarkdown": "The private leaderboard has now been finalized and published.  Thank you for your patience, and congratulations to the winners!  This was a challenging competition with a unique custom metric and progressive area of study in detecting satellite imagery features.  Our community did not disappoint and clearly rose to the challenge.  Kudos to all of you for your impassioned participation and active sharing across forums.  We look forward to diving into the winners' solutions!",
      "votes": null
    },
    {
      "id": "167906",
      "postDate": "03/15/2017 19:31:51",
      "content": "<p>I would like to give sincere thanks to both Dstl and to all of you that participated. From the very beginning of this project, we knew that this was a challenging problem. The way to measure the performance of the algorithms were difficult, and was an engineering challenge to make it run efficiently given the data size. There were also a lot of challenges to ensure the quality of the hand-labels for ground truth data and to ensure that they form valid polygons and multipolygons. Dstl is a great partner of Kaggle. We are thankful for the professionalism and patience they showed in working with us preparing the data and throughout the competition. </p>\n\n<p>Kaggle competitions community had not worked with multi-spectral satellite imagery data before, and it was not an easy task to even set up the environment to get started. I offered a Docker image that worked for me and was thrilled to see Kagglers quickly collaborated and contributed to merge it into Kaggle's Docker for Kernels. Throughout the competition, we continue to see great sharing of tools/methods/insights. We're extremely proud of our community working on a tough problem together. </p>\n\n<p>Instance segmentation (where the algorithm marks exactly where the object is with its outline) is an advanced topic. It remains a tough problem in today's computer vision research. We would like to congratulate everyone for finishing! </p>",
      "rawMarkdown": "I would like to give sincere thanks to both Dstl and to all of you that participated. From the very beginning of this project, we knew that this was a challenging problem. The way to measure the performance of the algorithms were difficult, and was an engineering challenge to make it run efficiently given the data size. There were also a lot of challenges to ensure the quality of the hand-labels for ground truth data and to ensure that they form valid polygons and multipolygons. Dstl is a great partner of Kaggle. We are thankful for the professionalism and patience they showed in working with us preparing the data and throughout the competition. \n\nKaggle competitions community had not worked with multi-spectral satellite imagery data before, and it was not an easy task to even set up the environment to get started. I offered a Docker image that worked for me and was thrilled to see Kagglers quickly collaborated and contributed to merge it into Kaggle's Docker for Kernels. Throughout the competition, we continue to see great sharing of tools/methods/insights. We're extremely proud of our community working on a tough problem together. \n\nInstance segmentation (where the algorithm marks exactly where the object is with its outline) is an advanced topic. It remains a tough problem in today's computer vision research. We would like to congratulate everyone for finishing!",
      "votes": null
    },
    {
      "id": "167918",
      "postDate": "03/15/2017 20:11:20",
      "content": "<p>It was a very exciting problem. Thank you for hosting it. I hope we will see more similar problems in future.</p>\n\n<p>May I ask why did you split in such a way that there is only 25 images in train? Even 50 will make this problem significantly easier.</p>",
      "rawMarkdown": "It was a very exciting problem. Thank you for hosting it. I hope we will see more similar problems in future.\n\nMay I ask why did you split in such a way that there is only 25 images in train? Even 50 will make this problem significantly easier.",
      "votes": null
    },
    {
      "id": "167957",
      "postDate": "03/15/2017 22:29:28",
      "content": "<p>Interesting question. You may be surprised, but the total number of labeled images is only 57. So we did a split of 25/32. Since we were concerned that having only 32 images in the test set would attract people to hand label, we had to \"salt\" the test set with unlabeled images that are ignored when scoring. </p>\n\n<p>This decision was to ensure we have a successful competition. We also knew that adding these ignored image would slow down people's iteration time in developing the algorithm. It also adds to the total size of the submission files as they do get large very quickly! Eventually, we figured it's more important to keep the competition fair, so we still kept the \"salted\" images in there. </p>\n\n<p>It may seem like a very small hand-labeled dataset but it's also an example of how valuable the ground truth data is. It was a lot of work to get those images labeled correctly. </p>",
      "rawMarkdown": "Interesting question. You may be surprised, but the total number of labeled images is only 57. So we did a split of 25/32. Since we were concerned that having only 32 images in the test set would attract people to hand label, we had to \"salt\" the test set with unlabeled images that are ignored when scoring. \n\nThis decision was to ensure we have a successful competition. We also knew that adding these ignored image would slow down people's iteration time in developing the algorithm. It also adds to the total size of the submission files as they do get large very quickly! Eventually, we figured it's more important to keep the competition fair, so we still kept the \"salted\" images in there. \n\nIt may seem like a very small hand-labeled dataset but it's also an example of how valuable the ground truth data is. It was a lot of work to get those images labeled correctly.",
      "votes": null
    },
    {
      "id": "167980",
      "postDate": "03/15/2017 23:58:28",
      "content": "<p>This was a really great competition.</p>\n\n<p>Even tough I could  only submit a dummy solution (so I will have to provide my solution at another place - e.g. academic paper e.a. ) I really enjoyed participating in this Kaggle competition.</p>\n\n<p>At the same time it looks like there is a substantial flaw in your scoring system - It can't be that there is a difference in position if the score is the same.</p>\n\n<p>Within the lower scores where you have the same scores repeating it seems most obvious e.g.  with 0.27837 or even more extreme with the score 0.07181 when - with the same score - you are either on 259th place or on 391st place randomly - let's call that fuzzy as the best euphemism  I could come of with.</p>\n\n<p>Overall you are discrediting yourself as a data science platform if such a thing is not remedied.</p>\n\n<p>At the same time - given all the data cleansing e.a. challenges in this competition - as the sponsor I would be strongly disappointed ( and these are - as a UK tax payer - my taxes not yours at work) - getting back solution(s) that is / are doing basically solely a brute force approach and still are coming back below 50%.</p>\n\n<p>Nevertheless, most importantly in this business we all learn from things we do and not from things we only talk about or watch others doing.</p>\n\n<p>With that alone this competition is a success even though it has not immediately forwarded potential results - this can most likely be achieved in a follow up when those who have invested substantially in that competition will be prepared to actually challenge ( look at the number of submissions for some of the competitors as a first indicator to understand this better).</p>\n\n<p>Many thanks again for this opportunity - I certainly learned a lot - now I will have to see where to apply this newly found knowledge - a $2'000 competition certainly will not be the place where I can see people to come forward with their skills or newly shaped experience. Like with other competitions the next stage will have to at least double the price money. </p>",
      "rawMarkdown": "This was a really great competition.\n\nEven tough I could  only submit a dummy solution (so I will have to provide my solution at another place - e.g. academic paper e.a. ) I really enjoyed participating in this Kaggle competition.\n\nAt the same time it looks like there is a substantial flaw in your scoring system - It can't be that there is a difference in position if the score is the same.\n\nWithin the lower scores where you have the same scores repeating it seems most obvious e.g.  with 0.27837 or even more extreme with the score 0.07181 when - with the same score - you are either on 259th place or on 391st place randomly - let's call that fuzzy as the best euphemism  I could come of with.\n\nOverall you are discrediting yourself as a data science platform if such a thing is not remedied.\n\nAt the same time - given all the data cleansing e.a. challenges in this competition - as the sponsor I would be strongly disappointed ( and these are - as a UK tax payer - my taxes not yours at work) - getting back solution(s) that is / are doing basically solely a brute force approach and still are coming back below 50%.\n\nNevertheless, most importantly in this business we all learn from things we do and not from things we only talk about or watch others doing.\n\nWith that alone this competition is a success even though it has not immediately forwarded potential results - this can most likely be achieved in a follow up when those who have invested substantially in that competition will be prepared to actually challenge ( look at the number of submissions for some of the competitors as a first indicator to understand this better).\n\nMany thanks again for this opportunity - I certainly learned a lot - now I will have to see where to apply this newly found knowledge - a $2'000 competition certainly will not be the place where I can see people to come forward with their skills or newly shaped experience. Like with other competitions the next stage will have to at least double the price money.",
      "votes": null
    },
    {
      "id": "167985",
      "postDate": "03/16/2017 00:22:48",
      "content": "<p>Has anybody done a validation of the ground truth like I've seen with other projects e.g. do a secondary validation to get an indication how much of the \"ground truth\" is actually valid / properly classified - example 6070_2_3 where large areas are classified as single areas of trees - IMHO each of those large areas are a mix of shadows, some water, trees e.a. a figure of 80% would be friendly - additionally some class 2 structures provided as training data in other scenes that are merely 1m2 in size again below a shadow / covered partially by another object e.g. trees - I guess it was intended to be a challenge but some of that makes it rather random if you're not looking at results above e.g. 80%+. </p>",
      "rawMarkdown": "Has anybody done a validation of the ground truth like I've seen with other projects e.g. do a secondary validation to get an indication how much of the \"ground truth\" is actually valid / properly classified - example 6070_2_3 where large areas are classified as single areas of trees - IMHO each of those large areas are a mix of shadows, some water, trees e.a. a figure of 80% would be friendly - additionally some class 2 structures provided as training data in other scenes that are merely 1m2 in size again below a shadow / covered partially by another object e.g. trees - I guess it was intended to be a challenge but some of that makes it rather random if you're not looking at results above e.g. 80%+.",
      "votes": null
    },
    {
      "id": "167988",
      "postDate": "03/16/2017 00:38:57",
      "content": "<p>Thank you for this competition. </p>\n\n<p>It was really enjoyable to have graphic predictions and a challenging subject on the technical side. Really liked those hours spent optimizing my models to have a decent execution time.</p>",
      "rawMarkdown": "Thank you for this competition. \n\nIt was really enjoyable to have graphic predictions and a challenging subject on the technical side. Really liked those hours spent optimizing my models to have a decent execution time.",
      "votes": null
    },
    {
      "id": "167999",
      "postDate": "03/16/2017 01:27:30",
      "content": "<p>Whilst I was looking for the press release about the results, I found that dstl put at least some of the work in creating the ground truth out to tender:</p>\n\n<p><a href=\"http://www.publictenders.net/node/3556044\">http://www.publictenders.net/node/3556044</a></p>\n\n<blockquote>\n  <p>This Requirement contributes to a work package of the Large Scale Data\n  Processing and OSINT (LSDP&amp;O) Project in the Intelligence Countering\n  Adversary Networks (ICAN) Programme by providing the ground truth\n  imagery needed to set a Kaggle Challenge.</p>\n  \n  <p>...</p>\n  \n  <p>Value excluding VAT: 148 821.00 GBP</p>\n</blockquote>\n\n<p>Currently about $180k, which is nearly double the whole prize pool!</p>\n\n<p>I thought, with 450 images, that is £330 per image, or £33 per class per image, not bad value...</p>\n\n<p>It's a bit disappointing to hear it was only 57 images, that's £2.6k per image.</p>\n\n<p>Given all the images fit into 18 5x5 grids, surely the original plan was for all of those to be annotated, and the subcontractor just ran out of time? ;)</p>\n\n<p>It also makes me wonder how they did the annotation - from working on entries for this competition I saw first hand how it would be a lot easier to use a well trained convnet and manually correct it's predictions than to create annotations from scratch.</p>",
      "rawMarkdown": "Whilst I was looking for the press release about the results, I found that dstl put at least some of the work in creating the ground truth out to tender:\n\nhttp://www.publictenders.net/node/3556044\n\n> This Requirement contributes to a work package of the Large Scale Data\n> Processing and OSINT (LSDP&O) Project in the Intelligence Countering\n> Adversary Networks (ICAN) Programme by providing the ground truth\n> imagery needed to set a Kaggle Challenge.\n\n> ...\n\n> Value excluding VAT: 148 821.00 GBP\n\nCurrently about $180k, which is nearly double the whole prize pool!\n\nI thought, with 450 images, that is £330 per image, or £33 per class per image, not bad value...\n\nIt's a bit disappointing to hear it was only 57 images, that's £2.6k per image.\n\nGiven all the images fit into 18 5x5 grids, surely the original plan was for all of those to be annotated, and the subcontractor just ran out of time? ;)\n\nIt also makes me wonder how they did the annotation - from working on entries for this competition I saw first hand how it would be a lot easier to use a well trained convnet and manually correct it's predictions than to create annotations from scratch.",
      "votes": null
    },
    {
      "id": "168006",
      "postDate": "03/16/2017 01:56:25",
      "content": "<p>If you will see another contract to annotate Satelite images especially as these ones from Nigeria - let me know. :)</p>",
      "rawMarkdown": "If you will see another contract to annotate Satelite images especially as these ones from Nigeria - let me know. :)",
      "votes": null
    },
    {
      "id": "168079",
      "postDate": "03/16/2017 07:02:21",
      "content": "<p>Thank you for managing interesting competition.<br>\nThere were many new settings for us (charactaristics of dataset, evaluation metrix, submission format and so on).\nI really enjoyed all of them.</p>",
      "rawMarkdown": "Thank you for managing interesting competition.<br>\nThere were many new settings for us (charactaristics of dataset, evaluation metrix, submission format and so on).\nI really enjoyed all of them.",
      "votes": null
    },
    {
      "id": "168098",
      "postDate": "03/16/2017 08:34:42",
      "content": "<p>Nice competition! </p>\n\n<p>I have a question, my last solution is still processing. I assume that it will never end? (It was submitted a day before the competition ended). Now,  I don't expect to be anything spectacular but it would be nice to know if throwing more compute power helped.</p>",
      "rawMarkdown": "Nice competition! \n\nI have a question, my last solution is still processing. I assume that it will never end? (It was submitted a day before the competition ended). Now,  I don't expect to be anything spectacular but it would be nice to know if throwing more compute power helped.",
      "votes": null
    },
    {
      "id": "168104",
      "postDate": "03/16/2017 09:01:10",
      "content": "<p>This sounds a little cheating. At least it should have mentioned actual number of test images that would be used for scoring, though it would have lead to more chaos.</p>\n\n<p>Still a better and more fair scoring can be done with a simple strategy.  i.e. take best performing submissions (or ensemble of them) <strong>per class</strong> and use them as truths to re score the submissions on entire data set.\nWhat do you think guys?</p>",
      "rawMarkdown": "This sounds a little cheating. At least it should have mentioned actual number of test images that would be used for scoring, though it would have lead to more chaos.\n\nStill a better and more fair scoring can be done with a simple strategy.  i.e. take best performing submissions (or ensemble of them) **per class** and use them as truths to re score the submissions on entire data set.\nWhat do you think guys?",
      "votes": null
    },
    {
      "id": "168113",
      "postDate": "03/16/2017 09:21:19",
      "content": "<p>I think it's a pity making a submission for 450 images and being evaluated only on 32. <br>\nI don't think it's a representative number to measure the goodness of the models.  </p>",
      "rawMarkdown": "I think it's a pity making a submission for 450 images and being evaluated only on 32.  \nI don't think it's a representative number to measure the goodness of the models.",
      "votes": null
    },
    {
      "id": "168388",
      "postDate": "03/17/2017 00:57:13",
      "content": "<p>Similar thing for me. Last submission ended with error. It would be nice to know what score it gets.</p>",
      "rawMarkdown": "Similar thing for me. Last submission ended with error. It would be nice to know what score it gets.",
      "votes": null
    },
    {
      "id": "168616",
      "postDate": "03/17/2017 15:50:05",
      "content": "<p>Have the same problem. my last solution submitted 3 days back it's still processing....</p>",
      "rawMarkdown": "Have the same problem. my last solution submitted 3 days back it's still processing....",
      "votes": null
    },
    {
      "id": "168742",
      "postDate": "03/17/2017 23:34:55",
      "content": "<p>@visoft, @Thian, @Naaga, </p>\n\n<p>If you want to know your score, please submit again. I think it most likely timed out but didn't get to throw that error because the site was busy, but you can always try again to see the score. </p>",
      "rawMarkdown": "visoft, @Thian, @Naaga, \n\nIf you want to know your score, please submit again. I think it most likely timed out but didn't get to throw that error because the site was busy, but you can always try again to see the score.",
      "votes": null
    },
    {
      "id": "168865",
      "postDate": "03/18/2017 12:20:20",
      "content": "<p>Thanks! I will try to locate the submision file (I deleted most of the files already).</p>",
      "rawMarkdown": "Thanks! I will try to locate the submision file (I deleted most of the files already).",
      "votes": null
    },
    {
      "id": "168888",
      "postDate": "03/18/2017 14:23:49",
      "content": "<p>It is a 11G (~4g zip) csv file. I don't think it will be evaluated. But i did the submit, just in case.</p>",
      "rawMarkdown": "It is a 11G (~4g zip) csv file. I don't think it will be evaluated. But i did the submit, just in case.",
      "votes": null
    },
    {
      "id": "251956",
      "postDate": "12/02/2017 00:26:27",
      "content": "<p>It is very interesting problem.\nWould you please publish the labeled test set (32 images) for whom are going to practice on this problem and evaluate their works?</p>",
      "rawMarkdown": "It is very interesting problem.\nWould you please publish the labeled test set (32 images) for whom are going to practice on this problem and evaluate their works?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 167906,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "03/15/2017 19:31:51",
      "content": "<p>I would like to give sincere thanks to both Dstl and to all of you that participated. From the very beginning of this project, we knew that this was a challenging problem. The way to measure the performance of the algorithms were difficult, and was an engineering challenge to make it run efficiently given the data size. There were also a lot of challenges to ensure the quality of the hand-labels for ground truth data and to ensure that they form valid polygons and multipolygons. Dstl is a great partner of Kaggle. We are thankful for the professionalism and patience they showed in working with us preparing the data and throughout the competition. </p>\n\n<p>Kaggle competitions community had not worked with multi-spectral satellite imagery data before, and it was not an easy task to even set up the environment to get started. I offered a Docker image that worked for me and was thrilled to see Kagglers quickly collaborated and contributed to merge it into Kaggle's Docker for Kernels. Throughout the competition, we continue to see great sharing of tools/methods/insights. We're extremely proud of our community working on a tough problem together. </p>\n\n<p>Instance segmentation (where the algorithm marks exactly where the object is with its outline) is an advanced topic. It remains a tough problem in today's computer vision research. We would like to congratulate everyone for finishing! </p>",
      "votes": null,
      "replies": [
        {
          "id": 167918,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "03/15/2017 20:11:20",
          "content": "<p>It was a very exciting problem. Thank you for hosting it. I hope we will see more similar problems in future.</p>\n\n<p>May I ask why did you split in such a way that there is only 25 images in train? Even 50 will make this problem significantly easier.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 167957,
          "author_name": "wendykan",
          "author_url": "",
          "post_date": "03/15/2017 22:29:28",
          "content": "<p>Interesting question. You may be surprised, but the total number of labeled images is only 57. So we did a split of 25/32. Since we were concerned that having only 32 images in the test set would attract people to hand label, we had to \"salt\" the test set with unlabeled images that are ignored when scoring. </p>\n\n<p>This decision was to ensure we have a successful competition. We also knew that adding these ignored image would slow down people's iteration time in developing the algorithm. It also adds to the total size of the submission files as they do get large very quickly! Eventually, we figured it's more important to keep the competition fair, so we still kept the \"salted\" images in there. </p>\n\n<p>It may seem like a very small hand-labeled dataset but it's also an example of how valuable the ground truth data is. It was a lot of work to get those images labeled correctly. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 167985,
          "author_name": "fppkaggle",
          "author_url": "",
          "post_date": "03/16/2017 00:22:48",
          "content": "<p>Has anybody done a validation of the ground truth like I've seen with other projects e.g. do a secondary validation to get an indication how much of the \"ground truth\" is actually valid / properly classified - example 6070_2_3 where large areas are classified as single areas of trees - IMHO each of those large areas are a mix of shadows, some water, trees e.a. a figure of 80% would be friendly - additionally some class 2 structures provided as training data in other scenes that are merely 1m2 in size again below a shadow / covered partially by another object e.g. trees - I guess it was intended to be a challenge but some of that makes it rather random if you're not looking at results above e.g. 80%+. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 167999,
          "author_name": "jtrotman",
          "author_url": "",
          "post_date": "03/16/2017 01:27:30",
          "content": "<p>Whilst I was looking for the press release about the results, I found that dstl put at least some of the work in creating the ground truth out to tender:</p>\n\n<p><a href=\"http://www.publictenders.net/node/3556044\">http://www.publictenders.net/node/3556044</a></p>\n\n<blockquote>\n  <p>This Requirement contributes to a work package of the Large Scale Data\n  Processing and OSINT (LSDP&amp;O) Project in the Intelligence Countering\n  Adversary Networks (ICAN) Programme by providing the ground truth\n  imagery needed to set a Kaggle Challenge.</p>\n  \n  <p>...</p>\n  \n  <p>Value excluding VAT: 148 821.00 GBP</p>\n</blockquote>\n\n<p>Currently about $180k, which is nearly double the whole prize pool!</p>\n\n<p>I thought, with 450 images, that is £330 per image, or £33 per class per image, not bad value...</p>\n\n<p>It's a bit disappointing to hear it was only 57 images, that's £2.6k per image.</p>\n\n<p>Given all the images fit into 18 5x5 grids, surely the original plan was for all of those to be annotated, and the subcontractor just ran out of time? ;)</p>\n\n<p>It also makes me wonder how they did the annotation - from working on entries for this competition I saw first hand how it would be a lot easier to use a well trained convnet and manually correct it's predictions than to create annotations from scratch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 168006,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "03/16/2017 01:56:25",
          "content": "<p>If you will see another contract to annotate Satelite images especially as these ones from Nigeria - let me know. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 168104,
          "author_name": "kghanareddy",
          "author_url": "",
          "post_date": "03/16/2017 09:01:10",
          "content": "<p>This sounds a little cheating. At least it should have mentioned actual number of test images that would be used for scoring, though it would have lead to more chaos.</p>\n\n<p>Still a better and more fair scoring can be done with a simple strategy.  i.e. take best performing submissions (or ensemble of them) <strong>per class</strong> and use them as truths to re score the submissions on entire data set.\nWhat do you think guys?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 168113,
          "author_name": "ironbar",
          "author_url": "",
          "post_date": "03/16/2017 09:21:19",
          "content": "<p>I think it's a pity making a submission for 450 images and being evaluated only on 32. <br>\nI don't think it's a representative number to measure the goodness of the models.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 167980,
      "author_name": "fppkaggle",
      "author_url": "",
      "post_date": "03/15/2017 23:58:28",
      "content": "<p>This was a really great competition.</p>\n\n<p>Even tough I could  only submit a dummy solution (so I will have to provide my solution at another place - e.g. academic paper e.a. ) I really enjoyed participating in this Kaggle competition.</p>\n\n<p>At the same time it looks like there is a substantial flaw in your scoring system - It can't be that there is a difference in position if the score is the same.</p>\n\n<p>Within the lower scores where you have the same scores repeating it seems most obvious e.g.  with 0.27837 or even more extreme with the score 0.07181 when - with the same score - you are either on 259th place or on 391st place randomly - let's call that fuzzy as the best euphemism  I could come of with.</p>\n\n<p>Overall you are discrediting yourself as a data science platform if such a thing is not remedied.</p>\n\n<p>At the same time - given all the data cleansing e.a. challenges in this competition - as the sponsor I would be strongly disappointed ( and these are - as a UK tax payer - my taxes not yours at work) - getting back solution(s) that is / are doing basically solely a brute force approach and still are coming back below 50%.</p>\n\n<p>Nevertheless, most importantly in this business we all learn from things we do and not from things we only talk about or watch others doing.</p>\n\n<p>With that alone this competition is a success even though it has not immediately forwarded potential results - this can most likely be achieved in a follow up when those who have invested substantially in that competition will be prepared to actually challenge ( look at the number of submissions for some of the competitors as a first indicator to understand this better).</p>\n\n<p>Many thanks again for this opportunity - I certainly learned a lot - now I will have to see where to apply this newly found knowledge - a $2'000 competition certainly will not be the place where I can see people to come forward with their skills or newly shaped experience. Like with other competitions the next stage will have to at least double the price money. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 167988,
      "author_name": "mxdbld",
      "author_url": "",
      "post_date": "03/16/2017 00:38:57",
      "content": "<p>Thank you for this competition. </p>\n\n<p>It was really enjoyable to have graphic predictions and a challenging subject on the technical side. Really liked those hours spent optimizing my models to have a decent execution time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 168079,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "03/16/2017 07:02:21",
      "content": "<p>Thank you for managing interesting competition.<br>\nThere were many new settings for us (charactaristics of dataset, evaluation metrix, submission format and so on).\nI really enjoyed all of them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 168098,
      "author_name": "visoft",
      "author_url": "",
      "post_date": "03/16/2017 08:34:42",
      "content": "<p>Nice competition! </p>\n\n<p>I have a question, my last solution is still processing. I assume that it will never end? (It was submitted a day before the competition ended). Now,  I don't expect to be anything spectacular but it would be nice to know if throwing more compute power helped.</p>",
      "votes": null,
      "replies": [
        {
          "id": 168388,
          "author_name": "cheowthianliang",
          "author_url": "",
          "post_date": "03/17/2017 00:57:13",
          "content": "<p>Similar thing for me. Last submission ended with error. It would be nice to know what score it gets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 168616,
          "author_name": "naagaharikrishna",
          "author_url": "",
          "post_date": "03/17/2017 15:50:05",
          "content": "<p>Have the same problem. my last solution submitted 3 days back it's still processing....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 168742,
          "author_name": "wendykan",
          "author_url": "",
          "post_date": "03/17/2017 23:34:55",
          "content": "<p>@visoft, @Thian, @Naaga, </p>\n\n<p>If you want to know your score, please submit again. I think it most likely timed out but didn't get to throw that error because the site was busy, but you can always try again to see the score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 168865,
          "author_name": "visoft",
          "author_url": "",
          "post_date": "03/18/2017 12:20:20",
          "content": "<p>Thanks! I will try to locate the submision file (I deleted most of the files already).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 168888,
          "author_name": "visoft",
          "author_url": "",
          "post_date": "03/18/2017 14:23:49",
          "content": "<p>It is a 11G (~4g zip) csv file. I don't think it will be evaluated. But i did the submit, just in case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 251956,
      "author_name": "rkaviani",
      "author_url": "",
      "post_date": "12/02/2017 00:26:27",
      "content": "<p>It is very interesting problem.\nWould you please publish the labeled test set (32 images) for whom are going to practice on this problem and evaluate their works?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "167901": "The private leaderboard has now been finalized and published.  Thank you for your patience, and congratulations to the winners!  This was a challenging competition with a unique custom metric and progressive area of study in detecting satellite imagery features.  Our community did not disappoint and clearly rose to the challenge.  Kudos to all of you for your impassioned participation and active sharing across forums.  We look forward to diving into the winners' solutions!",
    "167906": "I would like to give sincere thanks to both Dstl and to all of you that participated. From the very beginning of this project, we knew that this was a challenging problem. The way to measure the performance of the algorithms were difficult, and was an engineering challenge to make it run efficiently given the data size. There were also a lot of challenges to ensure the quality of the hand-labels for ground truth data and to ensure that they form valid polygons and multipolygons. Dstl is a great partner of Kaggle. We are thankful for the professionalism and patience they showed in working with us preparing the data and throughout the competition. \n\nKaggle competitions community had not worked with multi-spectral satellite imagery data before, and it was not an easy task to even set up the environment to get started. I offered a Docker image that worked for me and was thrilled to see Kagglers quickly collaborated and contributed to merge it into Kaggle's Docker for Kernels. Throughout the competition, we continue to see great sharing of tools/methods/insights. We're extremely proud of our community working on a tough problem together. \n\nInstance segmentation (where the algorithm marks exactly where the object is with its outline) is an advanced topic. It remains a tough problem in today's computer vision research. We would like to congratulate everyone for finishing!",
    "167918": "It was a very exciting problem. Thank you for hosting it. I hope we will see more similar problems in future.\n\nMay I ask why did you split in such a way that there is only 25 images in train? Even 50 will make this problem significantly easier.",
    "167957": "Interesting question. You may be surprised, but the total number of labeled images is only 57. So we did a split of 25/32. Since we were concerned that having only 32 images in the test set would attract people to hand label, we had to \"salt\" the test set with unlabeled images that are ignored when scoring. \n\nThis decision was to ensure we have a successful competition. We also knew that adding these ignored image would slow down people's iteration time in developing the algorithm. It also adds to the total size of the submission files as they do get large very quickly! Eventually, we figured it's more important to keep the competition fair, so we still kept the \"salted\" images in there. \n\nIt may seem like a very small hand-labeled dataset but it's also an example of how valuable the ground truth data is. It was a lot of work to get those images labeled correctly.",
    "167980": "This was a really great competition.\n\nEven tough I could  only submit a dummy solution (so I will have to provide my solution at another place - e.g. academic paper e.a. ) I really enjoyed participating in this Kaggle competition.\n\nAt the same time it looks like there is a substantial flaw in your scoring system - It can't be that there is a difference in position if the score is the same.\n\nWithin the lower scores where you have the same scores repeating it seems most obvious e.g.  with 0.27837 or even more extreme with the score 0.07181 when - with the same score - you are either on 259th place or on 391st place randomly - let's call that fuzzy as the best euphemism  I could come of with.\n\nOverall you are discrediting yourself as a data science platform if such a thing is not remedied.\n\nAt the same time - given all the data cleansing e.a. challenges in this competition - as the sponsor I would be strongly disappointed ( and these are - as a UK tax payer - my taxes not yours at work) - getting back solution(s) that is / are doing basically solely a brute force approach and still are coming back below 50%.\n\nNevertheless, most importantly in this business we all learn from things we do and not from things we only talk about or watch others doing.\n\nWith that alone this competition is a success even though it has not immediately forwarded potential results - this can most likely be achieved in a follow up when those who have invested substantially in that competition will be prepared to actually challenge ( look at the number of submissions for some of the competitors as a first indicator to understand this better).\n\nMany thanks again for this opportunity - I certainly learned a lot - now I will have to see where to apply this newly found knowledge - a $2'000 competition certainly will not be the place where I can see people to come forward with their skills or newly shaped experience. Like with other competitions the next stage will have to at least double the price money.",
    "167985": "Has anybody done a validation of the ground truth like I've seen with other projects e.g. do a secondary validation to get an indication how much of the \"ground truth\" is actually valid / properly classified - example 6070_2_3 where large areas are classified as single areas of trees - IMHO each of those large areas are a mix of shadows, some water, trees e.a. a figure of 80% would be friendly - additionally some class 2 structures provided as training data in other scenes that are merely 1m2 in size again below a shadow / covered partially by another object e.g. trees - I guess it was intended to be a challenge but some of that makes it rather random if you're not looking at results above e.g. 80%+.",
    "167988": "Thank you for this competition. \n\nIt was really enjoyable to have graphic predictions and a challenging subject on the technical side. Really liked those hours spent optimizing my models to have a decent execution time.",
    "167999": "Whilst I was looking for the press release about the results, I found that dstl put at least some of the work in creating the ground truth out to tender:\n\nhttp://www.publictenders.net/node/3556044\n\n> This Requirement contributes to a work package of the Large Scale Data\n> Processing and OSINT (LSDP&O) Project in the Intelligence Countering\n> Adversary Networks (ICAN) Programme by providing the ground truth\n> imagery needed to set a Kaggle Challenge.\n\n> ...\n\n> Value excluding VAT: 148 821.00 GBP\n\nCurrently about $180k, which is nearly double the whole prize pool!\n\nI thought, with 450 images, that is £330 per image, or £33 per class per image, not bad value...\n\nIt's a bit disappointing to hear it was only 57 images, that's £2.6k per image.\n\nGiven all the images fit into 18 5x5 grids, surely the original plan was for all of those to be annotated, and the subcontractor just ran out of time? ;)\n\nIt also makes me wonder how they did the annotation - from working on entries for this competition I saw first hand how it would be a lot easier to use a well trained convnet and manually correct it's predictions than to create annotations from scratch.",
    "168006": "If you will see another contract to annotate Satelite images especially as these ones from Nigeria - let me know. :)",
    "168079": "Thank you for managing interesting competition.<br>\nThere were many new settings for us (charactaristics of dataset, evaluation metrix, submission format and so on).\nI really enjoyed all of them.",
    "168098": "Nice competition! \n\nI have a question, my last solution is still processing. I assume that it will never end? (It was submitted a day before the competition ended). Now,  I don't expect to be anything spectacular but it would be nice to know if throwing more compute power helped.",
    "168104": "This sounds a little cheating. At least it should have mentioned actual number of test images that would be used for scoring, though it would have lead to more chaos.\n\nStill a better and more fair scoring can be done with a simple strategy.  i.e. take best performing submissions (or ensemble of them) **per class** and use them as truths to re score the submissions on entire data set.\nWhat do you think guys?",
    "168113": "I think it's a pity making a submission for 450 images and being evaluated only on 32.  \nI don't think it's a representative number to measure the goodness of the models.",
    "168388": "Similar thing for me. Last submission ended with error. It would be nice to know what score it gets.",
    "168616": "Have the same problem. my last solution submitted 3 days back it's still processing....",
    "168742": "visoft, @Thian, @Naaga, \n\nIf you want to know your score, please submit again. I think it most likely timed out but didn't get to throw that error because the site was busy, but you can always try again to see the score.",
    "168865": "Thanks! I will try to locate the submision file (I deleted most of the files already).",
    "168888": "It is a 11G (~4g zip) csv file. I don't think it will be evaluated. But i did the submit, just in case.",
    "251956": "It is very interesting problem.\nWould you please publish the labeled test set (32 images) for whom are going to practice on this problem and evaluate their works?"
  },
  "source": "meta"
}