{
  "id": 64355,
  "title": "Severe leak",
  "url": "/competitions/airbus-ship-detection/discussion/64355",
  "author_name": "Andrés Miguel Torrubia Sáez",
  "post_date": "2018-08-28T14:18:42.570000",
  "votes": 134,
  "comment_count": 47,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>As you may have seen we got 0.98+ LB with just one sub. Although we had a lot of good ideas for the challenge my partner Pavel found a leak which renders this competition almost useless.</p>\n\n<p>The images in test are just shifted versions of images in train. The problem also happens in train, it looks like airbus had bigger images and then the 768x768 are random crops of the bigger images; but it looks they didn't check whether there were any overlaps.</p>\n\n<p>e.g. train/dd38da47f.jpg and train/02776139a.jpg</p>\n\n<p>How to find:</p>\n\n<ul>\n<li>Run nearest neighbors on all images </li>\n<li>For each image take N closest neighbors and find where it overlaps</li>\n</ul>\n\n<p>You can quickly recover almost (all?) of test images.</p>",
  "messages": [
    {
      "id": 377037,
      "postDate": "2018-08-28T14:18:42.570Z",
      "content": "<p>Hi all,</p>\n\n<p>As you may have seen we got 0.98+ LB with just one sub. Although we had a lot of good ideas for the challenge my partner Pavel found a leak which renders this competition almost useless.</p>\n\n<p>The images in test are just shifted versions of images in train. The problem also happens in train, it looks like airbus had bigger images and then the 768x768 are random crops of the bigger images; but it looks they didn't check whether there were any overlaps.</p>\n\n<p>e.g. train/dd38da47f.jpg and train/02776139a.jpg</p>\n\n<p>How to find:</p>\n\n<ul>\n<li>Run nearest neighbors on all images </li>\n<li>For each image take N closest neighbors and find where it overlaps</li>\n</ul>\n\n<p>You can quickly recover almost (all?) of test images.</p>",
      "rawMarkdown": "Hi all,\n\nAs you may have seen we got 0.98+ LB with just one sub. Although we had a lot of good ideas for the challenge my partner Pavel found a leak which renders this competition almost useless.\n\nThe images in test are just shifted versions of images in train. The problem also happens in train, it looks like airbus had bigger images and then the 768x768 are random crops of the bigger images; but it looks they didn't check whether there were any overlaps.\n\ne.g. train/dd38da47f.jpg and train/02776139a.jpg\n\nHow to find:\n\n - Run nearest neighbors on all images \n - For each image take N closest neighbors and find where it overlaps\n\nYou can quickly recover almost (all?) of test images.\n\n",
      "votes": 134
    },
    {
      "id": 377084,
      "postDate": "2018-08-28T15:03:10.293Z",
      "content": "<p>Well, I hope they use a slightly more careful technique to check data while building aircrafts.</p>",
      "rawMarkdown": "Well, I hope they use a slightly more careful technique to check data while building aircrafts.",
      "votes": 21
    },
    {
      "id": 377068,
      "postDate": "2018-08-28T14:44:18.183Z",
      "content": "<p>Pictures shifted relative to each other by 256 pixels. It looks like organizers slice large pictures with a step of 256 and randomly divided them into train and test sets. Cutting all the images 256x256 they can be found easier because they match almost exactly (except for compression artifacts).</p>",
      "rawMarkdown": "Pictures shifted relative to each other by 256 pixels. It looks like organizers slice large pictures with a step of 256 and randomly divided them into train and test sets. Cutting all the images 256x256 they can be found easier because they match almost exactly (except for compression artifacts).",
      "votes": 21
    },
    {
      "id": 377059,
      "postDate": "2018-08-28T14:30:04.663Z",
      "content": "<p>I heard Kaggle charges competition sponsors with hefty sums to check everything and so on. In particular, it was one of the reasons they refused to award points on competitions where they didn't get their part (because, you know, it was \"unverified\" and possibly low quality data). So, given rate of leaks (and such freaking outrageously simple ones) on \"checked\" data I think they are just deceiving their clients claiming that they check anything and in reality just publishing the data as is. </p>",
      "rawMarkdown": "I heard Kaggle charges competition sponsors with hefty sums to check everything and so on. In particular, it was one of the reasons they refused to award points on competitions where they didn't get their part (because, you know, it was \"unverified\" and possibly low quality data). So, given rate of leaks (and such freaking outrageously simple ones) on \"checked\" data I think they are just deceiving their clients claiming that they check anything and in reality just publishing the data as is. ",
      "votes": 27,
      "replies": [
        {
          "id": 377127,
          "postDate": "2018-08-28T15:57:51.500Z",
          "content": "<p>Well, the leaks become more and more outrageous)\nFeaturing the SAME images is too much even for Kaggle =)</p>\n\n<p>Maybe if the crowd smells blood, it will force them to make some <em>staff related</em> decisions?</p>",
          "rawMarkdown": "Well, the leaks become more and more outrageous)\nFeaturing the SAME images is too much even for Kaggle =)\n\nMaybe if the crowd smells blood, it will force them to make some *staff related* decisions?",
          "votes": 7
        },
        {
          "id": 377134,
          "postDate": "2018-08-28T16:14:56.780Z",
          "content": "<p>Hi Sergey,</p>\n\n<p>I did not make this competition, but I have a lot of experience on both sides of competitions. I was a top competitor prior to joining Kaggle and have probably created more ML competitions than anybody on the planet. I don't say this to convince you I'm smart (the jury is very out on that front!), but rather to make the case that I am experienced. I have suffered through the throes of leakage many times, each time vowing that this will never happen again! But, it does, and we end up in a forum thread like this, with hard-working Kagglers wondering how a self-respecting data scientist could let this through.</p>\n\n<p>Finding leakage is surprisingly hard. Leakage looks trivial through the 20/20 lens of hindsight, but it's an infinite minefield when walking forward. Leakage changes shapes. Leakage hides in black boxes that we--we who did not create the data--often don't have the luxury of seeing. Leakage evades attempts to systemically detect it. You check exact duplicates? You missed fuzzy matches. You checked fuzzy matches? You missed reflected matches. You check fuzzy reflected matches? You missed scaled matches. You silently prevented 4 huge sources of leakage that Kagglers will never know they were spared from? Oops, there was a 5th. If you enumerated a check against every type of leakage that has ever plagued a dataset on Kaggle, you'd still likely miss the next time around. There is simply no substitute (that we've found) for 500 people spending a month hammering on a dataset.</p>\n\n<p>Kaggle and host data scientists are humans. We have the same time crunches and imperfect working demands that any data scientist role involves. These aren't excuses, but I do hope this reply gives you some understanding of the complexity that leakage presents on this side of the fence.</p>\n\n<p>Finally, no matter how many times leakage gets the better of us, we do have a track record of doing what we can to make things as right as possible, and you, the community, have an equally admirable track record of persisting through these course corrections. We can't ask for more than that, and continue to do our best to make quality problems.</p>",
          "rawMarkdown": "Hi Sergey,\n\nI did not make this competition, but I have a lot of experience on both sides of competitions. I was a top competitor prior to joining Kaggle and have probably created more ML competitions than anybody on the planet. I don't say this to convince you I'm smart (the jury is very out on that front!), but rather to make the case that I am experienced. I have suffered through the throes of leakage many times, each time vowing that this will never happen again! But, it does, and we end up in a forum thread like this, with hard-working Kagglers wondering how a self-respecting data scientist could let this through.\n\nFinding leakage is surprisingly hard. Leakage looks trivial through the 20/20 lens of hindsight, but it's an infinite minefield when walking forward. Leakage changes shapes. Leakage hides in black boxes that we--we who did not create the data--often don't have the luxury of seeing. Leakage evades attempts to systemically detect it. You check exact duplicates? You missed fuzzy matches. You checked fuzzy matches? You missed reflected matches. You check fuzzy reflected matches? You missed scaled matches. You silently prevented 4 huge sources of leakage that Kagglers will never know they were spared from? Oops, there was a 5th. If you enumerated a check against every type of leakage that has ever plagued a dataset on Kaggle, you'd still likely miss the next time around. There is simply no substitute (that we've found) for 500 people spending a month hammering on a dataset.\n\nKaggle and host data scientists are humans. We have the same time crunches and imperfect working demands that any data scientist role involves. These aren't excuses, but I do hope this reply gives you some understanding of the complexity that leakage presents on this side of the fence.\n\nFinally, no matter how many times leakage gets the better of us, we do have a track record of doing what we can to make things as right as possible, and you, the community, have an equally admirable track record of persisting through these course corrections. We can't ask for more than that, and continue to do our best to make quality problems.",
          "votes": 46
        },
        {
          "id": 377136,
          "postDate": "2018-08-28T16:19:22.140Z",
          "content": "<p>What is your next move? Are you going to restart the competition?</p>",
          "rawMarkdown": "What is your next move? Are you going to restart the competition?",
          "votes": 1
        },
        {
          "id": 377141,
          "postDate": "2018-08-28T16:27:29.760Z",
          "content": "<p>You can of course incentivize leak sharing e.g. during the first month or so.</p>\n\n<p>As you know some experienced Kagglers are now recommending holding off until the competition is running for a while to avoid wasting time training ML models. So why not make this an integral part of the process? </p>\n\n<p>EDAs, data exploration, leak hunting... as phase #1.</p>",
          "rawMarkdown": "You can of course incentivize leak sharing e.g. during the first month or so.\n\nAs you know some experienced Kagglers are now recommending holding off until the competition is running for a while to avoid wasting time training ML models. So why not make this an integral part of the process? \n\nEDAs, data exploration, leak hunting... as phase #1.",
          "votes": 16
        },
        {
          "id": 377148,
          "postDate": "2018-08-28T16:39:31.980Z",
          "content": "<p>You can ask for the data preparation algorithm, and do not check the data itself.</p>",
          "rawMarkdown": "You can ask for the data preparation algorithm, and do not check the data itself.",
          "votes": 5
        },
        {
          "id": 377156,
          "postDate": "2018-08-28T16:50:15.843Z",
          "content": "<p>I appreciate all the effort that went into checking this data set on kaggle's side. I'm mostly amazed by the naiveté of competition sponsors, like Airbus in this case or Santander in the recently finished competition. In both competitions the sponsors must have known about the presence of the leak and just hoped that the kaggle community wouldn't notice their futile attempts to obfuscate the data. It's an interesting vote of no-confidence in the community's abilities from the sponsor side.</p>",
          "rawMarkdown": "I appreciate all the effort that went into checking this data set on kaggle's side. I'm mostly amazed by the naiveté of competition sponsors, like Airbus in this case or Santander in the recently finished competition. In both competitions the sponsors must have known about the presence of the leak and just hoped that the kaggle community wouldn't notice their futile attempts to obfuscate the data. It's an interesting vote of no-confidence in the community's abilities from the sponsor side.",
          "votes": 3
        },
        {
          "id": 377163,
          "postDate": "2018-08-28T17:00:38.150Z",
          "content": "<p>Hi, William. I'm sorry if I was too emotional and too harsh on your efforts (as a team). I understand a level of efforts required to prepare a good dataset and competition and limited resources are always a thing at small startup, so many things were forgiven when Kaggle was a small team of like-minded guys. However, this is no longer the case and you are a Google now. I bet no such leaks occurs in production in your own company. It means there is a possibility to have some process to eliminate or at least greatly reduce amount of leaks. </p>\n\n<p>This process should be in the same spirit as usual quality assurance in software development -- yes, there are bugs in software but we now have all kind of means to eliminate them and validate the quality of shipped code. And this process (should) starts from the very beginning not be something to hand-off to QA guys. Like not only trying some checklist but also asking how the data was created in the first place, auditing that process and making sure it works fine. Also, I hope you using all kind of (semi-)automated \"regression\"  tests from previous competitions. Because, you know, it wasn't 500 people looking on a data for a month to spot this leak. It was couple of guys who used some techniques known from previous competitions (like knn clustering to find similar images, sift matching and so on, check this video for some examples (in russian with eng slides and auto translated subs <a href=\"https://www.youtube.com/watch?v=TwIi7zoiiPs\">https://www.youtube.com/watch?v=TwIi7zoiiPs</a>)) and this is really astounding to have full test being just copy of the train. I have not started to participate yet in this competition but I can really feel the pain of those who spent a lot of time (and energy) and found their effort completely useless. </p>\n\n<p>As a second thought, maybe, if Kaggle couldn't manage to check the data for leaks with available resources then you should just host raw data as is as \"leak finding and checking pre-competition stage\" and assign your part of $$ as prize for this?  Everybody will be happy: there will be no leaks in the main stage, anyone who find leak will get their share and there will be no embarrassment whatsoever. </p>",
          "rawMarkdown": "Hi, William. I'm sorry if I was too emotional and too harsh on your efforts (as a team). I understand a level of efforts required to prepare a good dataset and competition and limited resources are always a thing at small startup, so many things were forgiven when Kaggle was a small team of like-minded guys. However, this is no longer the case and you are a Google now. I bet no such leaks occurs in production in your own company. It means there is a possibility to have some process to eliminate or at least greatly reduce amount of leaks. \n\nThis process should be in the same spirit as usual quality assurance in software development -- yes, there are bugs in software but we now have all kind of means to eliminate them and validate the quality of shipped code. And this process (should) starts from the very beginning not be something to hand-off to QA guys. Like not only trying some checklist but also asking how the data was created in the first place, auditing that process and making sure it works fine. Also, I hope you using all kind of (semi-)automated \"regression\"  tests from previous competitions. Because, you know, it wasn't 500 people looking on a data for a month to spot this leak. It was couple of guys who used some techniques known from previous competitions (like knn clustering to find similar images, sift matching and so on, check this video for some examples (in russian with eng slides and auto translated subs https://www.youtube.com/watch?v=TwIi7zoiiPs)) and this is really astounding to have full test being just copy of the train. I have not started to participate yet in this competition but I can really feel the pain of those who spent a lot of time (and energy) and found their effort completely useless. \n\nAs a second thought, maybe, if Kaggle couldn't manage to check the data for leaks with available resources then you should just host raw data as is as \"leak finding and checking pre-competition stage\" and assign your part of $$ as prize for this?  Everybody will be happy: there will be no leaks in the main stage, anyone who find leak will get their share and there will be no embarrassment whatsoever. ",
          "votes": 22
        },
        {
          "id": 377166,
          "postDate": "2018-08-28T17:05:01.983Z",
          "content": "<p>Andres Torrubia, I'm starting to believe in synchronicity =) While I was writing my angry reply you managed to say almost the same thing I realized in the last minute or so :)</p>",
          "rawMarkdown": "Andres Torrubia, I'm starting to believe in synchronicity =) While I was writing my angry reply you managed to say almost the same thing I realized in the last minute or so :)",
          "votes": 5
        },
        {
          "id": 377309,
          "postDate": "2018-08-28T23:55:16.350Z",
          "content": "<p>Think it's a great proposal for a solution.</p>",
          "rawMarkdown": "Think it's a great proposal for a solution."
        }
      ]
    },
    {
      "id": 377593,
      "postDate": "2018-08-29T12:19:41.047Z",
      "content": "<p>At least we don't need to do image translation as part of data augmentation, given it is already done ;)</p>",
      "rawMarkdown": "At least we don't need to do image translation as part of data augmentation, given it is already done ;)",
      "votes": 15,
      "replies": [
        {
          "id": 379020,
          "postDate": "2018-08-30T18:53:00.433Z",
          "content": "<p>Lol!! Soo true.</p>",
          "rawMarkdown": "Lol!! Soo true.",
          "votes": 1
        }
      ]
    },
    {
      "id": 377376,
      "postDate": "2018-08-29T03:51:56.957Z",
      "content": "<p>Dear Kaggle,</p>\n\n<p>I am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.</p>\n\n<p>Also I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.</p>\n\n<p>So, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the <strong>ML/DS competition platforms:</strong></p>\n\n<ul>\n<li>Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;</li>\n<li>DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;</li>\n<li>CrowdAI - give us your code, write tests for us, receive nothing in exchange!</li>\n<li>Codalab - it is just ... weird;</li>\n<li>Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;</li>\n</ul>\n\n<p><strong>Most cringy cases in my memory:</strong></p>\n\n<ul>\n<li>Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;</li>\n<li>DSB 2018 test set;</li>\n<li>Santander;</li>\n<li>Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;</li>\n<li>This leak so far takes the cake;</li>\n<li>(I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )</li>\n</ul>\n\n<p>So, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.</p>\n\n<p><strong>- Dataset preparation:</strong></p>\n\n<ul>\n<li><p>Build a FIREWALL between train / test / delayed test sets on the basis of FILES;</p>\n\n<ul><li>Do not try to fool the audience by pretending that you have more data than you have;</li>\n<li>Let people decide how they want to slice their data - provide the data as raw as possible;</li>\n<li>A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";</li></ul></li>\n</ul>\n\n<p><strong>- Stratifying you dataset:</strong></p>\n\n<ul>\n<li>For each competition you have to find out what drives the predictions and the challenge;</li>\n<li>For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;</li>\n<li>Actually do stratify the train / test / delayed test sets;</li>\n<li>Yest it is difficult. But why not invest all the money into making datasets palatable?;</li>\n</ul>\n\n<p><strong>- Making data SCIENCE results usable and reproducible</strong></p>\n\n<ul>\n<li>Put tough calculation limits, so that people would actually THINK before stacking 9000 models;</li>\n<li>Always forbid any external annotation efforts;</li>\n<li>Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n<ul><li>Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;</li>\n<li>And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;</li>\n<li>And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;</li>\n<li>This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;</li></ul></li>\n</ul>\n\n<p>That is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.</p>",
      "rawMarkdown": "Dear Kaggle,\n\nI am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.\n\nAlso I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.\n\nSo, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the **ML/DS competition platforms:**\n\n - Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;\n - DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;\n - CrowdAI - give us your code, write tests for us, receive nothing in exchange!\n - Codalab - it is just ... weird;\n - Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;\n\n**Most cringy cases in my memory:**\n\n - Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;\n - DSB 2018 test set;\n - Santander;\n - Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;\n - This leak so far takes the cake;\n - (I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )\n\n\nSo, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.\n\n**- Dataset preparation:**\n  \n\n - Build a FIREWALL between train / test / delayed test sets on the basis of FILES;\n\n \n  - Do not try to fool the audience by pretending that you have more data than you have;\n  - Let people decide how they want to slice their data - provide the data as raw as possible;\n  - A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";\n\n**- Stratifying you dataset:**\n\n - For each competition you have to find out what drives the predictions and the challenge;\n - For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;\n - Actually do stratify the train / test / delayed test sets;\n - Yest it is difficult. But why not invest all the money into making datasets palatable?;\n\n**- Making data SCIENCE results usable and reproducible**\n\n- Put tough calculation limits, so that people would actually THINK before stacking 9000 models;\n- Always forbid any external annotation efforts;\n- Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n  - Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;\n  - And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;\n  - And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;\n  - This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;\n\nThat is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.",
      "votes": 16,
      "replies": [
        {
          "id": 377391,
          "postDate": "2018-08-29T04:48:31.030Z",
          "content": "<p>Decided to move to a separate thread\n<a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/64393\">https://www.kaggle.com/c/airbus-ship-detection/discussion/64393</a></p>",
          "rawMarkdown": "Decided to move to a separate thread\nhttps://www.kaggle.com/c/airbus-ship-detection/discussion/64393\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 377126,
      "postDate": "2018-08-28T15:57:17.813Z",
      "content": "<p>Hi to all,</p>\n\n<p>Andreas, thank you for sharing this. This is of course not intended. \nWe are looking into that now and will get back to all participants really soon.</p>\n\n<p>Jeff.</p>",
      "rawMarkdown": "Hi to all,\n\nAndreas, thank you for sharing this. This is of course not intended. \nWe are looking into that now and will get back to all participants really soon.\n\nJeff.",
      "votes": 9,
      "replies": [
        {
          "id": 377133,
          "postDate": "2018-08-28T16:06:18.167Z",
          "content": "<p>You're welcome.</p>\n\n<p>BTW - It's Andres, not Andreas.</p>",
          "rawMarkdown": "You're welcome.\n\nBTW - It's Andres, not Andreas.",
          "votes": 9
        },
        {
          "id": 377249,
          "postDate": "2018-08-28T20:26:36.387Z",
          "content": "<p>Andres. Sorry for the misspelling. </p>\n\n<p>Kaggle will communicate soon on the competition status. \nAirbus will have to deliver a new validation dataset. We are currently looking into that.</p>\n\n<p>Kind regards,\nJeff.</p>",
          "rawMarkdown": "Andres. Sorry for the misspelling. \n\nKaggle will communicate soon on the competition status. \nAirbus will have to deliver a new validation dataset. We are currently looking into that.\n\nKind regards,\nJeff.\n\n",
          "votes": 7
        },
        {
          "id": 379453,
          "postDate": "2018-08-31T11:30:26.773Z",
          "content": "<p>Hi Jeff,</p>\n\n<p>Since data leak event occurred, is there a possibility to move the competition end date? I assume that your team will provide new set of train and test dataset for this competition? Am I correct? </p>",
          "rawMarkdown": "Hi Jeff,\n\nSince data leak event occurred, is there a possibility to move the competition end date? I assume that your team will provide new set of train and test dataset for this competition? Am I correct? ",
          "votes": 4
        },
        {
          "id": 381620,
          "postDate": "2018-09-04T21:39:14.240Z",
          "content": "<p>May be a good time to look at leakage through file creation/last edit timestamps. At least in the train set this is an informative feature. If not on purpose, and random split, this may also \"generalize\" to the test set.</p>",
          "rawMarkdown": "May be a good time to look at leakage through file creation/last edit timestamps. At least in the train set this is an informative feature. If not on purpose, and random split, this may also \"generalize\" to the test set."
        }
      ]
    },
    {
      "id": 377521,
      "postDate": "2018-08-29T10:04:35.840Z",
      "content": "<p>Dear Kaggle,</p>\n\n<p>I want to propose to give the team of Andres Torrubia and Pavel Gonchar 25k.\nI can propose this without any conflict of interest since I don't know them.</p>\n\n<p>This way Kaggle sets a precedent on rewarding finding leaks and reporting them as early as possible.\nIt also gives Kaggle an incentive to be even more careful next time. \nIt could even give the community some reassurance seeing that this sort of mistakes can cost Kaggle, apart from extra time and effort, also some money. </p>\n\n<p>I talk about Kaggle here as the representative of the sponsor, who actually pays the money is something Kaggle and Airbus can decide between them. </p>\n\n<p>Edit: removed a remark that was outdated at the time of writing</p>",
      "rawMarkdown": "Dear Kaggle,\n\nI want to propose to give the team of Andres Torrubia and Pavel Gonchar 25k.\nI can propose this without any conflict of interest since I don't know them.\n\nThis way Kaggle sets a precedent on rewarding finding leaks and reporting them as early as possible.\nIt also gives Kaggle an incentive to be even more careful next time. \nIt could even give the community some reassurance seeing that this sort of mistakes can cost Kaggle, apart from extra time and effort, also some money. \n\nI talk about Kaggle here as the representative of the sponsor, who actually pays the money is something Kaggle and Airbus can decide between them. \n\nEdit: removed a remark that was outdated at the time of writing\n\n\n\n",
      "votes": 8,
      "replies": [
        {
          "id": 377533,
          "postDate": "2018-08-29T10:31:03.183Z",
          "content": "<p>There is indeed a lot of value in reporting the leak, more than being frustrated by not winning the competition I would have found very depressing (and I imagine Airbus even more so) of not knowing what SOTA was  for that type of tasks should we had known after the competition ended. \nThere should indeed be a way to reward this, as many others knew and were just waiting to see where it went, as it is allowed per competition rule (not judging here a game is a game).</p>",
          "rawMarkdown": "There is indeed a lot of value in reporting the leak, more than being frustrated by not winning the competition I would have found very depressing (and I imagine Airbus even more so) of not knowing what SOTA was  for that type of tasks should we had known after the competition ended. \nThere should indeed be a way to reward this, as many others knew and were just waiting to see where it went, as it is allowed per competition rule (not judging here a game is a game).",
          "votes": 2
        },
        {
          "id": 377548,
          "postDate": "2018-08-29T10:45:28.887Z",
          "content": "<p>I do not have spare US$25k, but I would chip in US$100 just for the sake of it</p>",
          "rawMarkdown": "I do not have spare US$25k, but I would chip in US$100 just for the sake of it",
          "votes": 2
        },
        {
          "id": 377556,
          "postDate": "2018-08-29T10:55:15.823Z",
          "content": "<blockquote>\n  <p><strong>Alexis Letulier wrote</strong></p>\n  \n  <blockquote>\n    <p>There should indeed be a way to reward this, as many others knew and were just waiting to see where it went, as it is allowed per competition rule (not judging here a game is a game).</p>\n  </blockquote>\n</blockquote>\n\n<p>If Kaggle rewards early reporting of leaks it can save a lot of Kagglers a lot of work and some Kagglers a lot of uncertainty.</p>",
          "rawMarkdown": "\n&gt; **Alexis Letulier wrote**\n&gt; \n&gt; &gt; \n&gt; There should indeed be a way to reward this, as many others knew and were just waiting to see where it went, as it is allowed per competition rule (not judging here a game is a game).\n\nIf Kaggle rewards early reporting of leaks it can save a lot of Kagglers a lot of work and some Kagglers a lot of uncertainty.",
          "votes": 3
        },
        {
          "id": 377576,
          "postDate": "2018-08-29T11:34:54.853Z",
          "content": "<p>$25k would be awesome 😂but so will be this:</p>\n\n<p><img src=\"https://i.imgur.com/Dx15bVo.png\" alt=\"Leak TShirt\"></p>",
          "rawMarkdown": "$25k would be awesome 😂but so will be this:\n\n![Leak TShirt][1]\n\n\n  [1]: https://i.imgur.com/Dx15bVo.png",
          "votes": 22
        }
      ]
    },
    {
      "id": 377073,
      "postDate": "2018-08-28T14:47:37.833Z",
      "content": "<p>I suppose the only way for them to redeem themselves is:</p>\n\n<ul>\n<li>Extend the duration of the competition; </li>\n<li>Change the dataset; </li>\n<li>Do   something about the prizes to account for double GPU time in  case of  dataset changes;</li>\n</ul>",
      "rawMarkdown": "I suppose the only way for them to redeem themselves is:\n\n - Extend the duration of the competition; \n - Change the dataset; \n - Do   something about the prizes to account for double GPU time in  case of  dataset changes;\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 379455,
          "postDate": "2018-08-31T11:33:48.180Z",
          "content": "<p>I do agree on your proposals! </p>",
          "rawMarkdown": "I do agree on your proposals! ",
          "votes": 4
        }
      ]
    },
    {
      "id": 377122,
      "postDate": "2018-08-28T15:53:05.250Z",
      "content": "<p>Oh, shoot....\nI will give up on this competition if there is no solution to this. \nWhat a waste of my GPU time.</p>",
      "rawMarkdown": "Oh, shoot....\nI will give up on this competition if there is no solution to this. \nWhat a waste of my GPU time.",
      "votes": 3,
      "replies": [
        {
          "id": 377631,
          "postDate": "2018-08-29T13:41:53.833Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 377102,
      "postDate": "2018-08-28T15:22:49.433Z",
      "content": "<p>I was thinking that 0.98+ may not be possible even with hand labeling:) Thanks for sharing.</p>\n\n<p>This competition can only be saved if it is converted to a kernel competition with strict time and memory constraints. Even this may not be enough, because now we know that it is better to overfit on train images. So they may replace the test set with the one that they saved for Algorithm Speed Challenge.</p>",
      "rawMarkdown": "I was thinking that 0.98+ may not be possible even with hand labeling:) Thanks for sharing.\n\nThis competition can only be saved if it is converted to a kernel competition with strict time and memory constraints. Even this may not be enough, because now we know that it is better to overfit on train images. So they may replace the test set with the one that they saved for Algorithm Speed Challenge.",
      "votes": 4
    },
    {
      "id": 377078,
      "postDate": "2018-08-28T14:52:37.647Z",
      "content": "<p>What a disappointment, at least I got a model that detects quite well ships (even from non Airbus images), that I will be able to use for my job...</p>",
      "rawMarkdown": "What a disappointment, at least I got a model that detects quite well ships (even from non Airbus images), that I will be able to use for my job...",
      "votes": 3,
      "replies": [
        {
          "id": 377822,
          "postDate": "2018-08-29T19:50:02.860Z",
          "content": "<p>You had a strong model very quickly! Did you already have one lying around from your work that you used for the competition? Very impressive score.</p>",
          "rawMarkdown": "You had a strong model very quickly! Did you already have one lying around from your work that you used for the competition? Very impressive score.",
          "votes": 1
        },
        {
          "id": 377825,
          "postDate": "2018-08-29T19:57:18.010Z",
          "content": "<p>No, its the other way around, I participated in order to have a model I could use in my work and was genuinely surprised to lead. We´ll see with the new dataset if the score holds or if I was only overfitting. </p>",
          "rawMarkdown": "No, its the other way around, I participated in order to have a model I could use in my work and was genuinely surprised to lead. We´ll see with the new dataset if the score holds or if I was only overfitting. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 377067,
      "postDate": "2018-08-28T14:42:30.237Z",
      "content": "<p>Amazing. I was able to get into top-7 with only one commit and now this. \nI feel appalled. My righteous indignation has no limits.\nAnd you know what - I did not even check for such bs.\n People are correct that Kaggle's admins do not know what the hell the are doing.</p>",
      "rawMarkdown": "Amazing. I was able to get into top-7 with only one commit and now this. \nI feel appalled. My righteous indignation has no limits.\nAnd you know what - I did not even check for such bs.\n People are correct that Kaggle's admins do not know what the hell the are doing.",
      "votes": 1
    },
    {
      "id": 378370,
      "postDate": "2018-08-30T06:51:46.990Z",
      "content": "<p>I am happy I did not spent time in this competition yet. Planned to use the stuff I learned from tgs-salt-competition at the end of this competition. (As might some others) Let's see what organizers will do...</p>\n\n<p>I also want to thank Andres and his team for their honesty. </p>",
      "rawMarkdown": "I am happy I did not spent time in this competition yet. Planned to use the stuff I learned from tgs-salt-competition at the end of this competition. (As might some others) Let's see what organizers will do...\n\nI also want to thank Andres and his team for their honesty. ",
      "votes": 2
    },
    {
      "id": 377081,
      "postDate": "2018-08-28T14:55:11.563Z",
      "content": "<p>So <strong>all</strong> images from test are from train? Surely you can't be serious!</p>",
      "rawMarkdown": "So **all** images from test are from train? Surely you can't be serious!",
      "votes": 2,
      "replies": [
        {
          "id": 377087,
          "postDate": "2018-08-28T15:03:26.077Z",
          "content": "<p>Obviously, there is a room for improvement for several thousandth as score isn't a perfect 1 =)</p>",
          "rawMarkdown": "Obviously, there is a room for improvement for several thousandth as score isn't a perfect 1 =)",
          "votes": 7
        }
      ]
    },
    {
      "id": 377094,
      "postDate": "2018-08-28T15:19:07.180Z",
      "content": "<p>So is it possible to restart competition without leak? But I guess there is no possibility for that, because we have almost 100% of labels from whole dataset.</p>",
      "rawMarkdown": "So is it possible to restart competition without leak? But I guess there is no possibility for that, because we have almost 100% of labels from whole dataset.",
      "votes": -1,
      "replies": [
        {
          "id": 377099,
          "postDate": "2018-08-28T15:22:12.790Z",
          "content": "<p>To do that you will need a completely new test set, and this is not something that can be created quickly nor cheaply.</p>",
          "rawMarkdown": "To do that you will need a completely new test set, and this is not something that can be created quickly nor cheaply.",
          "votes": 3
        },
        {
          "id": 377114,
          "postDate": "2018-08-28T15:37:38.603Z",
          "content": "<p>As for creating datasets ... if you just go to Google Earth to SF / HK / Singapore etc - you will be amazed how easy actually collecting such a dataset is.</p>\n\n<p>Labeling is a different thing ofc.</p>",
          "rawMarkdown": "As for creating datasets ... if you just go to Google Earth to SF / HK / Singapore etc - you will be amazed how easy actually collecting such a dataset is.\n\nLabeling is a different thing ofc.\n",
          "votes": 1
        },
        {
          "id": 377117,
          "postDate": "2018-08-28T15:42:01.643Z",
          "content": "<p>I was of course discussing about labeling.</p>",
          "rawMarkdown": "I was of course discussing about labeling.\n"
        }
      ]
    },
    {
      "id": 379107,
      "postDate": "2018-08-30T21:53:04.087Z",
      "content": "<p>I came to join this competition and found your post. Thanks @Andres Torrubia for informing the community promptly.</p>",
      "rawMarkdown": "I came to join this competition and found your post. Thanks @Andres Torrubia for informing the community promptly."
    },
    {
      "id": 377334,
      "postDate": "2018-08-29T01:19:02.473Z",
      "content": "<p>Wow, how did you even find this leak? maybe ML?</p>",
      "rawMarkdown": "Wow, how did you even find this leak? maybe ML?",
      "replies": [
        {
          "id": 377339,
          "postDate": "2018-08-29T01:51:46.060Z",
          "content": "<p>I'm \"auditing\" the Coursera class  \"How to Win a Data Science Competition: Learn from Top Kagglers\" -- the instructors discuss various ways of looking for leaks.</p>\n\n<p>The instructors even say that using leaks is valid in the competition (i.e. fitting the letter of the law, though not the spirit since you can't depend on leaks when solving the <em>real</em> problem) and suggest looking for leaks before spending too much time with machine learning.  Finding leaks in their title competition at <a href=\"https://www.kaggle.com/c/competitive-data-science-final-project\">https://www.kaggle.com/c/competitive-data-science-final-project</a>  is even an assignment during the course.</p>",
          "rawMarkdown": "I'm \"auditing\" the Coursera class  \"How to Win a Data Science Competition: Learn from Top Kagglers\" -- the instructors discuss various ways of looking for leaks.\n\nThe instructors even say that using leaks is valid in the competition (i.e. fitting the letter of the law, though not the spirit since you can't depend on leaks when solving the *real* problem) and suggest looking for leaks before spending too much time with machine learning.  Finding leaks in their title competition at https://www.kaggle.com/c/competitive-data-science-final-project  is even an assignment during the course."
        },
        {
          "id": 377364,
          "postDate": "2018-08-29T03:08:46.960Z",
          "content": "<p>I also knew this leak. </p>\n\n<p>I felt I saw the same image several times in the dataset when I visually checked how my model performs. I hypothesized the organizer used photos of the same place at different time, and ran KNN. Actually, KNN didn't work so well because they are shifted and the majority of them are just the sea, but it didn't take so much time to find pairs of shifted images in it.</p>",
          "rawMarkdown": "I also knew this leak. \n\nI felt I saw the same image several times in the dataset when I visually checked how my model performs. I hypothesized the organizer used photos of the same place at different time, and ran KNN. Actually, KNN didn't work so well because they are shifted and the majority of them are just the sea, but it didn't take so much time to find pairs of shifted images in it.",
          "votes": 2
        }
      ]
    },
    {
      "id": 377120,
      "postDate": "2018-08-28T15:44:45.283Z",
      "content": "<p>We can always hope that there has been a publication mistake and they have another labeled dataset completely different that can be used quickly as new test set.</p>",
      "rawMarkdown": "We can always hope that there has been a publication mistake and they have another labeled dataset completely different that can be used quickly as new test set."
    }
  ],
  "comments": [
    {
      "id": 377084,
      "author_name": "Azat Akhtyamov",
      "author_url": "",
      "post_date": "2018-08-28T15:03:10.293000",
      "content": "<p>Well, I hope they use a slightly more careful technique to check data while building aircrafts.</p>",
      "votes": 21,
      "replies": []
    },
    {
      "id": 377068,
      "author_name": "Pavel Iakubovskii (qubvel)",
      "author_url": "",
      "post_date": "2018-08-28T14:44:18.183000",
      "content": "<p>Pictures shifted relative to each other by 256 pixels. It looks like organizers slice large pictures with a step of 256 and randomly divided them into train and test sets. Cutting all the images 256x256 they can be found easier because they match almost exactly (except for compression artifacts).</p>",
      "votes": 21,
      "replies": []
    },
    {
      "id": 377059,
      "author_name": "Cut Onion",
      "author_url": "",
      "post_date": "2018-08-28T14:30:04.663000",
      "content": "<p>I heard Kaggle charges competition sponsors with hefty sums to check everything and so on. In particular, it was one of the reasons they refused to award points on competitions where they didn't get their part (because, you know, it was \"unverified\" and possibly low quality data). So, given rate of leaks (and such freaking outrageously simple ones) on \"checked\" data I think they are just deceiving their clients claiming that they check anything and in reality just publishing the data as is. </p>",
      "votes": 27,
      "replies": [
        {
          "id": 377127,
          "author_name": "Alexander Veysov",
          "author_url": "",
          "post_date": "2018-08-28T15:57:51.500000",
          "content": "<p>Well, the leaks become more and more outrageous)\nFeaturing the SAME images is too much even for Kaggle =)</p>\n\n<p>Maybe if the crowd smells blood, it will force them to make some <em>staff related</em> decisions?</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 377134,
          "author_name": "Will Cukierski",
          "author_url": "",
          "post_date": "2018-08-28T16:14:56.780000",
          "content": "<p>Hi Sergey,</p>\n\n<p>I did not make this competition, but I have a lot of experience on both sides of competitions. I was a top competitor prior to joining Kaggle and have probably created more ML competitions than anybody on the planet. I don't say this to convince you I'm smart (the jury is very out on that front!), but rather to make the case that I am experienced. I have suffered through the throes of leakage many times, each time vowing that this will never happen again! But, it does, and we end up in a forum thread like this, with hard-working Kagglers wondering how a self-respecting data scientist could let this through.</p>\n\n<p>Finding leakage is surprisingly hard. Leakage looks trivial through the 20/20 lens of hindsight, but it's an infinite minefield when walking forward. Leakage changes shapes. Leakage hides in black boxes that we--we who did not create the data--often don't have the luxury of seeing. Leakage evades attempts to systemically detect it. You check exact duplicates? You missed fuzzy matches. You checked fuzzy matches? You missed reflected matches. You check fuzzy reflected matches? You missed scaled matches. You silently prevented 4 huge sources of leakage that Kagglers will never know they were spared from? Oops, there was a 5th. If you enumerated a check against every type of leakage that has ever plagued a dataset on Kaggle, you'd still likely miss the next time around. There is simply no substitute (that we've found) for 500 people spending a month hammering on a dataset.</p>\n\n<p>Kaggle and host data scientists are humans. We have the same time crunches and imperfect working demands that any data scientist role involves. These aren't excuses, but I do hope this reply gives you some understanding of the complexity that leakage presents on this side of the fence.</p>\n\n<p>Finally, no matter how many times leakage gets the better of us, we do have a track record of doing what we can to make things as right as possible, and you, the community, have an equally admirable track record of persisting through these course corrections. We can't ask for more than that, and continue to do our best to make quality problems.</p>",
          "votes": 46,
          "replies": []
        },
        {
          "id": 377136,
          "author_name": "Insaf Ashrapov",
          "author_url": "",
          "post_date": "2018-08-28T16:19:22.140000",
          "content": "<p>What is your next move? Are you going to restart the competition?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 377141,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-08-28T16:27:29.760000",
          "content": "<p>You can of course incentivize leak sharing e.g. during the first month or so.</p>\n\n<p>As you know some experienced Kagglers are now recommending holding off until the competition is running for a while to avoid wasting time training ML models. So why not make this an integral part of the process? </p>\n\n<p>EDAs, data exploration, leak hunting... as phase #1.</p>",
          "votes": 16,
          "replies": []
        },
        {
          "id": 377148,
          "author_name": "Taras Baranyuk",
          "author_url": "",
          "post_date": "2018-08-28T16:39:31.980000",
          "content": "<p>You can ask for the data preparation algorithm, and do not check the data itself.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 377156,
          "author_name": "Limited Release",
          "author_url": "",
          "post_date": "2018-08-28T16:50:15.843000",
          "content": "<p>I appreciate all the effort that went into checking this data set on kaggle's side. I'm mostly amazed by the naiveté of competition sponsors, like Airbus in this case or Santander in the recently finished competition. In both competitions the sponsors must have known about the presence of the leak and just hoped that the kaggle community wouldn't notice their futile attempts to obfuscate the data. It's an interesting vote of no-confidence in the community's abilities from the sponsor side.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 377163,
          "author_name": "Cut Onion",
          "author_url": "",
          "post_date": "2018-08-28T17:00:38.150000",
          "content": "<p>Hi, William. I'm sorry if I was too emotional and too harsh on your efforts (as a team). I understand a level of efforts required to prepare a good dataset and competition and limited resources are always a thing at small startup, so many things were forgiven when Kaggle was a small team of like-minded guys. However, this is no longer the case and you are a Google now. I bet no such leaks occurs in production in your own company. It means there is a possibility to have some process to eliminate or at least greatly reduce amount of leaks. </p>\n\n<p>This process should be in the same spirit as usual quality assurance in software development -- yes, there are bugs in software but we now have all kind of means to eliminate them and validate the quality of shipped code. And this process (should) starts from the very beginning not be something to hand-off to QA guys. Like not only trying some checklist but also asking how the data was created in the first place, auditing that process and making sure it works fine. Also, I hope you using all kind of (semi-)automated \"regression\"  tests from previous competitions. Because, you know, it wasn't 500 people looking on a data for a month to spot this leak. It was couple of guys who used some techniques known from previous competitions (like knn clustering to find similar images, sift matching and so on, check this video for some examples (in russian with eng slides and auto translated subs <a href=\"https://www.youtube.com/watch?v=TwIi7zoiiPs\">https://www.youtube.com/watch?v=TwIi7zoiiPs</a>)) and this is really astounding to have full test being just copy of the train. I have not started to participate yet in this competition but I can really feel the pain of those who spent a lot of time (and energy) and found their effort completely useless. </p>\n\n<p>As a second thought, maybe, if Kaggle couldn't manage to check the data for leaks with available resources then you should just host raw data as is as \"leak finding and checking pre-competition stage\" and assign your part of $$ as prize for this?  Everybody will be happy: there will be no leaks in the main stage, anyone who find leak will get their share and there will be no embarrassment whatsoever. </p>",
          "votes": 22,
          "replies": []
        },
        {
          "id": 377166,
          "author_name": "Cut Onion",
          "author_url": "",
          "post_date": "2018-08-28T17:05:01.983000",
          "content": "<p>Andres Torrubia, I'm starting to believe in synchronicity =) While I was writing my angry reply you managed to say almost the same thing I realized in the last minute or so :)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 377309,
          "author_name": "Bobby Jo",
          "author_url": "",
          "post_date": "2018-08-28T23:55:16.350000",
          "content": "<p>Think it's a great proposal for a solution.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 377593,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-08-29T12:19:41.047000",
      "content": "<p>At least we don't need to do image translation as part of data augmentation, given it is already done ;)</p>",
      "votes": 15,
      "replies": [
        {
          "id": 379020,
          "author_name": "PiGram",
          "author_url": "",
          "post_date": "2018-08-30T18:53:00.433000",
          "content": "<p>Lol!! Soo true.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 377376,
      "author_name": "Alexander Veysov",
      "author_url": "",
      "post_date": "2018-08-29T03:51:56.957000",
      "content": "<p>Dear Kaggle,</p>\n\n<p>I am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.</p>\n\n<p>Also I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.</p>\n\n<p>So, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the <strong>ML/DS competition platforms:</strong></p>\n\n<ul>\n<li>Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;</li>\n<li>DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;</li>\n<li>CrowdAI - give us your code, write tests for us, receive nothing in exchange!</li>\n<li>Codalab - it is just ... weird;</li>\n<li>Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;</li>\n</ul>\n\n<p><strong>Most cringy cases in my memory:</strong></p>\n\n<ul>\n<li>Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;</li>\n<li>DSB 2018 test set;</li>\n<li>Santander;</li>\n<li>Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;</li>\n<li>This leak so far takes the cake;</li>\n<li>(I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )</li>\n</ul>\n\n<p>So, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.</p>\n\n<p><strong>- Dataset preparation:</strong></p>\n\n<ul>\n<li><p>Build a FIREWALL between train / test / delayed test sets on the basis of FILES;</p>\n\n<ul><li>Do not try to fool the audience by pretending that you have more data than you have;</li>\n<li>Let people decide how they want to slice their data - provide the data as raw as possible;</li>\n<li>A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";</li></ul></li>\n</ul>\n\n<p><strong>- Stratifying you dataset:</strong></p>\n\n<ul>\n<li>For each competition you have to find out what drives the predictions and the challenge;</li>\n<li>For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;</li>\n<li>Actually do stratify the train / test / delayed test sets;</li>\n<li>Yest it is difficult. But why not invest all the money into making datasets palatable?;</li>\n</ul>\n\n<p><strong>- Making data SCIENCE results usable and reproducible</strong></p>\n\n<ul>\n<li>Put tough calculation limits, so that people would actually THINK before stacking 9000 models;</li>\n<li>Always forbid any external annotation efforts;</li>\n<li>Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n<ul><li>Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;</li>\n<li>And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;</li>\n<li>And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;</li>\n<li>This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;</li></ul></li>\n</ul>\n\n<p>That is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 377391,
          "author_name": "Alexander Veysov",
          "author_url": "",
          "post_date": "2018-08-29T04:48:31.030000",
          "content": "<p>Decided to move to a separate thread\n<a href=\"https://www.kaggle.com/c/airbus-ship-detection/discussion/64393\">https://www.kaggle.com/c/airbus-ship-detection/discussion/64393</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 377126,
      "author_name": "Jeff Faudi",
      "author_url": "",
      "post_date": "2018-08-28T15:57:17.813000",
      "content": "<p>Hi to all,</p>\n\n<p>Andreas, thank you for sharing this. This is of course not intended. \nWe are looking into that now and will get back to all participants really soon.</p>\n\n<p>Jeff.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 377133,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-08-28T16:06:18.167000",
          "content": "<p>You're welcome.</p>\n\n<p>BTW - It's Andres, not Andreas.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 377249,
          "author_name": "Jeff Faudi",
          "author_url": "",
          "post_date": "2018-08-28T20:26:36.387000",
          "content": "<p>Andres. Sorry for the misspelling. </p>\n\n<p>Kaggle will communicate soon on the competition status. \nAirbus will have to deliver a new validation dataset. We are currently looking into that.</p>\n\n<p>Kind regards,\nJeff.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 379453,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-31T11:30:26.773000",
          "content": "<p>Hi Jeff,</p>\n\n<p>Since data leak event occurred, is there a possibility to move the competition end date? I assume that your team will provide new set of train and test dataset for this competition? Am I correct? </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 381620,
          "author_name": "Triskelion",
          "author_url": "",
          "post_date": "2018-09-04T21:39:14.240000",
          "content": "<p>May be a good time to look at leakage through file creation/last edit timestamps. At least in the train set this is an informative feature. If not on purpose, and random split, this may also \"generalize\" to the test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 377521,
      "author_name": "Jules",
      "author_url": "",
      "post_date": "2018-08-29T10:04:35.840000",
      "content": "<p>Dear Kaggle,</p>\n\n<p>I want to propose to give the team of Andres Torrubia and Pavel Gonchar 25k.\nI can propose this without any conflict of interest since I don't know them.</p>\n\n<p>This way Kaggle sets a precedent on rewarding finding leaks and reporting them as early as possible.\nIt also gives Kaggle an incentive to be even more careful next time. \nIt could even give the community some reassurance seeing that this sort of mistakes can cost Kaggle, apart from extra time and effort, also some money. </p>\n\n<p>I talk about Kaggle here as the representative of the sponsor, who actually pays the money is something Kaggle and Airbus can decide between them. </p>\n\n<p>Edit: removed a remark that was outdated at the time of writing</p>",
      "votes": 8,
      "replies": [
        {
          "id": 377533,
          "author_name": "Alexis Letulier",
          "author_url": "",
          "post_date": "2018-08-29T10:31:03.183000",
          "content": "<p>There is indeed a lot of value in reporting the leak, more than being frustrated by not winning the competition I would have found very depressing (and I imagine Airbus even more so) of not knowing what SOTA was  for that type of tasks should we had known after the competition ended. \nThere should indeed be a way to reward this, as many others knew and were just waiting to see where it went, as it is allowed per competition rule (not judging here a game is a game).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 377548,
          "author_name": "Alexander Veysov",
          "author_url": "",
          "post_date": "2018-08-29T10:45:28.887000",
          "content": "<p>I do not have spare US$25k, but I would chip in US$100 just for the sake of it</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 377556,
          "author_name": "Jules",
          "author_url": "",
          "post_date": "2018-08-29T10:55:15.823000",
          "content": "<blockquote>\n  <p><strong>Alexis Letulier wrote</strong></p>\n  \n  <blockquote>\n    <p>There should indeed be a way to reward this, as many others knew and were just waiting to see where it went, as it is allowed per competition rule (not judging here a game is a game).</p>\n  </blockquote>\n</blockquote>\n\n<p>If Kaggle rewards early reporting of leaks it can save a lot of Kagglers a lot of work and some Kagglers a lot of uncertainty.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 377576,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2018-08-29T11:34:54.853000",
          "content": "<p>$25k would be awesome 😂but so will be this:</p>\n\n<p><img src=\"https://i.imgur.com/Dx15bVo.png\" alt=\"Leak TShirt\"></p>",
          "votes": 22,
          "replies": []
        }
      ]
    },
    {
      "id": 377073,
      "author_name": "Alexander Veysov",
      "author_url": "",
      "post_date": "2018-08-28T14:47:37.833000",
      "content": "<p>I suppose the only way for them to redeem themselves is:</p>\n\n<ul>\n<li>Extend the duration of the competition; </li>\n<li>Change the dataset; </li>\n<li>Do   something about the prizes to account for double GPU time in  case of  dataset changes;</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 379455,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-31T11:33:48.180000",
          "content": "<p>I do agree on your proposals! </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 377122,
      "author_name": "Wudi Wang",
      "author_url": "",
      "post_date": "2018-08-28T15:53:05.250000",
      "content": "<p>Oh, shoot....\nI will give up on this competition if there is no solution to this. \nWhat a waste of my GPU time.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 377631,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-29T13:41:53.833000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 377102,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2018-08-28T15:22:49.433000",
      "content": "<p>I was thinking that 0.98+ may not be possible even with hand labeling:) Thanks for sharing.</p>\n\n<p>This competition can only be saved if it is converted to a kernel competition with strict time and memory constraints. Even this may not be enough, because now we know that it is better to overfit on train images. So they may replace the test set with the one that they saved for Algorithm Speed Challenge.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 377078,
      "author_name": "Alexis Letulier",
      "author_url": "",
      "post_date": "2018-08-28T14:52:37.647000",
      "content": "<p>What a disappointment, at least I got a model that detects quite well ships (even from non Airbus images), that I will be able to use for my job...</p>",
      "votes": 3,
      "replies": [
        {
          "id": 377822,
          "author_name": "dhammack",
          "author_url": "",
          "post_date": "2018-08-29T19:50:02.860000",
          "content": "<p>You had a strong model very quickly! Did you already have one lying around from your work that you used for the competition? Very impressive score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 377825,
          "author_name": "Alexis Letulier",
          "author_url": "",
          "post_date": "2018-08-29T19:57:18.010000",
          "content": "<p>No, its the other way around, I participated in order to have a model I could use in my work and was genuinely surprised to lead. We´ll see with the new dataset if the score holds or if I was only overfitting. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 377067,
      "author_name": "Alexander Veysov",
      "author_url": "",
      "post_date": "2018-08-28T14:42:30.237000",
      "content": "<p>Amazing. I was able to get into top-7 with only one commit and now this. \nI feel appalled. My righteous indignation has no limits.\nAnd you know what - I did not even check for such bs.\n People are correct that Kaggle's admins do not know what the hell the are doing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 378370,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-08-30T06:51:46.990000",
      "content": "<p>I am happy I did not spent time in this competition yet. Planned to use the stuff I learned from tgs-salt-competition at the end of this competition. (As might some others) Let's see what organizers will do...</p>\n\n<p>I also want to thank Andres and his team for their honesty. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 377081,
      "author_name": "krop",
      "author_url": "",
      "post_date": "2018-08-28T14:55:11.563000",
      "content": "<p>So <strong>all</strong> images from test are from train? Surely you can't be serious!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 377087,
          "author_name": "Cut Onion",
          "author_url": "",
          "post_date": "2018-08-28T15:03:26.077000",
          "content": "<p>Obviously, there is a room for improvement for several thousandth as score isn't a perfect 1 =)</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 377094,
      "author_name": "n01z3",
      "author_url": "",
      "post_date": "2018-08-28T15:19:07.180000",
      "content": "<p>So is it possible to restart competition without leak? But I guess there is no possibility for that, because we have almost 100% of labels from whole dataset.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 377099,
          "author_name": "Alexis Letulier",
          "author_url": "",
          "post_date": "2018-08-28T15:22:12.790000",
          "content": "<p>To do that you will need a completely new test set, and this is not something that can be created quickly nor cheaply.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 377114,
          "author_name": "Alexander Veysov",
          "author_url": "",
          "post_date": "2018-08-28T15:37:38.603000",
          "content": "<p>As for creating datasets ... if you just go to Google Earth to SF / HK / Singapore etc - you will be amazed how easy actually collecting such a dataset is.</p>\n\n<p>Labeling is a different thing ofc.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 377117,
          "author_name": "Alexis Letulier",
          "author_url": "",
          "post_date": "2018-08-28T15:42:01.643000",
          "content": "<p>I was of course discussing about labeling.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 379107,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-08-30T21:53:04.087000",
      "content": "<p>I came to join this competition and found your post. Thanks @Andres Torrubia for informing the community promptly.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 377334,
      "author_name": "Master",
      "author_url": "",
      "post_date": "2018-08-29T01:19:02.473000",
      "content": "<p>Wow, how did you even find this leak? maybe ML?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 377339,
          "author_name": "Todd D. Vance",
          "author_url": "",
          "post_date": "2018-08-29T01:51:46.060000",
          "content": "<p>I'm \"auditing\" the Coursera class  \"How to Win a Data Science Competition: Learn from Top Kagglers\" -- the instructors discuss various ways of looking for leaks.</p>\n\n<p>The instructors even say that using leaks is valid in the competition (i.e. fitting the letter of the law, though not the spirit since you can't depend on leaks when solving the <em>real</em> problem) and suggest looking for leaks before spending too much time with machine learning.  Finding leaks in their title competition at <a href=\"https://www.kaggle.com/c/competitive-data-science-final-project\">https://www.kaggle.com/c/competitive-data-science-final-project</a>  is even an assignment during the course.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 377364,
          "author_name": "Komaki",
          "author_url": "",
          "post_date": "2018-08-29T03:08:46.960000",
          "content": "<p>I also knew this leak. </p>\n\n<p>I felt I saw the same image several times in the dataset when I visually checked how my model performs. I hypothesized the organizer used photos of the same place at different time, and ran KNN. Actually, KNN didn't work so well because they are shifted and the majority of them are just the sea, but it didn't take so much time to find pairs of shifted images in it.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 377120,
      "author_name": "Alexis Letulier",
      "author_url": "",
      "post_date": "2018-08-28T15:44:45.283000",
      "content": "<p>We can always hope that there has been a publication mistake and they have another labeled dataset completely different that can be used quickly as new test set.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "377037": "Hi all,\n\nAs you may have seen we got 0.98+ LB with just one sub. Although we had a lot of good ideas for the challenge my partner Pavel found a leak which renders this competition almost useless.\n\nThe images in test are just shifted versions of images in train. The problem also happens in train, it looks like airbus had bigger images and then the 768x768 are random crops of the bigger images; but it looks they didn't check whether there were any overlaps.\n\ne.g. train/dd38da47f.jpg and train/02776139a.jpg\n\nHow to find:\n\n - Run nearest neighbors on all images \n - For each image take N closest neighbors and find where it overlaps\n\nYou can quickly recover almost (all?) of test images.\n\n",
    "377084": "Well, I hope they use a slightly more careful technique to check data while building aircrafts.",
    "377068": "Pictures shifted relative to each other by 256 pixels. It looks like organizers slice large pictures with a step of 256 and randomly divided them into train and test sets. Cutting all the images 256x256 they can be found easier because they match almost exactly (except for compression artifacts).",
    "377059": "I heard Kaggle charges competition sponsors with hefty sums to check everything and so on. In particular, it was one of the reasons they refused to award points on competitions where they didn't get their part (because, you know, it was \"unverified\" and possibly low quality data). So, given rate of leaks (and such freaking outrageously simple ones) on \"checked\" data I think they are just deceiving their clients claiming that they check anything and in reality just publishing the data as is. ",
    "377593": "At least we don't need to do image translation as part of data augmentation, given it is already done ;)",
    "377376": "Dear Kaggle,\n\nI am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.\n\nAlso I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.\n\nSo, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the **ML/DS competition platforms:**\n\n - Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;\n - DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;\n - CrowdAI - give us your code, write tests for us, receive nothing in exchange!\n - Codalab - it is just ... weird;\n - Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;\n\n**Most cringy cases in my memory:**\n\n - Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;\n - DSB 2018 test set;\n - Santander;\n - Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;\n - This leak so far takes the cake;\n - (I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )\n\n\nSo, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.\n\n**- Dataset preparation:**\n  \n\n - Build a FIREWALL between train / test / delayed test sets on the basis of FILES;\n\n \n  - Do not try to fool the audience by pretending that you have more data than you have;\n  - Let people decide how they want to slice their data - provide the data as raw as possible;\n  - A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";\n\n**- Stratifying you dataset:**\n\n - For each competition you have to find out what drives the predictions and the challenge;\n - For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;\n - Actually do stratify the train / test / delayed test sets;\n - Yest it is difficult. But why not invest all the money into making datasets palatable?;\n\n**- Making data SCIENCE results usable and reproducible**\n\n- Put tough calculation limits, so that people would actually THINK before stacking 9000 models;\n- Always forbid any external annotation efforts;\n- Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n  - Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;\n  - And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;\n  - And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;\n  - This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;\n\nThat is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.",
    "377126": "Hi to all,\n\nAndreas, thank you for sharing this. This is of course not intended. \nWe are looking into that now and will get back to all participants really soon.\n\nJeff.",
    "377521": "Dear Kaggle,\n\nI want to propose to give the team of Andres Torrubia and Pavel Gonchar 25k.\nI can propose this without any conflict of interest since I don't know them.\n\nThis way Kaggle sets a precedent on rewarding finding leaks and reporting them as early as possible.\nIt also gives Kaggle an incentive to be even more careful next time. \nIt could even give the community some reassurance seeing that this sort of mistakes can cost Kaggle, apart from extra time and effort, also some money. \n\nI talk about Kaggle here as the representative of the sponsor, who actually pays the money is something Kaggle and Airbus can decide between them. \n\nEdit: removed a remark that was outdated at the time of writing\n\n\n\n",
    "377073": "I suppose the only way for them to redeem themselves is:\n\n - Extend the duration of the competition; \n - Change the dataset; \n - Do   something about the prizes to account for double GPU time in  case of  dataset changes;\n\n",
    "377122": "Oh, shoot....\nI will give up on this competition if there is no solution to this. \nWhat a waste of my GPU time.",
    "377102": "I was thinking that 0.98+ may not be possible even with hand labeling:) Thanks for sharing.\n\nThis competition can only be saved if it is converted to a kernel competition with strict time and memory constraints. Even this may not be enough, because now we know that it is better to overfit on train images. So they may replace the test set with the one that they saved for Algorithm Speed Challenge.",
    "377078": "What a disappointment, at least I got a model that detects quite well ships (even from non Airbus images), that I will be able to use for my job...",
    "377067": "Amazing. I was able to get into top-7 with only one commit and now this. \nI feel appalled. My righteous indignation has no limits.\nAnd you know what - I did not even check for such bs.\n People are correct that Kaggle's admins do not know what the hell the are doing.",
    "378370": "I am happy I did not spent time in this competition yet. Planned to use the stuff I learned from tgs-salt-competition at the end of this competition. (As might some others) Let's see what organizers will do...\n\nI also want to thank Andres and his team for their honesty. ",
    "377081": "So **all** images from test are from train? Surely you can't be serious!",
    "377094": "So is it possible to restart competition without leak? But I guess there is no possibility for that, because we have almost 100% of labels from whole dataset.",
    "379107": "I came to join this competition and found your post. Thanks @Andres Torrubia for informing the community promptly.",
    "377334": "Wow, how did you even find this leak? maybe ML?",
    "377120": "We can always hope that there has been a publication mistake and they have another labeled dataset completely different that can be used quickly as new test set."
  }
}