{
  "id": 64393,
  "title": "Dear Kaggle! ",
  "url": "/competitions/airbus-ship-detection/discussion/64393",
  "author_name": "",
  "post_date": "2018-08-29T04:38:16.336763800Z",
  "votes": 12,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Dear Kaggle,</p>\n\n<p>I am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.</p>\n\n<p>Also I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.</p>\n\n<p>So, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the <strong>ML/DS competition platforms:</strong></p>\n\n<ul>\n<li>Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;</li>\n<li>DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;</li>\n<li>CrowdAI - give us your code, write tests for us, receive nothing in exchange!</li>\n<li>Codalab - it is just ... weird;</li>\n<li>Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;</li>\n</ul>\n\n<p><strong>Most cringy cases in my memory:</strong></p>\n\n<ul>\n<li>Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;</li>\n<li>DSB 2018 test set;</li>\n<li>Santander;</li>\n<li>Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;</li>\n<li>This leak so far takes the cake;</li>\n<li>(I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )</li>\n</ul>\n\n<p>So, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.</p>\n\n<p><strong>- Dataset preparation:</strong></p>\n\n<ul>\n<li><p>Build a FIREWALL between train / test / delayed test sets on the basis of FILES;</p>\n\n<ul><li>Do not try to fool the audience by pretending that you have more data than you have;</li>\n<li>Let people decide how they want to slice their data - provide the data as raw as possible;</li>\n<li>A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";</li></ul></li>\n</ul>\n\n<p><strong>- Stratifying you dataset:</strong></p>\n\n<ul>\n<li>For each competition you have to find out what drives the predictions and the challenge;</li>\n<li>For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;</li>\n<li>Actually do stratify the train / test / delayed test sets;</li>\n<li>Yest it is difficult. But why not invest all the money into making datasets palatable?;</li>\n</ul>\n\n<p><strong>- Making data SCIENCE results usable and reproducible</strong></p>\n\n<ul>\n<li>Put tough calculation limits, so that people would actually THINK before stacking 9000 models;</li>\n<li>Always forbid any external annotation efforts;</li>\n<li>Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n<ul><li>Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;</li>\n<li>And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;</li>\n<li>And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;</li>\n<li>This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;</li></ul></li>\n</ul>\n\n<p>That is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.</p>",
  "messages": [
    {
      "id": "377385",
      "postDate": "08/29/2018 04:38:16",
      "content": "<p>Dear Kaggle,</p>\n\n<p>I am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.</p>\n\n<p>Also I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.</p>\n\n<p>So, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the <strong>ML/DS competition platforms:</strong></p>\n\n<ul>\n<li>Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;</li>\n<li>DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;</li>\n<li>CrowdAI - give us your code, write tests for us, receive nothing in exchange!</li>\n<li>Codalab - it is just ... weird;</li>\n<li>Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;</li>\n</ul>\n\n<p><strong>Most cringy cases in my memory:</strong></p>\n\n<ul>\n<li>Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;</li>\n<li>DSB 2018 test set;</li>\n<li>Santander;</li>\n<li>Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;</li>\n<li>This leak so far takes the cake;</li>\n<li>(I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )</li>\n</ul>\n\n<p>So, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.</p>\n\n<p><strong>- Dataset preparation:</strong></p>\n\n<ul>\n<li><p>Build a FIREWALL between train / test / delayed test sets on the basis of FILES;</p>\n\n<ul><li>Do not try to fool the audience by pretending that you have more data than you have;</li>\n<li>Let people decide how they want to slice their data - provide the data as raw as possible;</li>\n<li>A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";</li></ul></li>\n</ul>\n\n<p><strong>- Stratifying you dataset:</strong></p>\n\n<ul>\n<li>For each competition you have to find out what drives the predictions and the challenge;</li>\n<li>For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;</li>\n<li>Actually do stratify the train / test / delayed test sets;</li>\n<li>Yest it is difficult. But why not invest all the money into making datasets palatable?;</li>\n</ul>\n\n<p><strong>- Making data SCIENCE results usable and reproducible</strong></p>\n\n<ul>\n<li>Put tough calculation limits, so that people would actually THINK before stacking 9000 models;</li>\n<li>Always forbid any external annotation efforts;</li>\n<li>Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n<ul><li>Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;</li>\n<li>And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;</li>\n<li>And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;</li>\n<li>This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;</li></ul></li>\n</ul>\n\n<p>That is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.</p>",
      "rawMarkdown": "Dear Kaggle,\n\nI am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.\n\nAlso I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.\n\nSo, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the **ML/DS competition platforms:**\n\n - Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;\n - DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;\n - CrowdAI - give us your code, write tests for us, receive nothing in exchange!\n - Codalab - it is just ... weird;\n - Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;\n\n**Most cringy cases in my memory:**\n\n - Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;\n - DSB 2018 test set;\n - Santander;\n - Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;\n - This leak so far takes the cake;\n - (I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )\n\n\nSo, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.\n\n**- Dataset preparation:**\n  \n\n - Build a FIREWALL between train / test / delayed test sets on the basis of FILES;\n\n \n  - Do not try to fool the audience by pretending that you have more data than you have;\n  - Let people decide how they want to slice their data - provide the data as raw as possible;\n  - A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";\n\n**- Stratifying you dataset:**\n\n - For each competition you have to find out what drives the predictions and the challenge;\n - For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;\n - Actually do stratify the train / test / delayed test sets;\n - Yest it is difficult. But why not invest all the money into making datasets palatable?;\n\n**- Making data SCIENCE results usable and reproducible**\n\n- Put tough calculation limits, so that people would actually THINK before stacking 9000 models;\n- Always forbid any external annotation efforts;\n- Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n  - Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;\n  - And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;\n  - And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;\n  - This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;\n\nThat is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.",
      "votes": null
    },
    {
      "id": "377527",
      "postDate": "08/29/2018 10:19:40",
      "content": "<p>Lots of common sense in your post, but some items are debatable still IMHO.</p>\n\n<ol>\n<li><p>Leaks can still happen with all the things you list.  </p></li>\n<li><p>You assume sponsors want production ready models.  They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.</p></li>\n<li><p>Sponsors can also get interesting modeling ideas, and interesting tooling ideas from a solution that stacks 9000 models.</p></li>\n<li><p>Last, what is the issue with using 10 ResNet models for passenger screening?  What if using 10 instead of 1 saves some lives?  Don't say that you cannot deploy 10 models in a production environment ;)</p></li>\n</ol>\n\n<p>I fully agree with all the rest, and I really enjoyed reading your message!  And don't assume Kaggle team will not hear you either ;)</p>",
      "rawMarkdown": "Lots of common sense in your post, but some items are debatable still IMHO.\n\n 1. Leaks can still happen with all the things you list.  \n\n 2.  You assume sponsors want production ready models.  They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.\n\n 3. Sponsors can also get interesting modeling ideas, and interesting tooling ideas from a solution that stacks 9000 models.\n\n 4. Last, what is the issue with using 10 ResNet models for passenger screening?  What if using 10 instead of 1 saves some lives?  Don't say that you cannot deploy 10 models in a production environment ;)\n\nI fully agree with all the rest, and I really enjoyed reading your message!  And don't assume Kaggle team will not hear you either ;)",
      "votes": null
    },
    {
      "id": "377545",
      "postDate": "08/29/2018 10:43:32",
      "content": "<p>Many thanks for your comments.</p>\n\n<blockquote>\n  <p>Leaks can still happen with all the things you list. </p>\n</blockquote>\n\n<p>Ofc. And this would only work for images. For tables I honestly have no idea how to handle train/test split - it is too dataset specific. You have to understand the drivers / dataset well.\nBut you at least would make a reasonable and open effort to avoid it.\nAn effort that would be obvious to the public. It would also level the playing field a bit and discourage any techniques like:\n- LB probing\n- Leak expoitation\n- Metric exploitation \n- Pseudo-labelling\netc etc\nI am not saying these are not valid techniques. They are just not in the spirit of knowledge pursuit.\nBut that is my bias.</p>\n\n<blockquote>\n  <p>You assume sponsors want production ready models. They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.</p>\n</blockquote>\n\n<p>Naturally yes. But overall this is detrimental to the broader ML community.\nI understand that whoever pays for the music orders the music.\nBut unless the goals match the goals listed above - the ML community loses in general.</p>\n\n<blockquote>\n  <p>Last, what is the issue with using 10 ResNet models for passenger screening? What if using 10 instead of 1 saves some lives? Don't say that you cannot deploy 10 models in a production environment ;)</p>\n</blockquote>\n\n<p>Your criticism is valid. But I used this just to illustrate my point.\nI should have written something like - \"if there was some anti-stacking rule, it would force people to come up with actually clever and interesting ideas, like proper loss weighting, new losses, new layers etc etc\"\nJust imagine, that each of these ResNets were trained 5% better, together it will give 5+% more saved lives.</p>\n\n<blockquote>\n  <p>I really enjoyed reading your message!</p>\n</blockquote>\n\n<p>I was a bit passionate about my message, but I am happy someone shares my ideas)</p>",
      "rawMarkdown": "Many thanks for your comments.\n\n&gt; Leaks can still happen with all the things you list. \n\nOfc. And this would only work for images. For tables I honestly have no idea how to handle train/test split - it is too dataset specific. You have to understand the drivers / dataset well.\nBut you at least would make a reasonable and open effort to avoid it.\nAn effort that would be obvious to the public. It would also level the playing field a bit and discourage any techniques like:\n- LB probing\n- Leak expoitation\n- Metric exploitation \n- Pseudo-labelling\netc etc\nI am not saying these are not valid techniques. They are just not in the spirit of knowledge pursuit.\nBut that is my bias.\n\n&gt; You assume sponsors want production ready models. They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.\n\nNaturally yes. But overall this is detrimental to the broader ML community.\nI understand that whoever pays for the music orders the music.\nBut unless the goals match the goals listed above - the ML community loses in general.\n\n&gt; Last, what is the issue with using 10 ResNet models for passenger screening? What if using 10 instead of 1 saves some lives? Don't say that you cannot deploy 10 models in a production environment ;)\n\nYour criticism is valid. But I used this just to illustrate my point.\nI should have written something like - \"if there was some anti-stacking rule, it would force people to come up with actually clever and interesting ideas, like proper loss weighting, new losses, new layers etc etc\"\nJust imagine, that each of these ResNets were trained 5% better, together it will give 5+% more saved lives.\n\n&gt; I really enjoyed reading your message!\n\nI was a bit passionate about my message, but I am happy someone shares my ideas)",
      "votes": null
    },
    {
      "id": "377591",
      "postDate": "08/29/2018 12:16:01",
      "content": "<p>Re simple model vs complex stacks I have an anecdote.  In Talking Data recent competition, I had the simplest solution among top 10 teams: a single LightGBM with 48 features than ranked 6th.  Yet the sponsor expressed no interest in knowing more about it.</p>",
      "rawMarkdown": "Re simple model vs complex stacks I have an anecdote.  In Talking Data recent competition, I had the simplest solution among top 10 teams: a single LightGBM with 48 features than ranked 6th.  Yet the sponsor expressed no interest in knowing more about it.",
      "votes": null
    },
    {
      "id": "377599",
      "postDate": "08/29/2018 12:26:22",
      "content": "<p>I have a similar funny episode.\nThere was a jungle competition - detecting animals on jungle videos with 1TB of data.</p>\n\n<ul>\n<li>Dmytro won with a simple and elegant solution;</li>\n<li>ZFTurbo passed us with 4-5 levels of stacking;</li>\n<li>We stacked 2 levels;</li>\n</ul>\n\n<p>I then wrote to the hosts and said that all this stacking is BS, we are ready to implement a MOBILE OPTIMIZED model for a really small fee, that would run on small devices - and <strong>would actually be deployable in the jungle</strong>. And you know what? They did not care.\nBut then, after 2-3+ months - they packaged the stacked models under another set of abstractions and published it. No one cared of course - because you will not have a GPU in the jungle.</p>\n\n<p>It felt so good.</p>",
      "rawMarkdown": "I have a similar funny episode.\nThere was a jungle competition - detecting animals on jungle videos with 1TB of data.\n\n -  Dmytro won with a simple and elegant solution;\n -   ZFTurbo passed us with 4-5 levels of stacking;\n -   We stacked 2 levels;\n\nI then wrote to the hosts and said that all this stacking is BS, we are ready to implement a MOBILE OPTIMIZED model for a really small fee, that would run on small devices - and **would actually be deployable in the jungle**. And you know what? They did not care.\nBut then, after 2-3+ months - they packaged the stacked models under another set of abstractions and published it. No one cared of course - because you will not have a GPU in the jungle.\n\nIt felt so good.",
      "votes": null
    },
    {
      "id": "377604",
      "postDate": "08/29/2018 12:41:04",
      "content": "<p>Was it a Kaggle competition?</p>",
      "rawMarkdown": "Was it a Kaggle competition?",
      "votes": null
    },
    {
      "id": "377607",
      "postDate": "08/29/2018 12:44:32",
      "content": "<p>No it was on DrivenData</p>",
      "rawMarkdown": "No it was on DrivenData",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 377527,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/29/2018 10:19:40",
      "content": "<p>Lots of common sense in your post, but some items are debatable still IMHO.</p>\n\n<ol>\n<li><p>Leaks can still happen with all the things you list.  </p></li>\n<li><p>You assume sponsors want production ready models.  They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.</p></li>\n<li><p>Sponsors can also get interesting modeling ideas, and interesting tooling ideas from a solution that stacks 9000 models.</p></li>\n<li><p>Last, what is the issue with using 10 ResNet models for passenger screening?  What if using 10 instead of 1 saves some lives?  Don't say that you cannot deploy 10 models in a production environment ;)</p></li>\n</ol>\n\n<p>I fully agree with all the rest, and I really enjoyed reading your message!  And don't assume Kaggle team will not hear you either ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 377545,
          "author_name": "snakers41",
          "author_url": "",
          "post_date": "08/29/2018 10:43:32",
          "content": "<p>Many thanks for your comments.</p>\n\n<blockquote>\n  <p>Leaks can still happen with all the things you list. </p>\n</blockquote>\n\n<p>Ofc. And this would only work for images. For tables I honestly have no idea how to handle train/test split - it is too dataset specific. You have to understand the drivers / dataset well.\nBut you at least would make a reasonable and open effort to avoid it.\nAn effort that would be obvious to the public. It would also level the playing field a bit and discourage any techniques like:\n- LB probing\n- Leak expoitation\n- Metric exploitation \n- Pseudo-labelling\netc etc\nI am not saying these are not valid techniques. They are just not in the spirit of knowledge pursuit.\nBut that is my bias.</p>\n\n<blockquote>\n  <p>You assume sponsors want production ready models. They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.</p>\n</blockquote>\n\n<p>Naturally yes. But overall this is detrimental to the broader ML community.\nI understand that whoever pays for the music orders the music.\nBut unless the goals match the goals listed above - the ML community loses in general.</p>\n\n<blockquote>\n  <p>Last, what is the issue with using 10 ResNet models for passenger screening? What if using 10 instead of 1 saves some lives? Don't say that you cannot deploy 10 models in a production environment ;)</p>\n</blockquote>\n\n<p>Your criticism is valid. But I used this just to illustrate my point.\nI should have written something like - \"if there was some anti-stacking rule, it would force people to come up with actually clever and interesting ideas, like proper loss weighting, new losses, new layers etc etc\"\nJust imagine, that each of these ResNets were trained 5% better, together it will give 5+% more saved lives.</p>\n\n<blockquote>\n  <p>I really enjoyed reading your message!</p>\n</blockquote>\n\n<p>I was a bit passionate about my message, but I am happy someone shares my ideas)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377591,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/29/2018 12:16:01",
          "content": "<p>Re simple model vs complex stacks I have an anecdote.  In Talking Data recent competition, I had the simplest solution among top 10 teams: a single LightGBM with 48 features than ranked 6th.  Yet the sponsor expressed no interest in knowing more about it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377599,
          "author_name": "snakers41",
          "author_url": "",
          "post_date": "08/29/2018 12:26:22",
          "content": "<p>I have a similar funny episode.\nThere was a jungle competition - detecting animals on jungle videos with 1TB of data.</p>\n\n<ul>\n<li>Dmytro won with a simple and elegant solution;</li>\n<li>ZFTurbo passed us with 4-5 levels of stacking;</li>\n<li>We stacked 2 levels;</li>\n</ul>\n\n<p>I then wrote to the hosts and said that all this stacking is BS, we are ready to implement a MOBILE OPTIMIZED model for a really small fee, that would run on small devices - and <strong>would actually be deployable in the jungle</strong>. And you know what? They did not care.\nBut then, after 2-3+ months - they packaged the stacked models under another set of abstractions and published it. No one cared of course - because you will not have a GPU in the jungle.</p>\n\n<p>It felt so good.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377604,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/29/2018 12:41:04",
          "content": "<p>Was it a Kaggle competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377607,
          "author_name": "snakers41",
          "author_url": "",
          "post_date": "08/29/2018 12:44:32",
          "content": "<p>No it was on DrivenData</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "377385": "Dear Kaggle,\n\nI am writing this message with only a minor hope that I will be heard.\nIn the world of information bubbles and bozo explosion, sometimes it seems\nthat indeed the blind lead the way in 95% of cases.\n\nAlso I kind of hope that the missions of such competitions is to i) find the best solution on the market ii) actually use it in production, and not actually hype / BS / PR.\n\nSo, at first a bit a sketch on how I (and maybe some people will even agree with me)\nsee the **ML/DS competition platforms:**\n\n - Kaggle - the home of dataset leaks (no major recent / high profile competition in my memory was without cringe), poor administration, stacking and \"stupid\" solutions - below I will explain why;\n - DrivenData - they started well, but became cringy in the end. Plus no interesting competitions now;\n - CrowdAI - give us your code, write tests for us, receive nothing in exchange!\n - Codalab - it is just ... weird;\n - Topcoder - once in a while they have a decent ML competition and in SpaceNet their rules were actually really clever;\n\n**Most cringy cases in my memory:**\n\n - Passenger screening challenge that forbid prizes for non-US citizens ... and top solution was 10 ResNets;\n - DSB 2018 test set;\n - Santander;\n - Camera challenge - where people with plain python scraper could get 10x more data than the hosts - it laughable;\n - This leak so far takes the cake;\n - (I omit all the TF challenges filled with blatant marketing, but Google bought Kaggle - it is business baby =) )\n\n\nSo, actually enough ranting, how to fix it?\nBasically you can adopt a pipeline used by Topcoder's admin from SpaceNet with slight modifications to lower barriers to entry.\nI understand that in case of table data competitions leaks can be really tricky, but in case of images this should just work.\nBut table data competitions are over9000 LightGBM / XGBoosts anyway.\n\n**- Dataset preparation:**\n  \n\n - Build a FIREWALL between train / test / delayed test sets on the basis of FILES;\n\n \n  - Do not try to fool the audience by pretending that you have more data than you have;\n  - Let people decide how they want to slice their data - provide the data as raw as possible;\n  - A decent recent example was CrowAI map house detection, but in the end their stage 2 instructions were \"meh\";\n\n**- Stratifying you dataset:**\n\n - For each competition you have to find out what drives the predictions and the challenge;\n - For ships it is: i) balanced share of images w or wo ships ii) images with merged ships iii) images with small / big ships;\n - Actually do stratify the train / test / delayed test sets;\n - Yest it is difficult. But why not invest all the money into making datasets palatable?;\n\n**- Making data SCIENCE results usable and reproducible**\n\n- Put tough calculation limits, so that people would actually THINK before stacking 9000 models;\n- Always forbid any external annotation efforts;\n- Forbid usage of any hardcoded leaky stuff by means of asking top 10-20 contenders in the private phase to:\n  - Provide a container that i) has to train on a NEW train dataset from scratch to similar performance ii) has to validate on a new dataset well;\n  - And yes, in this case it is easy to fall in a pitfall like CrowdAI did - force some half baked solution on everyone - basically you should just provide general instructions on dockerization and then just let each member of community decide what they do inside;\n  - And yes - keep barriers to entry low for first stage of competition, but keep barriers to actually win high, requiring high skill set and DECENT effort;\n  - This way you will motivate people to come up with solutions that actually benefit the community - and not this usual blending / stacking crap;\n\nThat is all. I highly doubt that you will hear me, since it looks like Kaggle is a long way on the \"silver spoon\" and \"bozo explosion\" path. But a man can have a dream.",
    "377527": "Lots of common sense in your post, but some items are debatable still IMHO.\n\n 1. Leaks can still happen with all the things you list.  \n\n 2.  You assume sponsors want production ready models.  They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.\n\n 3. Sponsors can also get interesting modeling ideas, and interesting tooling ideas from a solution that stacks 9000 models.\n\n 4. Last, what is the issue with using 10 ResNet models for passenger screening?  What if using 10 instead of 1 saves some lives?  Don't say that you cannot deploy 10 models in a production environment ;)\n\nI fully agree with all the rest, and I really enjoyed reading your message!  And don't assume Kaggle team will not hear you either ;)",
    "377545": "Many thanks for your comments.\n\n&gt; Leaks can still happen with all the things you list. \n\nOfc. And this would only work for images. For tables I honestly have no idea how to handle train/test split - it is too dataset specific. You have to understand the drivers / dataset well.\nBut you at least would make a reasonable and open effort to avoid it.\nAn effort that would be obvious to the public. It would also level the playing field a bit and discourage any techniques like:\n- LB probing\n- Leak expoitation\n- Metric exploitation \n- Pseudo-labelling\netc etc\nI am not saying these are not valid techniques. They are just not in the spirit of knowledge pursuit.\nBut that is my bias.\n\n&gt; You assume sponsors want production ready models. They may rather want to know what is the limit of accuracy that one can get with stacking 9000 models so that they can assess the quality of their own models.\n\nNaturally yes. But overall this is detrimental to the broader ML community.\nI understand that whoever pays for the music orders the music.\nBut unless the goals match the goals listed above - the ML community loses in general.\n\n&gt; Last, what is the issue with using 10 ResNet models for passenger screening? What if using 10 instead of 1 saves some lives? Don't say that you cannot deploy 10 models in a production environment ;)\n\nYour criticism is valid. But I used this just to illustrate my point.\nI should have written something like - \"if there was some anti-stacking rule, it would force people to come up with actually clever and interesting ideas, like proper loss weighting, new losses, new layers etc etc\"\nJust imagine, that each of these ResNets were trained 5% better, together it will give 5+% more saved lives.\n\n&gt; I really enjoyed reading your message!\n\nI was a bit passionate about my message, but I am happy someone shares my ideas)",
    "377591": "Re simple model vs complex stacks I have an anecdote.  In Talking Data recent competition, I had the simplest solution among top 10 teams: a single LightGBM with 48 features than ranked 6th.  Yet the sponsor expressed no interest in knowing more about it.",
    "377599": "I have a similar funny episode.\nThere was a jungle competition - detecting animals on jungle videos with 1TB of data.\n\n -  Dmytro won with a simple and elegant solution;\n -   ZFTurbo passed us with 4-5 levels of stacking;\n -   We stacked 2 levels;\n\nI then wrote to the hosts and said that all this stacking is BS, we are ready to implement a MOBILE OPTIMIZED model for a really small fee, that would run on small devices - and **would actually be deployable in the jungle**. And you know what? They did not care.\nBut then, after 2-3+ months - they packaged the stacked models under another set of abstractions and published it. No one cared of course - because you will not have a GPU in the jungle.\n\nIt felt so good.",
    "377604": "Was it a Kaggle competition?",
    "377607": "No it was on DrivenData"
  },
  "source": "meta"
}