{
  "id": 48482,
  "title": "Are huge ensembles getting out of control in Kaggle?",
  "url": "/competitions/sp-society-camera-model-identification/discussion/48482",
  "author_name": "",
  "post_date": "2018-01-28T13:14:46.882870Z",
  "votes": 48,
  "comment_count": 50,
  "views": 0,
  "content": "<p>I just read the top solutions of the Statoil/C-CORE Iceberg Classifier Challenge <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge\">https://www.kaggle.com/c/statoil-iceberg-classifier-challenge</a> - the top solution used an ensemble of 100 networks, and the 2nd and 3rd also used ensembles although to a lesser degree.</p>\n\n<p>I think this is totally out of control. For the organizer it would be impractical to run such a huge solution in the majority of cases, and for Kaggle as a platform it does not really level the playing field, the more GPUs you or your team have... the better off you are.</p>\n\n<p>In this competition (IEEE) there's a lot of students who are asking how to get started in deep learning with limited resources (mid-range GPUs). Sorry guys, you are going to see that the winners will be either people with a lot of GPUs or teams who merge at the last minute to ensemble disparate solutions (very easy to do).</p>\n\n<p>I believe this is a great problem. If you have been following my progress I am trying to strive for relative simplicity in my solution, otherwise competitions will become a brute-force approach.</p>\n\n<p>I ask Kaggle to consider imposing limits on the complexity of solutions, in light of the trend in previous competitions. Otherwise we will see less novel approaches than what the Kaggle community at large could provide.</p>\n\n<p>To put my money where my mouth is:</p>\n\n<p><strong>If you use my code (see other thread), you have to participate SOLO, not as part of a team (I will be updating license now)</strong>. Although not perfect, this is the best way I can think of my code is not used in ensembles by a bigger team.</p>",
  "messages": [
    {
      "id": "275199",
      "postDate": "01/28/2018 13:14:46",
      "content": "<p>I just read the top solutions of the Statoil/C-CORE Iceberg Classifier Challenge <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge\">https://www.kaggle.com/c/statoil-iceberg-classifier-challenge</a> - the top solution used an ensemble of 100 networks, and the 2nd and 3rd also used ensembles although to a lesser degree.</p>\n\n<p>I think this is totally out of control. For the organizer it would be impractical to run such a huge solution in the majority of cases, and for Kaggle as a platform it does not really level the playing field, the more GPUs you or your team have... the better off you are.</p>\n\n<p>In this competition (IEEE) there's a lot of students who are asking how to get started in deep learning with limited resources (mid-range GPUs). Sorry guys, you are going to see that the winners will be either people with a lot of GPUs or teams who merge at the last minute to ensemble disparate solutions (very easy to do).</p>\n\n<p>I believe this is a great problem. If you have been following my progress I am trying to strive for relative simplicity in my solution, otherwise competitions will become a brute-force approach.</p>\n\n<p>I ask Kaggle to consider imposing limits on the complexity of solutions, in light of the trend in previous competitions. Otherwise we will see less novel approaches than what the Kaggle community at large could provide.</p>\n\n<p>To put my money where my mouth is:</p>\n\n<p><strong>If you use my code (see other thread), you have to participate SOLO, not as part of a team (I will be updating license now)</strong>. Although not perfect, this is the best way I can think of my code is not used in ensembles by a bigger team.</p>",
      "rawMarkdown": "I just read the top solutions of the Statoil/C-CORE Iceberg Classifier Challenge https://www.kaggle.com/c/statoil-iceberg-classifier-challenge - the top solution used an ensemble of 100 networks, and the 2nd and 3rd also used ensembles although to a lesser degree.\n\nI think this is totally out of control. For the organizer it would be impractical to run such a huge solution in the majority of cases, and for Kaggle as a platform it does not really level the playing field, the more GPUs you or your team have... the better off you are.\n\nIn this competition (IEEE) there's a lot of students who are asking how to get started in deep learning with limited resources (mid-range GPUs). Sorry guys, you are going to see that the winners will be either people with a lot of GPUs or teams who merge at the last minute to ensemble disparate solutions (very easy to do).\n\nI believe this is a great problem. If you have been following my progress I am trying to strive for relative simplicity in my solution, otherwise competitions will become a brute-force approach.\n\nI ask Kaggle to consider imposing limits on the complexity of solutions, in light of the trend in previous competitions. Otherwise we will see less novel approaches than what the Kaggle community at large could provide.\n\nTo put my money where my mouth is:\n\n**If you use my code (see other thread), you have to participate SOLO, not as part of a team (I will be updating license now)**. Although not perfect, this is the best way I can think of my code is not used in ensembles by a bigger team.",
      "votes": null
    },
    {
      "id": "275221",
      "postDate": "01/28/2018 14:13:04",
      "content": "<p>I think you are totally right, I have started compete only after I got two powerfull GPUS((((</p>",
      "rawMarkdown": "I think you are totally right, I have started compete only after I got two powerfull GPUS((((",
      "votes": null
    },
    {
      "id": "275228",
      "postDate": "01/28/2018 14:22:50",
      "content": "<p>That's a great point! However, one may argue if it is possible to change the licence conditions after publication.</p>",
      "rawMarkdown": "That's a great point! However, one may argue if it is possible to change the licence conditions after publication.",
      "votes": null
    },
    {
      "id": "275229",
      "postDate": "01/28/2018 14:23:57",
      "content": "<p>Ensembling is a good systematic way to get better results. You can have better results by using more models.\nYou are also using a kind of systematic approach by using always deeper and deeper networks. Is it so different ? </p>",
      "rawMarkdown": "Ensembling is a good systematic way to get better results. You can have better results by using more models.\nYou are also using a kind of systematic approach by using always deeper and deeper networks. Is it so different ?",
      "votes": null
    },
    {
      "id": "275236",
      "postDate": "01/28/2018 14:46:02",
      "content": "<p>It would be good to have Kaggle comment on what they think about this? </p>\n\n<p>Many teams will have forked your repo and made a submission before you added a licence  - What do those teams do?</p>",
      "rawMarkdown": "It would be good to have Kaggle comment on what they think about this? \n\nMany teams will have forked your repo and made a submission before you added a licence  - What do those teams do?",
      "votes": null
    },
    {
      "id": "275237",
      "postDate": "01/28/2018 14:46:25",
      "content": "<p>yes and yes again, we must make this topic big, in order to the people from Kaggle start think about it. </p>",
      "rawMarkdown": "yes and yes again, we must make this topic big, in order to the people from Kaggle start think about it.",
      "votes": null
    },
    {
      "id": "275244",
      "postDate": "01/28/2018 15:03:59",
      "content": "<p>Large, and huge, ensembles are natural. Try to extrapolate from where this is going and beat everybody? In the meanwhile, why not change the rules so that, e.g., Statoil can sort the entries and pick the first best one that they find to be most useful?</p>",
      "rawMarkdown": "Large, and huge, ensembles are natural. Try to extrapolate from where this is going and beat everybody? In the meanwhile, why not change the rules so that, e.g., Statoil can sort the entries and pick the first best one that they find to be most useful?",
      "votes": null
    },
    {
      "id": "275252",
      "postDate": "01/28/2018 15:33:05",
      "content": "<p>People from Kaggle have already stated that in 2018 one of their priorities will be in-kernel competitions</p>",
      "rawMarkdown": "People from Kaggle have already stated that in 2018 one of their priorities will be in-kernel competitions",
      "votes": null
    },
    {
      "id": "275257",
      "postDate": "01/28/2018 15:56:00",
      "content": "<p>Good point. If you forked the repo BEFORE the addition of the condition in the README, you are OK, but not if afterwards and also further commits fall under the new condition.</p>",
      "rawMarkdown": "Good point. If you forked the repo BEFORE the addition of the condition in the README, you are OK, but not if afterwards and also further commits fall under the new condition.",
      "votes": null
    },
    {
      "id": "275258",
      "postDate": "01/28/2018 15:56:51",
      "content": "<p>I'm not a lawyer, but since the code was released under GPL 3 <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/blob/f7d40b8aab82295d4a5a3e2ada93319e5eacdd3b/license.txt\">https://github.com/antorsae/sp-society-camera-model-identification/blob/f7d40b8aab82295d4a5a3e2ada93319e5eacdd3b/license.txt</a>, teams that use code released with that license are fine. This looks like standard GPL license, with no additional SOLO clause.</p>",
      "rawMarkdown": "I'm not a lawyer, but since the code was released under GPL 3 https://github.com/antorsae/sp-society-camera-model-identification/blob/f7d40b8aab82295d4a5a3e2ada93319e5eacdd3b/license.txt, teams that use code released with that license are fine. This looks like standard GPL license, with no additional SOLO clause.",
      "votes": null
    },
    {
      "id": "275259",
      "postDate": "01/28/2018 15:57:04",
      "content": "<p>Ensembling is easy. Scaling a model deep-wise was not feasible until resnets (or LSTMs for RNNs), it required innovation.</p>",
      "rawMarkdown": "Ensembling is easy. Scaling a model deep-wise was not feasible until resnets (or LSTMs for RNNs), it required innovation.",
      "votes": null
    },
    {
      "id": "275260",
      "postDate": "01/28/2018 15:58:25",
      "content": "<p>If you did it before I added it's OK. If you did/do it later is not OK as it is an additional condition I'm asking to be fulfilled.</p>",
      "rawMarkdown": "If you did it before I added it's OK. If you did/do it later is not OK as it is an additional condition I'm asking to be fulfilled.",
      "votes": null
    },
    {
      "id": "275262",
      "postDate": "01/28/2018 16:03:14",
      "content": "<p>Yup it's easier to just brute-force your way to a good score when you can throw unlimited models at the problem and stack those models with hundreds of models etc. etc. Hence why I think the Mercari comp is more fun.</p>",
      "rawMarkdown": "Yup it's easier to just brute-force your way to a good score when you can throw unlimited models at the problem and stack those models with hundreds of models etc. etc. Hence why I think the Mercari comp is more fun.",
      "votes": null
    },
    {
      "id": "275265",
      "postDate": "01/28/2018 16:06:23",
      "content": "<p>The very simple solution exists for exclusion many nets stacking: do inference on kaggle side (docker containers e.g.) with constraints on inference time and don't provide test dataset to participants. If competition organizers need a really production-like solution they can spend a little time to care about it before competition start (I think already many organizers faced with it and their experience can be useful). </p>",
      "rawMarkdown": "The very simple solution exists for exclusion many nets stacking: do inference on kaggle side (docker containers e.g.) with constraints on inference time and don't provide test dataset to participants. If competition organizers need a really production-like solution they can spend a little time to care about it before competition start (I think already many organizers faced with it and their experience can be useful).",
      "votes": null
    },
    {
      "id": "275266",
      "postDate": "01/28/2018 16:07:39",
      "content": "<p>Each organization may have it's own rules and impose a general cap on resources used:</p>\n\n<p>For example, this one <a href=\"https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1015-functional-map-of-the-world-fmow\">https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1015-functional-map-of-the-world-fmow</a> required the top 10 competitors at the end of the 1st deadline to submit docker files for the organization to replicate the training/etc. under a certain budgeted environment, and they ran those trained models against the first test set (and later against a new hidden second test set).</p>\n\n<p>The Tensorflow Speech competition also had a prize for a low complexity model, so savvy organizations can choose this route already if they so desire.</p>",
      "rawMarkdown": "Each organization may have it's own rules and impose a general cap on resources used:\n\nFor example, this one https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1015-functional-map-of-the-world-fmow required the top 10 competitors at the end of the 1st deadline to submit docker files for the organization to replicate the training/etc. under a certain budgeted environment, and they ran those trained models against the first test set (and later against a new hidden second test set).\n\nThe Tensorflow Speech competition also had a prize for a low complexity model, so savvy organizations can choose this route already if they so desire.",
      "votes": null
    },
    {
      "id": "275267",
      "postDate": "01/28/2018 16:08:00",
      "content": "<p>@Andres I think legally it's still fine to use the original code, since it was released with a license without such a clause. I'm not speaking about whether it's ok or not to do it morally.</p>",
      "rawMarkdown": "Andres I think legally it's still fine to use the original code, since it was released with a license without such a clause. I'm not speaking about whether it's ok or not to do it morally.",
      "votes": null
    },
    {
      "id": "275271",
      "postDate": "01/28/2018 16:11:13",
      "content": "<p>I've stated my point, and Kaggle will have the final words if a winning team used it with the condition I'm stating is in the clear or not.</p>",
      "rawMarkdown": "I've stated my point, and Kaggle will have the final words if a winning team used it with the condition I'm stating is in the clear or not.",
      "votes": null
    },
    {
      "id": "275272",
      "postDate": "01/28/2018 16:11:41",
      "content": "<p>As one of the owners of the large networks you're referencing, I understand the general concern and I have a few thoughts...</p>\n\n<ol>\n<li>In our specific case, the 100+ networks train in about 6 hours which doesn't seem to cross into the range of impractical.</li>\n<li>When a challenge is posed and the stated goal is minimize loss (or maximize some measure of accuracy), the solutions that are going to be received are the ones that do just that.  Unless the problem statement is bounded by some measure of computational cost then complexity will be largely ignored.  For many problems the cost of computation is often negligible relative to the gain in accuracy.  For example Data Science Bowl, it's hard to justify the a limitation in computation cost vs the ability to identify a fatal disease.  I understand your point is that there's a practical limit is some cases and I agree.</li>\n<li>There are some solutions out there that attempt to get the gain from ensembling without the heavy cost of complexity.  I particularly like <a href=\"https://arxiv.org/pdf/1704.00109.pdf\">snapshot ensembles</a> where the idea is to find several local minima during the same training cycle by introducing large perturbations of the learning rate.  </li>\n</ol>",
      "rawMarkdown": "As one of the owners of the large networks you're referencing, I understand the general concern and I have a few thoughts...\n\n 1. In our specific case, the 100+ networks train in about 6 hours which doesn't seem to cross into the range of impractical.\n 2. When a challenge is posed and the stated goal is minimize loss (or maximize some measure of accuracy), the solutions that are going to be received are the ones that do just that.  Unless the problem statement is bounded by some measure of computational cost then complexity will be largely ignored.  For many problems the cost of computation is often negligible relative to the gain in accuracy.  For example Data Science Bowl, it's hard to justify the a limitation in computation cost vs the ability to identify a fatal disease.  I understand your point is that there's a practical limit is some cases and I agree.\n 3. There are some solutions out there that attempt to get the gain from ensembling without the heavy cost of complexity.  I particularly like [snapshot ensembles][1] where the idea is to find several local minima during the same training cycle by introducing large perturbations of the learning rate.  \n\n  [1]: https://arxiv.org/pdf/1704.00109.pdf",
      "votes": null
    },
    {
      "id": "275283",
      "postDate": "01/28/2018 16:35:53",
      "content": "<p>Hi, Pavel, could you provide more details about your information?</p>",
      "rawMarkdown": "Hi, Pavel, could you provide more details about your information?",
      "votes": null
    },
    {
      "id": "275287",
      "postDate": "01/28/2018 16:51:16",
      "content": "<p>Thanks for stepping in David, re:</p>\n\n<ol>\n<li><p>How many GPUs did you have available to train 100+ networks in 6 hours?</p></li>\n<li><p>Agreed, but right now resource-constrained competitions are not the norm (and the instances where unbounded resources are justifiable from the problem's perspective are the minority, imo).</p></li>\n<li><p>Agreed, that to me is an innovation and is the same model sampled at different points, and it doesn't tax on resources or bruteforcing (at least on training, inference is less of an issue).</p></li>\n</ol>",
      "rawMarkdown": "Thanks for stepping in David, re:\n\n1. How many GPUs did you have available to train 100+ networks in 6 hours?\n\n2. Agreed, but right now resource-constrained competitions are not the norm (and the instances where unbounded resources are justifiable from the problem's perspective are the minority, imo).\n\n3. Agreed, that to me is an innovation and is the same model sampled at different points, and it doesn't tax on resources or bruteforcing (at least on training, inference is less of an issue).",
      "votes": null
    },
    {
      "id": "275295",
      "postDate": "01/28/2018 17:45:05",
      "content": "<p>remind me, how many kernel competitions do we have vs others? </p>",
      "rawMarkdown": "remind me, how many kernel competitions do we have vs others?",
      "votes": null
    },
    {
      "id": "275328",
      "postDate": "01/28/2018 20:42:03",
      "content": "<p>Hi Andres, </p>\n\n<p>While I agree with the spirit of your argument, I suspect the issue is more with the nature of the problem we're dealing with - vanilla image classification is largely a 'solved' problem in that everyone uses a tweaked ImageNet pre-trained model. If a baseline model gets you to 90%, grinding up to the high nineties is an exercise in engineering rather than insight. </p>\n\n<p>Kernels-only competitions a la Mercari are a solution but there's still nothing stopping a team from using external compute resources to apply hyper-parameter optimization. </p>",
      "rawMarkdown": "Hi Andres, \n\nWhile I agree with the spirit of your argument, I suspect the issue is more with the nature of the problem we're dealing with - vanilla image classification is largely a 'solved' problem in that everyone uses a tweaked ImageNet pre-trained model. If a baseline model gets you to 90%, grinding up to the high nineties is an exercise in engineering rather than insight. \n\nKernels-only competitions a la Mercari are a solution but there's still nothing stopping a team from using external compute resources to apply hyper-parameter optimization.",
      "votes": null
    },
    {
      "id": "275351",
      "postDate": "01/28/2018 23:37:27",
      "content": "<p>If I were a sponsor of one of these competitions, the way I would be thinking about it is this: Give me the best unconstrained solution you have so I can see the upper limits of what's possible.  I'll then take those solutions and scale them back, if need be, based on the constraints I have on my infrastructure and use case.  If they were asked for a solution that was already resource constrained, there's a lot more work to do to get more accuracy out of the problem than there is to scale it back.  Given that the cost of computation goes down over time, I would want the flexibility of starting with the solutions that give me the most accuracy and scale as needed.</p>",
      "rawMarkdown": "If I were a sponsor of one of these competitions, the way I would be thinking about it is this: Give me the best unconstrained solution you have so I can see the upper limits of what's possible.  I'll then take those solutions and scale them back, if need be, based on the constraints I have on my infrastructure and use case.  If they were asked for a solution that was already resource constrained, there's a lot more work to do to get more accuracy out of the problem than there is to scale it back.  Given that the cost of computation goes down over time, I would want the flexibility of starting with the solutions that give me the most accuracy and scale as needed.",
      "votes": null
    },
    {
      "id": "275364",
      "postDate": "01/29/2018 00:35:48",
      "content": "<p>Hi Andres,</p>\n\n<p>I totally agree with you. I would also like to emphasize that with massive stacking models and/or with neural networks it is very difficult if not impossible to deduce properties on the data from the models, which in fact is one of the most important things that must be extracted from a model.</p>\n\n<p>I would also like to add that this type of models do not contribute anything at a theoretical level and do not extract any kind of general knowledge from the data.</p>\n\n<p>Having said that, I would like to mention that I agree with Chun Ming Lee as to whether these techniques are only for solving a specific type of problem such as image classification, from 90% upwards then ...</p>",
      "rawMarkdown": "Hi Andres,\n\nI totally agree with you. I would also like to emphasize that with massive stacking models and/or with neural networks it is very difficult if not impossible to deduce properties on the data from the models, which in fact is one of the most important things that must be extracted from a model.\n\nI would also like to add that this type of models do not contribute anything at a theoretical level and do not extract any kind of general knowledge from the data.\n\nHaving said that, I would like to mention that I agree with Chun Ming Lee as to whether these techniques are only for solving a specific type of problem such as image classification, from 90% upwards then ...",
      "votes": null
    },
    {
      "id": "275487",
      "postDate": "01/29/2018 09:31:53",
      "content": "<p>Well, with knowledge distillation you can train simpler model with predictions from your ensemble to get most of the large model accuracy.\n<a href=\"https://arxiv.org/abs/1503.02531\">https://arxiv.org/abs/1503.02531</a>\nSo it's not necessarily useless.</p>",
      "rawMarkdown": "Well, with knowledge distillation you can train simpler model with predictions from your ensemble to get most of the large model accuracy.\nhttps://arxiv.org/abs/1503.02531\nSo it's not necessarily useless.",
      "votes": null
    },
    {
      "id": "275573",
      "postDate": "01/29/2018 14:11:49",
      "content": "<p>Hi, Andres. \nI think, it would be better to share your code and approach after the end of the competition. </p>\n\n<p>Now there is too much noise around it.</p>",
      "rawMarkdown": "Hi, Andres. \nI think, it would be better to share your code and approach after the end of the competition. \n\nNow there is too much noise around it.",
      "votes": null
    },
    {
      "id": "275576",
      "postDate": "01/29/2018 14:14:16",
      "content": "<p>Agreed. </p>\n\n<p>I've already share a good baseline (enough to get ~0.95 or more in the LB) so there will be no further updates of code until the end.</p>",
      "rawMarkdown": "Agreed. \n\nI've already share a good baseline (enough to get ~0.95 or more in the LB) so there will be no further updates of code until the end.",
      "votes": null
    },
    {
      "id": "275628",
      "postDate": "01/29/2018 16:53:24",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "275733",
      "postDate": "01/29/2018 21:57:26",
      "content": "<p>Hi, Andres, I think you make a very good point.  Many of the top placers on kaggle are using these complex ensembles.  It would be nice to see some more \"elegant\" solutions to place higher on the leaderboards. </p>",
      "rawMarkdown": "Hi, Andres, I think you make a very good point.  Many of the top placers on kaggle are using these complex ensembles.  It would be nice to see some more \"elegant\" solutions to place higher on the leaderboards.",
      "votes": null
    },
    {
      "id": "275771",
      "postDate": "01/30/2018 00:39:02",
      "content": "<p>all of them started in 2017 actually, so its too early to count</p>",
      "rawMarkdown": "all of them started in 2017 actually, so its too early to count",
      "votes": null
    },
    {
      "id": "275772",
      "postDate": "01/30/2018 00:40:26",
      "content": "<p><a href=\"http://blog.kaggle.com/2018/01/22/reviewing-2017-and-previewing-2018/\">http://blog.kaggle.com/2018/01/22/reviewing-2017-and-previewing-2018/</a></p>",
      "rawMarkdown": "http://blog.kaggle.com/2018/01/22/reviewing-2017-and-previewing-2018/",
      "votes": null
    },
    {
      "id": "275879",
      "postDate": "01/30/2018 06:35:08",
      "content": "<p>Could it be that the end state of machine learning is infinitely nested models (turtles all the way down)?</p>",
      "rawMarkdown": "Could it be that the end state of machine learning is infinitely nested models (turtles all the way down)?",
      "votes": null
    },
    {
      "id": "275956",
      "postDate": "01/30/2018 10:17:53",
      "content": "<p>Good point, but it is hard to quantify the standard</p>",
      "rawMarkdown": "Good point, but it is hard to quantify the standard",
      "votes": null
    },
    {
      "id": "275966",
      "postDate": "01/30/2018 10:41:01",
      "content": "<p>I agree. I do not know about this competition, but I have not seen elegant solutions around.</p>\n\n<p>I've learned a lot from your approach and I'm still trying to get a better LB from my own approach.</p>\n\n<p>I think sharing your code helped a lot of people, but many just used to achieve a good result without knowing what they were doing.</p>",
      "rawMarkdown": "I agree. I do not know about this competition, but I have not seen elegant solutions around.\n\nI've learned a lot from your approach and I'm still trying to get a better LB from my own approach.\n\nI think sharing your code helped a lot of people, but many just used to achieve a good result without knowing what they were doing.",
      "votes": null
    },
    {
      "id": "276256",
      "postDate": "01/31/2018 03:14:29",
      "content": "<p>Human brain is using the ultimate complex, nested, ensembled approach to solving almost all problems that it encounters. It has many, many, MANY orders of magnitude more complex neural network than even the most complicated Kaggle solution. There is no <em>a priori</em> reason that we should expect some of these problems to have just a very simple and “elegant” solution.</p>",
      "rawMarkdown": "Human brain is using the ultimate complex, nested, ensembled approach to solving almost all problems that it encounters. It has many, many, MANY orders of magnitude more complex neural network than even the most complicated Kaggle solution. There is no *a priori* reason that we should expect some of these problems to have just a very simple and “elegant” solution.",
      "votes": null
    },
    {
      "id": "276341",
      "postDate": "01/31/2018 09:04:09",
      "content": "<p>Are ensembles effective? Of course they are, that's why people use them. But my point was that the excesive use of ensembles my slow down other novel (and by extension more creative) solutions.\nThey may not be elegant, but they surely pose more of an intellectual challenge than assembling GPUs.</p>",
      "rawMarkdown": "Are ensembles effective? Of course they are, that's why people use them. But my point was that the excesive use of ensembles my slow down other novel (and by extension more creative) solutions.\nThey may not be elegant, but they surely pose more of an intellectual challenge than assembling GPUs.",
      "votes": null
    },
    {
      "id": "276348",
      "postDate": "01/31/2018 09:38:26",
      "content": "<p>How ensembles impair ability to come with novel and creative solutions for those who want to do so? On the other side, every major step in computing ability gave major step in creative solutions. </p>",
      "rawMarkdown": "How ensembles impair ability to come with novel and creative solutions for those who want to do so? On the other side, every major step in computing ability gave major step in creative solutions.",
      "votes": null
    },
    {
      "id": "276354",
      "postDate": "01/31/2018 10:03:08",
      "content": "<p>In my case if I know I have the option to go for ensembles I would just devote less mental focus to come with other ideas. </p>\n\n<p>I had to (artificially, at least for this competition)... rule out ensembles off the equation, hence I need to focus on something different knowing that the other teams are going to be using ensembles + L2 networks, etc. </p>\n\n<p>I don't know how far my approach would lead me. So far I get 0.972 LB single model.</p>",
      "rawMarkdown": "In my case if I know I have the option to go for ensembles I would just devote less mental focus to come with other ideas. \n\nI had to (artificially, at least for this competition)... rule out ensembles off the equation, hence I need to focus on something different knowing that the other teams are going to be using ensembles + L2 networks, etc. \n\nI don't know how far my approach would lead me. So far I get 0.972 LB single model.",
      "votes": null
    },
    {
      "id": "276357",
      "postDate": "01/31/2018 10:09:10",
      "content": "<p>That's &lt;4 mins total training time per network (8 mins per network per GPU). Pretty nice. \nHow did you create 100+ networks? A few base networks + hyper-param (random) changes?</p>\n\n<p>I must admit \"there are not the ensembles I was looking for\" :-) and are more like macro-ensembles which are new in my book.</p>\n\n<p><img src=\"http://i.imgur.com/zoFj17z.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "That's &lt;4 mins total training time per network (8 mins per network per GPU). Pretty nice. \nHow did you create 100+ networks? A few base networks + hyper-param (random) changes?\n\nI must admit \"there are not the ensembles I was looking for\" :-) and are more like macro-ensembles which are new in my book.\n\n![enter image description here][1]\n\n\n  [1]: http://i.imgur.com/zoFj17z.png",
      "votes": null
    },
    {
      "id": "276363",
      "postDate": "01/31/2018 10:21:35",
      "content": "<p>@Andres, I sympathize with you - I wouldn't be surprised if my team gets knocked out of the top 3 by a mass ensemble solution. </p>\n\n<p>However, I still believe any solution should focus on making it easier for people with constrained resources (e.g., students) to compete, rather than trying to handicap teams. Put another way, it's easier to raise the floor rather than lower the ceiling by enforcing artificial constraints. </p>\n\n<p>Some ideas -</p>\n\n<ol>\n<li>Giving students credits for AWS/Google Cloud time</li>\n<li>Creating a separate prize category for the most compute-efficient solution </li>\n<li>Rewarding competitors who create useful kernels (think couple competitions have done this?)</li>\n</ol>\n\n<p>IMO, the best solution would be for Kaggle to spend more time designing competitions that are resistant to compute resources (an analogy would be cryptos that are resistant to ASICs :)) - my team hit public LB 0.963 three days after we entered in December so it was clear to me that this competition would end in a brutal slog. </p>",
      "rawMarkdown": "Andres, I sympathize with you - I wouldn't be surprised if my team gets knocked out of the top 3 by a mass ensemble solution. \n\nHowever, I still believe any solution should focus on making it easier for people with constrained resources (e.g., students) to compete, rather than trying to handicap teams. Put another way, it's easier to raise the floor rather than lower the ceiling by enforcing artificial constraints. \n\nSome ideas -\n\n 1. Giving students credits for AWS/Google Cloud time\n 2. Creating a separate prize category for the most compute-efficient solution \n 3. Rewarding competitors who create useful kernels (think couple competitions have done this?)\n\nIMO, the best solution would be for Kaggle to spend more time designing competitions that are resistant to compute resources (an analogy would be cryptos that are resistant to ASICs :)) - my team hit public LB 0.963 three days after we entered in December so it was clear to me that this competition would end in a brutal slog.",
      "votes": null
    },
    {
      "id": "276371",
      "postDate": "01/31/2018 10:47:32",
      "content": "<p>I agree with you. </p>",
      "rawMarkdown": "I agree with you.",
      "votes": null
    },
    {
      "id": "276394",
      "postDate": "01/31/2018 12:01:43",
      "content": "<p>Andres, you have plenty of GPU power actually! Two 1080ti is no joke and it sounds like \"Oh, I wanted to start small so I just trained DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't want to mean average several epochs predictions to not get 'unfair' ensembling advantage over guys with just 1050ti\". </p>\n\n<p>Btw, I trained MobileNet on spare 1070 for a day and it gets about 28 place right now. And our top score contains nothing fancy, just mean of networks trained during exploration which one will train better.</p>\n\n<p>So far, I've seen wins by big ensembles, l2 layer models and so on, only on boring competitions where task is largely solved already and the only difference is by chance and better cross-validation.  </p>\n\n<p>Also, there are plenty of competitions outside of Kaggle (much lesser known) and there are plenty of room to grow in them (and no one [successfully] approaches them with blind applying ton of ensembles)</p>",
      "rawMarkdown": "Andres, you have plenty of GPU power actually! Two 1080ti is no joke and it sounds like \"Oh, I wanted to start small so I just trained DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't want to mean average several epochs predictions to not get 'unfair' ensembling advantage over guys with just 1050ti\". \n\nBtw, I trained MobileNet on spare 1070 for a day and it gets about 28 place right now. And our top score contains nothing fancy, just mean of networks trained during exploration which one will train better.\n\nSo far, I've seen wins by big ensembles, l2 layer models and so on, only on boring competitions where task is largely solved already and the only difference is by chance and better cross-validation.  \n\nAlso, there are plenty of competitions outside of Kaggle (much lesser known) and there are plenty of room to grow in them (and no one [successfully] approaches them with blind applying ton of ensembles)",
      "votes": null
    },
    {
      "id": "276400",
      "postDate": "01/31/2018 12:14:40",
      "content": "<blockquote>\n  <p>Andres, you have plenty of GPU power actually! Two 1080ti is no joke\n  and it sounds like \"Oh, I wanted to start small so I just trained\n  DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't\n  want to mean average several epochs predictions to not get 'unfair'\n  ensembling advantage over guys with just 1050ti\".</p>\n</blockquote>\n\n<p>Touché.  :-)</p>\n\n<p>I actually have 2x1080Tis w/ Threadripper 1850x (16 cores) + 1x1070Ti + Ryzen 1600x (w/8 cores); but still far from teams who successfully assemble 20+ GPUs.</p>",
      "rawMarkdown": "&gt; Andres, you have plenty of GPU power actually! Two 1080ti is no joke\n&gt; and it sounds like \"Oh, I wanted to start small so I just trained\n&gt; DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't\n&gt; want to mean average several epochs predictions to not get 'unfair'\n&gt; ensembling advantage over guys with just 1050ti\".\n\nTouché.  :-)\n\nI actually have 2x1080Tis w/ Threadripper 1850x (16 cores) + 1x1070Ti + Ryzen 1600x (w/8 cores); but still far from teams who successfully assemble 20+ GPUs.",
      "votes": null
    },
    {
      "id": "276479",
      "postDate": "01/31/2018 15:38:46",
      "content": "<p>@Chun Ming Lee Honestly if kaggle kernels could provide even a crappy GPU-instance I don't see why there couldn't be more Kernel-only competitions like Mercari but for images as well. Would certainly be fun!</p>",
      "rawMarkdown": "Chun Ming Lee Honestly if kaggle kernels could provide even a crappy GPU-instance I don't see why there couldn't be more Kernel-only competitions like Mercari but for images as well. Would certainly be fun!",
      "votes": null
    },
    {
      "id": "276499",
      "postDate": "01/31/2018 16:19:57",
      "content": "<p>I agree with you.</p>",
      "rawMarkdown": "I agree with you.",
      "votes": null
    },
    {
      "id": "276543",
      "postDate": "01/31/2018 17:49:36",
      "content": "<p>IMO limiting computational resources won't necessary promote simplicity/novelty of the solutions. Instead, it will turn Kaggle competitions more into competitive programming; the winner is the one who will (at the low machine code level) write the most optimized code for the particular dataset/problem. That might be desirable in some cases, but it's not the ultimate solution.</p>",
      "rawMarkdown": "IMO limiting computational resources won't necessary promote simplicity/novelty of the solutions. Instead, it will turn Kaggle competitions more into competitive programming; the winner is the one who will (at the low machine code level) write the most optimized code for the particular dataset/problem. That might be desirable in some cases, but it's not the ultimate solution.",
      "votes": null
    },
    {
      "id": "276559",
      "postDate": "01/31/2018 18:43:27",
      "content": "<p>Hi @Andres,</p>\n\n<p>A bit off-topic, but have you considered entering another <a href=\"https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1070-geopolitical-forecasting-challenge\">IARPA competition</a>? How was your experience with the first one?</p>",
      "rawMarkdown": "Hi @Andres,\n\nA bit off-topic, but have you considered entering another [IARPA competition][1]? How was your experience with the first one?\n\n\n  [1]: https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1070-geopolitical-forecasting-challenge",
      "votes": null
    },
    {
      "id": "276572",
      "postDate": "01/31/2018 19:38:36",
      "content": "<p>I ended up #9 on that one and was tough! One reason I got #9 (I was #3 just few days early) is that I got a pretty bad flu and had to be in the hospital for a few days... almost impossible to do anything while there so the last days I lost a few positions... yeah that's my excuse! (if you look at my code in github I had a few unused architectures to train multi-image networks; a PITA in Keras; retrospectively Pytorch would have been better choice; I will be moving for Pytorch for this and other reasons).</p>\n\n<p>What I liked about IARPA FMoW competition:</p>\n\n<ul>\n<li>Challenging problem (small dataset 200-300 Gb, big one 3 Tb) </li>\n<li>Limited computational resources (limited doesn't mean small, it means constrained). </li>\n<li>Top 10 winners to submit dockers to re-run training and inference on test set, and a second hidden test set.</li>\n<li>They provided a baseline implementation (the code structure wasn't all that good imo but a great starting point, especially for me who joined pretty late)</li>\n<li>Organizers on top of things, respond in 24 hours or less. Where are you @inversion? There's questions about external datasets and manipulation left unanswered...</li>\n</ul>\n\n<p>What I didn't like:</p>\n\n<ul>\n<li>Not possible to share information, even publicly.</li>\n<li>Forum, ranking tools, etc. on topcoder not at the level of Kaggle.</li>\n</ul>",
      "rawMarkdown": "I ended up #9 on that one and was tough! One reason I got #9 (I was #3 just few days early) is that I got a pretty bad flu and had to be in the hospital for a few days... almost impossible to do anything while there so the last days I lost a few positions... yeah that's my excuse! (if you look at my code in github I had a few unused architectures to train multi-image networks; a PITA in Keras; retrospectively Pytorch would have been better choice; I will be moving for Pytorch for this and other reasons).\n\nWhat I liked about IARPA FMoW competition:\n\n - Challenging problem (small dataset 200-300 Gb, big one 3 Tb) \n - Limited computational resources (limited doesn't mean small, it means constrained). \n - Top 10 winners to submit dockers to re-run training and inference on test set, and a second hidden test set.\n - They provided a baseline implementation (the code structure wasn't all that good imo but a great starting point, especially for me who joined pretty late)\n - Organizers on top of things, respond in 24 hours or less. Where are you @inversion? There's questions about external datasets and manipulation left unanswered...\n\nWhat I didn't like:\n\n - Not possible to share information, even publicly.\n - Forum, ranking tools, etc. on topcoder not at the level of Kaggle.",
      "votes": null
    },
    {
      "id": "276650",
      "postDate": "02/01/2018 03:07:59",
      "content": "<p>sad</p>",
      "rawMarkdown": "sad",
      "votes": null
    },
    {
      "id": "276680",
      "postDate": "02/01/2018 06:11:29",
      "content": "<p>sad, my own computer runs even slower than the online kernel.</p>",
      "rawMarkdown": "sad, my own computer runs even slower than the online kernel.",
      "votes": null
    },
    {
      "id": "276754",
      "postDate": "02/01/2018 11:29:38",
      "content": "<p>I totally agree with your global statement that having massive computing power and using ensemble method will definitely give you a big advantage in several (if not all) competitions. Consequently it develops a sense a frustration for others (like me) with limited resources.\nIn the specific case of the iceberg challenge, the only goal was to get the best classifier. In a sense if we look at it as the industrial problem it is, the method did not matter and only the result was important. If the company has the power to run any model why trying to get more elegant solutions ?</p>\n\n<p>Nonetheless alternative and elegant solutions can often be interesting to investigate. I hence agree with the idea that we should find a way to promote them and not necessarily penalize big groups.\nI know that in some competitions/hackatons outside Kaggle there are often several ranking system : one for the best performances, one for the most innovative approaches, one considering the resource consumption, etc. It could something to develop on Kaggle, maybe not for all competitions but at least for more exploratory ones.</p>\n\n<p>The second thing that came to my mind is that, according to my consulting experience, there is often numerous constraints tied to a project. For instance limited memory consumption for phone application, limited computing power, running time constraint, etc. I assume that, when creating a competition, Kaggle offers the possibility to set such constraint (if not why ?) but I haven't seen it very often. Can we ask to encourage competitions with such constraint that I think are really close to a lot of data science and machine learning \"real world\" application ?</p>\n\n<p>I'd be happy to have anyone feedback on this especially Kaggle administrators</p>",
      "rawMarkdown": "I totally agree with your global statement that having massive computing power and using ensemble method will definitely give you a big advantage in several (if not all) competitions. Consequently it develops a sense a frustration for others (like me) with limited resources.\nIn the specific case of the iceberg challenge, the only goal was to get the best classifier. In a sense if we look at it as the industrial problem it is, the method did not matter and only the result was important. If the company has the power to run any model why trying to get more elegant solutions ?\n\nNonetheless alternative and elegant solutions can often be interesting to investigate. I hence agree with the idea that we should find a way to promote them and not necessarily penalize big groups.\nI know that in some competitions/hackatons outside Kaggle there are often several ranking system : one for the best performances, one for the most innovative approaches, one considering the resource consumption, etc. It could something to develop on Kaggle, maybe not for all competitions but at least for more exploratory ones.\n\nThe second thing that came to my mind is that, according to my consulting experience, there is often numerous constraints tied to a project. For instance limited memory consumption for phone application, limited computing power, running time constraint, etc. I assume that, when creating a competition, Kaggle offers the possibility to set such constraint (if not why ?) but I haven't seen it very often. Can we ask to encourage competitions with such constraint that I think are really close to a lot of data science and machine learning \"real world\" application ?\n\nI'd be happy to have anyone feedback on this especially Kaggle administrators",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 275221,
      "author_name": "maksimovka",
      "author_url": "",
      "post_date": "01/28/2018 14:13:04",
      "content": "<p>I think you are totally right, I have started compete only after I got two powerfull GPUS((((</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275228,
      "author_name": "nikolasent",
      "author_url": "",
      "post_date": "01/28/2018 14:22:50",
      "content": "<p>That's a great point! However, one may argue if it is possible to change the licence conditions after publication.</p>",
      "votes": null,
      "replies": [
        {
          "id": 275257,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 15:56:00",
          "content": "<p>Good point. If you forked the repo BEFORE the addition of the condition in the README, you are OK, but not if afterwards and also further commits fall under the new condition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275229,
      "author_name": "jeandebleau",
      "author_url": "",
      "post_date": "01/28/2018 14:23:57",
      "content": "<p>Ensembling is a good systematic way to get better results. You can have better results by using more models.\nYou are also using a kind of systematic approach by using always deeper and deeper networks. Is it so different ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 275259,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 15:57:04",
          "content": "<p>Ensembling is easy. Scaling a model deep-wise was not feasible until resnets (or LSTMs for RNNs), it required innovation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275236,
      "author_name": "craigglastonbury",
      "author_url": "",
      "post_date": "01/28/2018 14:46:02",
      "content": "<p>It would be good to have Kaggle comment on what they think about this? </p>\n\n<p>Many teams will have forked your repo and made a submission before you added a licence  - What do those teams do?</p>",
      "votes": null,
      "replies": [
        {
          "id": 275258,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "01/28/2018 15:56:51",
          "content": "<p>I'm not a lawyer, but since the code was released under GPL 3 <a href=\"https://github.com/antorsae/sp-society-camera-model-identification/blob/f7d40b8aab82295d4a5a3e2ada93319e5eacdd3b/license.txt\">https://github.com/antorsae/sp-society-camera-model-identification/blob/f7d40b8aab82295d4a5a3e2ada93319e5eacdd3b/license.txt</a>, teams that use code released with that license are fine. This looks like standard GPL license, with no additional SOLO clause.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275260,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 15:58:25",
          "content": "<p>If you did it before I added it's OK. If you did/do it later is not OK as it is an additional condition I'm asking to be fulfilled.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275267,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "01/28/2018 16:08:00",
          "content": "<p>@Andres I think legally it's still fine to use the original code, since it was released with a license without such a clause. I'm not speaking about whether it's ok or not to do it morally.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275271,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 16:11:13",
          "content": "<p>I've stated my point, and Kaggle will have the final words if a winning team used it with the condition I'm stating is in the clear or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275237,
      "author_name": "hireme",
      "author_url": "",
      "post_date": "01/28/2018 14:46:25",
      "content": "<p>yes and yes again, we must make this topic big, in order to the people from Kaggle start think about it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 275252,
          "author_name": "",
          "author_url": "",
          "post_date": "01/28/2018 15:33:05",
          "content": "<p>People from Kaggle have already stated that in 2018 one of their priorities will be in-kernel competitions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275283,
          "author_name": "shuhuagao",
          "author_url": "",
          "post_date": "01/28/2018 16:35:53",
          "content": "<p>Hi, Pavel, could you provide more details about your information?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275295,
          "author_name": "hireme",
          "author_url": "",
          "post_date": "01/28/2018 17:45:05",
          "content": "<p>remind me, how many kernel competitions do we have vs others? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275771,
          "author_name": "",
          "author_url": "",
          "post_date": "01/30/2018 00:39:02",
          "content": "<p>all of them started in 2017 actually, so its too early to count</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275772,
          "author_name": "",
          "author_url": "",
          "post_date": "01/30/2018 00:40:26",
          "content": "<p><a href=\"http://blog.kaggle.com/2018/01/22/reviewing-2017-and-previewing-2018/\">http://blog.kaggle.com/2018/01/22/reviewing-2017-and-previewing-2018/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275244,
      "author_name": "glimmung",
      "author_url": "",
      "post_date": "01/28/2018 15:03:59",
      "content": "<p>Large, and huge, ensembles are natural. Try to extrapolate from where this is going and beat everybody? In the meanwhile, why not change the rules so that, e.g., Statoil can sort the entries and pick the first best one that they find to be most useful?</p>",
      "votes": null,
      "replies": [
        {
          "id": 275266,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 16:07:39",
          "content": "<p>Each organization may have it's own rules and impose a general cap on resources used:</p>\n\n<p>For example, this one <a href=\"https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1015-functional-map-of-the-world-fmow\">https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1015-functional-map-of-the-world-fmow</a> required the top 10 competitors at the end of the 1st deadline to submit docker files for the organization to replicate the training/etc. under a certain budgeted environment, and they ran those trained models against the first test set (and later against a new hidden second test set).</p>\n\n<p>The Tensorflow Speech competition also had a prize for a low complexity model, so savvy organizations can choose this route already if they so desire.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275262,
      "author_name": "michaelsnell",
      "author_url": "",
      "post_date": "01/28/2018 16:03:14",
      "content": "<p>Yup it's easier to just brute-force your way to a good score when you can throw unlimited models at the problem and stack those models with hundreds of models etc. etc. Hence why I think the Mercari comp is more fun.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275265,
      "author_name": "cortwave",
      "author_url": "",
      "post_date": "01/28/2018 16:06:23",
      "content": "<p>The very simple solution exists for exclusion many nets stacking: do inference on kaggle side (docker containers e.g.) with constraints on inference time and don't provide test dataset to participants. If competition organizers need a really production-like solution they can spend a little time to care about it before competition start (I think already many organizers faced with it and their experience can be useful). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275272,
      "author_name": "tivfrvqhs5",
      "author_url": "",
      "post_date": "01/28/2018 16:11:41",
      "content": "<p>As one of the owners of the large networks you're referencing, I understand the general concern and I have a few thoughts...</p>\n\n<ol>\n<li>In our specific case, the 100+ networks train in about 6 hours which doesn't seem to cross into the range of impractical.</li>\n<li>When a challenge is posed and the stated goal is minimize loss (or maximize some measure of accuracy), the solutions that are going to be received are the ones that do just that.  Unless the problem statement is bounded by some measure of computational cost then complexity will be largely ignored.  For many problems the cost of computation is often negligible relative to the gain in accuracy.  For example Data Science Bowl, it's hard to justify the a limitation in computation cost vs the ability to identify a fatal disease.  I understand your point is that there's a practical limit is some cases and I agree.</li>\n<li>There are some solutions out there that attempt to get the gain from ensembling without the heavy cost of complexity.  I particularly like <a href=\"https://arxiv.org/pdf/1704.00109.pdf\">snapshot ensembles</a> where the idea is to find several local minima during the same training cycle by introducing large perturbations of the learning rate.  </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 275287,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/28/2018 16:51:16",
          "content": "<p>Thanks for stepping in David, re:</p>\n\n<ol>\n<li><p>How many GPUs did you have available to train 100+ networks in 6 hours?</p></li>\n<li><p>Agreed, but right now resource-constrained competitions are not the norm (and the instances where unbounded resources are justifiable from the problem's perspective are the minority, imo).</p></li>\n<li><p>Agreed, that to me is an innovation and is the same model sampled at different points, and it doesn't tax on resources or bruteforcing (at least on training, inference is less of an issue).</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275351,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "01/28/2018 23:37:27",
          "content": "<p>If I were a sponsor of one of these competitions, the way I would be thinking about it is this: Give me the best unconstrained solution you have so I can see the upper limits of what's possible.  I'll then take those solutions and scale them back, if need be, based on the constraints I have on my infrastructure and use case.  If they were asked for a solution that was already resource constrained, there's a lot more work to do to get more accuracy out of the problem than there is to scale it back.  Given that the cost of computation goes down over time, I would want the flexibility of starting with the solutions that give me the most accuracy and scale as needed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276357,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 10:09:10",
          "content": "<p>That's &lt;4 mins total training time per network (8 mins per network per GPU). Pretty nice. \nHow did you create 100+ networks? A few base networks + hyper-param (random) changes?</p>\n\n<p>I must admit \"there are not the ensembles I was looking for\" :-) and are more like macro-ensembles which are new in my book.</p>\n\n<p><img src=\"http://i.imgur.com/zoFj17z.png\" alt=\"enter image description here\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275328,
      "author_name": "leecming",
      "author_url": "",
      "post_date": "01/28/2018 20:42:03",
      "content": "<p>Hi Andres, </p>\n\n<p>While I agree with the spirit of your argument, I suspect the issue is more with the nature of the problem we're dealing with - vanilla image classification is largely a 'solved' problem in that everyone uses a tweaked ImageNet pre-trained model. If a baseline model gets you to 90%, grinding up to the high nineties is an exercise in engineering rather than insight. </p>\n\n<p>Kernels-only competitions a la Mercari are a solution but there's still nothing stopping a team from using external compute resources to apply hyper-parameter optimization. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275364,
      "author_name": "jmlago",
      "author_url": "",
      "post_date": "01/29/2018 00:35:48",
      "content": "<p>Hi Andres,</p>\n\n<p>I totally agree with you. I would also like to emphasize that with massive stacking models and/or with neural networks it is very difficult if not impossible to deduce properties on the data from the models, which in fact is one of the most important things that must be extracted from a model.</p>\n\n<p>I would also like to add that this type of models do not contribute anything at a theoretical level and do not extract any kind of general knowledge from the data.</p>\n\n<p>Having said that, I would like to mention that I agree with Chun Ming Lee as to whether these techniques are only for solving a specific type of problem such as image classification, from 90% upwards then ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275487,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "01/29/2018 09:31:53",
      "content": "<p>Well, with knowledge distillation you can train simpler model with predictions from your ensemble to get most of the large model accuracy.\n<a href=\"https://arxiv.org/abs/1503.02531\">https://arxiv.org/abs/1503.02531</a>\nSo it's not necessarily useless.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275573,
      "author_name": "dukhovnik",
      "author_url": "",
      "post_date": "01/29/2018 14:11:49",
      "content": "<p>Hi, Andres. \nI think, it would be better to share your code and approach after the end of the competition. </p>\n\n<p>Now there is too much noise around it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 275576,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/29/2018 14:14:16",
          "content": "<p>Agreed. </p>\n\n<p>I've already share a good baseline (enough to get ~0.95 or more in the LB) so there will be no further updates of code until the end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275966,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "01/30/2018 10:41:01",
          "content": "<p>I agree. I do not know about this competition, but I have not seen elegant solutions around.</p>\n\n<p>I've learned a lot from your approach and I'm still trying to get a better LB from my own approach.</p>\n\n<p>I think sharing your code helped a lot of people, but many just used to achieve a good result without knowing what they were doing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 275628,
      "author_name": "rhgrossm",
      "author_url": "",
      "post_date": "01/29/2018 16:53:24",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 275733,
      "author_name": "cgundlach",
      "author_url": "",
      "post_date": "01/29/2018 21:57:26",
      "content": "<p>Hi, Andres, I think you make a very good point.  Many of the top placers on kaggle are using these complex ensembles.  It would be nice to see some more \"elegant\" solutions to place higher on the leaderboards. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275879,
      "author_name": "stevethatsmyname",
      "author_url": "",
      "post_date": "01/30/2018 06:35:08",
      "content": "<p>Could it be that the end state of machine learning is infinitely nested models (turtles all the way down)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275956,
      "author_name": "yaozhenjie",
      "author_url": "",
      "post_date": "01/30/2018 10:17:53",
      "content": "<p>Good point, but it is hard to quantify the standard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 276256,
      "author_name": "tunguz",
      "author_url": "",
      "post_date": "01/31/2018 03:14:29",
      "content": "<p>Human brain is using the ultimate complex, nested, ensembled approach to solving almost all problems that it encounters. It has many, many, MANY orders of magnitude more complex neural network than even the most complicated Kaggle solution. There is no <em>a priori</em> reason that we should expect some of these problems to have just a very simple and “elegant” solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 276341,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 09:04:09",
          "content": "<p>Are ensembles effective? Of course they are, that's why people use them. But my point was that the excesive use of ensembles my slow down other novel (and by extension more creative) solutions.\nThey may not be elegant, but they surely pose more of an intellectual challenge than assembling GPUs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276348,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/31/2018 09:38:26",
          "content": "<p>How ensembles impair ability to come with novel and creative solutions for those who want to do so? On the other side, every major step in computing ability gave major step in creative solutions. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276354,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 10:03:08",
          "content": "<p>In my case if I know I have the option to go for ensembles I would just devote less mental focus to come with other ideas. </p>\n\n<p>I had to (artificially, at least for this competition)... rule out ensembles off the equation, hence I need to focus on something different knowing that the other teams are going to be using ensembles + L2 networks, etc. </p>\n\n<p>I don't know how far my approach would lead me. So far I get 0.972 LB single model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276363,
          "author_name": "leecming",
          "author_url": "",
          "post_date": "01/31/2018 10:21:35",
          "content": "<p>@Andres, I sympathize with you - I wouldn't be surprised if my team gets knocked out of the top 3 by a mass ensemble solution. </p>\n\n<p>However, I still believe any solution should focus on making it easier for people with constrained resources (e.g., students) to compete, rather than trying to handicap teams. Put another way, it's easier to raise the floor rather than lower the ceiling by enforcing artificial constraints. </p>\n\n<p>Some ideas -</p>\n\n<ol>\n<li>Giving students credits for AWS/Google Cloud time</li>\n<li>Creating a separate prize category for the most compute-efficient solution </li>\n<li>Rewarding competitors who create useful kernels (think couple competitions have done this?)</li>\n</ol>\n\n<p>IMO, the best solution would be for Kaggle to spend more time designing competitions that are resistant to compute resources (an analogy would be cryptos that are resistant to ASICs :)) - my team hit public LB 0.963 three days after we entered in December so it was clear to me that this competition would end in a brutal slog. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276394,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "01/31/2018 12:01:43",
          "content": "<p>Andres, you have plenty of GPU power actually! Two 1080ti is no joke and it sounds like \"Oh, I wanted to start small so I just trained DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't want to mean average several epochs predictions to not get 'unfair' ensembling advantage over guys with just 1050ti\". </p>\n\n<p>Btw, I trained MobileNet on spare 1070 for a day and it gets about 28 place right now. And our top score contains nothing fancy, just mean of networks trained during exploration which one will train better.</p>\n\n<p>So far, I've seen wins by big ensembles, l2 layer models and so on, only on boring competitions where task is largely solved already and the only difference is by chance and better cross-validation.  </p>\n\n<p>Also, there are plenty of competitions outside of Kaggle (much lesser known) and there are plenty of room to grow in them (and no one [successfully] approaches them with blind applying ton of ensembles)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276400,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 12:14:40",
          "content": "<blockquote>\n  <p>Andres, you have plenty of GPU power actually! Two 1080ti is no joke\n  and it sounds like \"Oh, I wanted to start small so I just trained\n  DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't\n  want to mean average several epochs predictions to not get 'unfair'\n  ensembling advantage over guys with just 1050ti\".</p>\n</blockquote>\n\n<p>Touché.  :-)</p>\n\n<p>I actually have 2x1080Tis w/ Threadripper 1850x (16 cores) + 1x1070Ti + Ryzen 1600x (w/8 cores); but still far from teams who successfully assemble 20+ GPUs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276479,
          "author_name": "michaelsnell",
          "author_url": "",
          "post_date": "01/31/2018 15:38:46",
          "content": "<p>@Chun Ming Lee Honestly if kaggle kernels could provide even a crappy GPU-instance I don't see why there couldn't be more Kernel-only competitions like Mercari but for images as well. Would certainly be fun!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276559,
          "author_name": "zaraks",
          "author_url": "",
          "post_date": "01/31/2018 18:43:27",
          "content": "<p>Hi @Andres,</p>\n\n<p>A bit off-topic, but have you considered entering another <a href=\"https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1070-geopolitical-forecasting-challenge\">IARPA competition</a>? How was your experience with the first one?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 276572,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "01/31/2018 19:38:36",
          "content": "<p>I ended up #9 on that one and was tough! One reason I got #9 (I was #3 just few days early) is that I got a pretty bad flu and had to be in the hospital for a few days... almost impossible to do anything while there so the last days I lost a few positions... yeah that's my excuse! (if you look at my code in github I had a few unused architectures to train multi-image networks; a PITA in Keras; retrospectively Pytorch would have been better choice; I will be moving for Pytorch for this and other reasons).</p>\n\n<p>What I liked about IARPA FMoW competition:</p>\n\n<ul>\n<li>Challenging problem (small dataset 200-300 Gb, big one 3 Tb) </li>\n<li>Limited computational resources (limited doesn't mean small, it means constrained). </li>\n<li>Top 10 winners to submit dockers to re-run training and inference on test set, and a second hidden test set.</li>\n<li>They provided a baseline implementation (the code structure wasn't all that good imo but a great starting point, especially for me who joined pretty late)</li>\n<li>Organizers on top of things, respond in 24 hours or less. Where are you @inversion? There's questions about external datasets and manipulation left unanswered...</li>\n</ul>\n\n<p>What I didn't like:</p>\n\n<ul>\n<li>Not possible to share information, even publicly.</li>\n<li>Forum, ranking tools, etc. on topcoder not at the level of Kaggle.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 276371,
      "author_name": "rmanaswi",
      "author_url": "",
      "post_date": "01/31/2018 10:47:32",
      "content": "<p>I agree with you. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 276499,
      "author_name": "songwei82",
      "author_url": "",
      "post_date": "01/31/2018 16:19:57",
      "content": "<p>I agree with you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 276543,
      "author_name": "herrahuu",
      "author_url": "",
      "post_date": "01/31/2018 17:49:36",
      "content": "<p>IMO limiting computational resources won't necessary promote simplicity/novelty of the solutions. Instead, it will turn Kaggle competitions more into competitive programming; the winner is the one who will (at the low machine code level) write the most optimized code for the particular dataset/problem. That might be desirable in some cases, but it's not the ultimate solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 276650,
      "author_name": "duanchaoqun",
      "author_url": "",
      "post_date": "02/01/2018 03:07:59",
      "content": "<p>sad</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 276680,
      "author_name": "",
      "author_url": "",
      "post_date": "02/01/2018 06:11:29",
      "content": "<p>sad, my own computer runs even slower than the online kernel.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 276754,
      "author_name": "pylablanche",
      "author_url": "",
      "post_date": "02/01/2018 11:29:38",
      "content": "<p>I totally agree with your global statement that having massive computing power and using ensemble method will definitely give you a big advantage in several (if not all) competitions. Consequently it develops a sense a frustration for others (like me) with limited resources.\nIn the specific case of the iceberg challenge, the only goal was to get the best classifier. In a sense if we look at it as the industrial problem it is, the method did not matter and only the result was important. If the company has the power to run any model why trying to get more elegant solutions ?</p>\n\n<p>Nonetheless alternative and elegant solutions can often be interesting to investigate. I hence agree with the idea that we should find a way to promote them and not necessarily penalize big groups.\nI know that in some competitions/hackatons outside Kaggle there are often several ranking system : one for the best performances, one for the most innovative approaches, one considering the resource consumption, etc. It could something to develop on Kaggle, maybe not for all competitions but at least for more exploratory ones.</p>\n\n<p>The second thing that came to my mind is that, according to my consulting experience, there is often numerous constraints tied to a project. For instance limited memory consumption for phone application, limited computing power, running time constraint, etc. I assume that, when creating a competition, Kaggle offers the possibility to set such constraint (if not why ?) but I haven't seen it very often. Can we ask to encourage competitions with such constraint that I think are really close to a lot of data science and machine learning \"real world\" application ?</p>\n\n<p>I'd be happy to have anyone feedback on this especially Kaggle administrators</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "275199": "I just read the top solutions of the Statoil/C-CORE Iceberg Classifier Challenge https://www.kaggle.com/c/statoil-iceberg-classifier-challenge - the top solution used an ensemble of 100 networks, and the 2nd and 3rd also used ensembles although to a lesser degree.\n\nI think this is totally out of control. For the organizer it would be impractical to run such a huge solution in the majority of cases, and for Kaggle as a platform it does not really level the playing field, the more GPUs you or your team have... the better off you are.\n\nIn this competition (IEEE) there's a lot of students who are asking how to get started in deep learning with limited resources (mid-range GPUs). Sorry guys, you are going to see that the winners will be either people with a lot of GPUs or teams who merge at the last minute to ensemble disparate solutions (very easy to do).\n\nI believe this is a great problem. If you have been following my progress I am trying to strive for relative simplicity in my solution, otherwise competitions will become a brute-force approach.\n\nI ask Kaggle to consider imposing limits on the complexity of solutions, in light of the trend in previous competitions. Otherwise we will see less novel approaches than what the Kaggle community at large could provide.\n\nTo put my money where my mouth is:\n\n**If you use my code (see other thread), you have to participate SOLO, not as part of a team (I will be updating license now)**. Although not perfect, this is the best way I can think of my code is not used in ensembles by a bigger team.",
    "275221": "I think you are totally right, I have started compete only after I got two powerfull GPUS((((",
    "275228": "That's a great point! However, one may argue if it is possible to change the licence conditions after publication.",
    "275229": "Ensembling is a good systematic way to get better results. You can have better results by using more models.\nYou are also using a kind of systematic approach by using always deeper and deeper networks. Is it so different ?",
    "275236": "It would be good to have Kaggle comment on what they think about this? \n\nMany teams will have forked your repo and made a submission before you added a licence  - What do those teams do?",
    "275237": "yes and yes again, we must make this topic big, in order to the people from Kaggle start think about it.",
    "275244": "Large, and huge, ensembles are natural. Try to extrapolate from where this is going and beat everybody? In the meanwhile, why not change the rules so that, e.g., Statoil can sort the entries and pick the first best one that they find to be most useful?",
    "275252": "People from Kaggle have already stated that in 2018 one of their priorities will be in-kernel competitions",
    "275257": "Good point. If you forked the repo BEFORE the addition of the condition in the README, you are OK, but not if afterwards and also further commits fall under the new condition.",
    "275258": "I'm not a lawyer, but since the code was released under GPL 3 https://github.com/antorsae/sp-society-camera-model-identification/blob/f7d40b8aab82295d4a5a3e2ada93319e5eacdd3b/license.txt, teams that use code released with that license are fine. This looks like standard GPL license, with no additional SOLO clause.",
    "275259": "Ensembling is easy. Scaling a model deep-wise was not feasible until resnets (or LSTMs for RNNs), it required innovation.",
    "275260": "If you did it before I added it's OK. If you did/do it later is not OK as it is an additional condition I'm asking to be fulfilled.",
    "275262": "Yup it's easier to just brute-force your way to a good score when you can throw unlimited models at the problem and stack those models with hundreds of models etc. etc. Hence why I think the Mercari comp is more fun.",
    "275265": "The very simple solution exists for exclusion many nets stacking: do inference on kaggle side (docker containers e.g.) with constraints on inference time and don't provide test dataset to participants. If competition organizers need a really production-like solution they can spend a little time to care about it before competition start (I think already many organizers faced with it and their experience can be useful).",
    "275266": "Each organization may have it's own rules and impose a general cap on resources used:\n\nFor example, this one https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1015-functional-map-of-the-world-fmow required the top 10 competitors at the end of the 1st deadline to submit docker files for the organization to replicate the training/etc. under a certain budgeted environment, and they ran those trained models against the first test set (and later against a new hidden second test set).\n\nThe Tensorflow Speech competition also had a prize for a low complexity model, so savvy organizations can choose this route already if they so desire.",
    "275267": "Andres I think legally it's still fine to use the original code, since it was released with a license without such a clause. I'm not speaking about whether it's ok or not to do it morally.",
    "275271": "I've stated my point, and Kaggle will have the final words if a winning team used it with the condition I'm stating is in the clear or not.",
    "275272": "As one of the owners of the large networks you're referencing, I understand the general concern and I have a few thoughts...\n\n 1. In our specific case, the 100+ networks train in about 6 hours which doesn't seem to cross into the range of impractical.\n 2. When a challenge is posed and the stated goal is minimize loss (or maximize some measure of accuracy), the solutions that are going to be received are the ones that do just that.  Unless the problem statement is bounded by some measure of computational cost then complexity will be largely ignored.  For many problems the cost of computation is often negligible relative to the gain in accuracy.  For example Data Science Bowl, it's hard to justify the a limitation in computation cost vs the ability to identify a fatal disease.  I understand your point is that there's a practical limit is some cases and I agree.\n 3. There are some solutions out there that attempt to get the gain from ensembling without the heavy cost of complexity.  I particularly like [snapshot ensembles][1] where the idea is to find several local minima during the same training cycle by introducing large perturbations of the learning rate.  \n\n  [1]: https://arxiv.org/pdf/1704.00109.pdf",
    "275283": "Hi, Pavel, could you provide more details about your information?",
    "275287": "Thanks for stepping in David, re:\n\n1. How many GPUs did you have available to train 100+ networks in 6 hours?\n\n2. Agreed, but right now resource-constrained competitions are not the norm (and the instances where unbounded resources are justifiable from the problem's perspective are the minority, imo).\n\n3. Agreed, that to me is an innovation and is the same model sampled at different points, and it doesn't tax on resources or bruteforcing (at least on training, inference is less of an issue).",
    "275295": "remind me, how many kernel competitions do we have vs others?",
    "275328": "Hi Andres, \n\nWhile I agree with the spirit of your argument, I suspect the issue is more with the nature of the problem we're dealing with - vanilla image classification is largely a 'solved' problem in that everyone uses a tweaked ImageNet pre-trained model. If a baseline model gets you to 90%, grinding up to the high nineties is an exercise in engineering rather than insight. \n\nKernels-only competitions a la Mercari are a solution but there's still nothing stopping a team from using external compute resources to apply hyper-parameter optimization.",
    "275351": "If I were a sponsor of one of these competitions, the way I would be thinking about it is this: Give me the best unconstrained solution you have so I can see the upper limits of what's possible.  I'll then take those solutions and scale them back, if need be, based on the constraints I have on my infrastructure and use case.  If they were asked for a solution that was already resource constrained, there's a lot more work to do to get more accuracy out of the problem than there is to scale it back.  Given that the cost of computation goes down over time, I would want the flexibility of starting with the solutions that give me the most accuracy and scale as needed.",
    "275364": "Hi Andres,\n\nI totally agree with you. I would also like to emphasize that with massive stacking models and/or with neural networks it is very difficult if not impossible to deduce properties on the data from the models, which in fact is one of the most important things that must be extracted from a model.\n\nI would also like to add that this type of models do not contribute anything at a theoretical level and do not extract any kind of general knowledge from the data.\n\nHaving said that, I would like to mention that I agree with Chun Ming Lee as to whether these techniques are only for solving a specific type of problem such as image classification, from 90% upwards then ...",
    "275487": "Well, with knowledge distillation you can train simpler model with predictions from your ensemble to get most of the large model accuracy.\nhttps://arxiv.org/abs/1503.02531\nSo it's not necessarily useless.",
    "275573": "Hi, Andres. \nI think, it would be better to share your code and approach after the end of the competition. \n\nNow there is too much noise around it.",
    "275576": "Agreed. \n\nI've already share a good baseline (enough to get ~0.95 or more in the LB) so there will be no further updates of code until the end.",
    "275628": "",
    "275733": "Hi, Andres, I think you make a very good point.  Many of the top placers on kaggle are using these complex ensembles.  It would be nice to see some more \"elegant\" solutions to place higher on the leaderboards.",
    "275771": "all of them started in 2017 actually, so its too early to count",
    "275772": "http://blog.kaggle.com/2018/01/22/reviewing-2017-and-previewing-2018/",
    "275879": "Could it be that the end state of machine learning is infinitely nested models (turtles all the way down)?",
    "275956": "Good point, but it is hard to quantify the standard",
    "275966": "I agree. I do not know about this competition, but I have not seen elegant solutions around.\n\nI've learned a lot from your approach and I'm still trying to get a better LB from my own approach.\n\nI think sharing your code helped a lot of people, but many just used to achieve a good result without knowing what they were doing.",
    "276256": "Human brain is using the ultimate complex, nested, ensembled approach to solving almost all problems that it encounters. It has many, many, MANY orders of magnitude more complex neural network than even the most complicated Kaggle solution. There is no *a priori* reason that we should expect some of these problems to have just a very simple and “elegant” solution.",
    "276341": "Are ensembles effective? Of course they are, that's why people use them. But my point was that the excesive use of ensembles my slow down other novel (and by extension more creative) solutions.\nThey may not be elegant, but they surely pose more of an intellectual challenge than assembling GPUs.",
    "276348": "How ensembles impair ability to come with novel and creative solutions for those who want to do so? On the other side, every major step in computing ability gave major step in creative solutions.",
    "276354": "In my case if I know I have the option to go for ensembles I would just devote less mental focus to come with other ideas. \n\nI had to (artificially, at least for this competition)... rule out ensembles off the equation, hence I need to focus on something different knowing that the other teams are going to be using ensembles + L2 networks, etc. \n\nI don't know how far my approach would lead me. So far I get 0.972 LB single model.",
    "276357": "That's &lt;4 mins total training time per network (8 mins per network per GPU). Pretty nice. \nHow did you create 100+ networks? A few base networks + hyper-param (random) changes?\n\nI must admit \"there are not the ensembles I was looking for\" :-) and are more like macro-ensembles which are new in my book.\n\n![enter image description here][1]\n\n\n  [1]: http://i.imgur.com/zoFj17z.png",
    "276363": "Andres, I sympathize with you - I wouldn't be surprised if my team gets knocked out of the top 3 by a mass ensemble solution. \n\nHowever, I still believe any solution should focus on making it easier for people with constrained resources (e.g., students) to compete, rather than trying to handicap teams. Put another way, it's easier to raise the floor rather than lower the ceiling by enforcing artificial constraints. \n\nSome ideas -\n\n 1. Giving students credits for AWS/Google Cloud time\n 2. Creating a separate prize category for the most compute-efficient solution \n 3. Rewarding competitors who create useful kernels (think couple competitions have done this?)\n\nIMO, the best solution would be for Kaggle to spend more time designing competitions that are resistant to compute resources (an analogy would be cryptos that are resistant to ASICs :)) - my team hit public LB 0.963 three days after we entered in December so it was clear to me that this competition would end in a brutal slog.",
    "276371": "I agree with you.",
    "276394": "Andres, you have plenty of GPU power actually! Two 1080ti is no joke and it sounds like \"Oh, I wanted to start small so I just trained DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't want to mean average several epochs predictions to not get 'unfair' ensembling advantage over guys with just 1050ti\". \n\nBtw, I trained MobileNet on spare 1070 for a day and it gets about 28 place right now. And our top score contains nothing fancy, just mean of networks trained during exploration which one will train better.\n\nSo far, I've seen wins by big ensembles, l2 layer models and so on, only on boring competitions where task is largely solved already and the only difference is by chance and better cross-validation.  \n\nAlso, there are plenty of competitions outside of Kaggle (much lesser known) and there are plenty of room to grow in them (and no one [successfully] approaches them with blind applying ton of ensembles)",
    "276400": "&gt; Andres, you have plenty of GPU power actually! Two 1080ti is no joke\n&gt; and it sounds like \"Oh, I wanted to start small so I just trained\n&gt; DenseNet201 for 200 epochs on my 2*1080ti 32 core beast but I don't\n&gt; want to mean average several epochs predictions to not get 'unfair'\n&gt; ensembling advantage over guys with just 1050ti\".\n\nTouché.  :-)\n\nI actually have 2x1080Tis w/ Threadripper 1850x (16 cores) + 1x1070Ti + Ryzen 1600x (w/8 cores); but still far from teams who successfully assemble 20+ GPUs.",
    "276479": "Chun Ming Lee Honestly if kaggle kernels could provide even a crappy GPU-instance I don't see why there couldn't be more Kernel-only competitions like Mercari but for images as well. Would certainly be fun!",
    "276499": "I agree with you.",
    "276543": "IMO limiting computational resources won't necessary promote simplicity/novelty of the solutions. Instead, it will turn Kaggle competitions more into competitive programming; the winner is the one who will (at the low machine code level) write the most optimized code for the particular dataset/problem. That might be desirable in some cases, but it's not the ultimate solution.",
    "276559": "Hi @Andres,\n\nA bit off-topic, but have you considered entering another [IARPA competition][1]? How was your experience with the first one?\n\n\n  [1]: https://www.iarpa.gov/index.php/working-with-iarpa/prize-challenges/1070-geopolitical-forecasting-challenge",
    "276572": "I ended up #9 on that one and was tough! One reason I got #9 (I was #3 just few days early) is that I got a pretty bad flu and had to be in the hospital for a few days... almost impossible to do anything while there so the last days I lost a few positions... yeah that's my excuse! (if you look at my code in github I had a few unused architectures to train multi-image networks; a PITA in Keras; retrospectively Pytorch would have been better choice; I will be moving for Pytorch for this and other reasons).\n\nWhat I liked about IARPA FMoW competition:\n\n - Challenging problem (small dataset 200-300 Gb, big one 3 Tb) \n - Limited computational resources (limited doesn't mean small, it means constrained). \n - Top 10 winners to submit dockers to re-run training and inference on test set, and a second hidden test set.\n - They provided a baseline implementation (the code structure wasn't all that good imo but a great starting point, especially for me who joined pretty late)\n - Organizers on top of things, respond in 24 hours or less. Where are you @inversion? There's questions about external datasets and manipulation left unanswered...\n\nWhat I didn't like:\n\n - Not possible to share information, even publicly.\n - Forum, ranking tools, etc. on topcoder not at the level of Kaggle.",
    "276650": "sad",
    "276680": "sad, my own computer runs even slower than the online kernel.",
    "276754": "I totally agree with your global statement that having massive computing power and using ensemble method will definitely give you a big advantage in several (if not all) competitions. Consequently it develops a sense a frustration for others (like me) with limited resources.\nIn the specific case of the iceberg challenge, the only goal was to get the best classifier. In a sense if we look at it as the industrial problem it is, the method did not matter and only the result was important. If the company has the power to run any model why trying to get more elegant solutions ?\n\nNonetheless alternative and elegant solutions can often be interesting to investigate. I hence agree with the idea that we should find a way to promote them and not necessarily penalize big groups.\nI know that in some competitions/hackatons outside Kaggle there are often several ranking system : one for the best performances, one for the most innovative approaches, one considering the resource consumption, etc. It could something to develop on Kaggle, maybe not for all competitions but at least for more exploratory ones.\n\nThe second thing that came to my mind is that, according to my consulting experience, there is often numerous constraints tied to a project. For instance limited memory consumption for phone application, limited computing power, running time constraint, etc. I assume that, when creating a competition, Kaggle offers the possibility to set such constraint (if not why ?) but I haven't seen it very often. Can we ask to encourage competitions with such constraint that I think are really close to a lot of data science and machine learning \"real world\" application ?\n\nI'd be happy to have anyone feedback on this especially Kaggle administrators"
  },
  "source": "meta"
}