{
  "id": 238758,
  "title": "This should be kernel competition ",
  "url": "/competitions/seti-breakthrough-listen/discussion/238758",
  "author_name": "Serigne ",
  "post_date": "2021-05-13T10:28:41.702000",
  "votes": 58,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Given the moderate size of the data, the metric, and the objective. </p>\n<p>Otherwise I'm afraid it will end up with endless ensembling and postprocessing tricks.  </p>",
  "messages": [
    {
      "id": 1305508,
      "postDate": "2021-05-13T10:28:41.703Z",
      "content": "<p>Given the moderate size of the data, the metric, and the objective. </p>\n<p>Otherwise I'm afraid it will end up with endless ensembling and postprocessing tricks.  </p>",
      "rawMarkdown": "Given the moderate size of the data, the metric, and the objective. \n\nOtherwise I'm afraid it will end up with endless ensembling and postprocessing tricks.  ",
      "votes": 58
    },
    {
      "id": 1305631,
      "postDate": "2021-05-13T11:58:04.563Z",
      "content": "<p>Your point is right !!! But kernel competition is sometimes disappointing. Personally i almost lost love with Kaggle after my Hubmap experience. I got a 0.000 LB because of memory error behind the hood when kernels are rerun. Yet i put so much effort in it. This kind of competition is not necessarily the solution. Maybe Kaggle Community will have to think further.</p>",
      "rawMarkdown": "Your point is right !!! But kernel competition is sometimes disappointing. Personally i almost lost love with Kaggle after my Hubmap experience. I got a 0.000 LB because of memory error behind the hood when kernels are rerun. Yet i put so much effort in it. This kind of competition is not necessarily the solution. Maybe Kaggle Community will have to think further.",
      "votes": 17,
      "replies": [
        {
          "id": 1306480,
          "postDate": "2021-05-13T19:35:16.407Z",
          "content": "<p>HubMap was quite disappointing for me too. Especially after the public LB went completely crazy after the hand labeling incident.</p>",
          "rawMarkdown": "HubMap was quite disappointing for me too. Especially after the public LB went completely crazy after the hand labeling incident.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1306103,
      "postDate": "2021-05-13T16:01:03.190Z",
      "content": "<p>I agree. Best solution might be around 0.99 very soon with hundred of models ensembled. And for the last 2 months people are going to fight for 0.001 digit. I may give it a try anyway to see if I can discover UFO.</p>",
      "rawMarkdown": "I agree. Best solution might be around 0.99 very soon with hundred of models ensembled. And for the last 2 months people are going to fight for 0.001 digit. I may give it a try anyway to see if I can discover UFO.",
      "votes": 18,
      "replies": [
        {
          "id": 1340447,
          "postDate": "2021-06-08T00:33:39.350Z",
          "content": "<p>Guess, you have discovered the UFO :) I will vote for kernel's only comp any day if data is midsize! Helps us come-up with better solutions that could be used in a feasible setting TBH.</p>\n<p>Seeing you around forums after Riiid! </p>",
          "rawMarkdown": "Guess, you have discovered the UFO :) I will vote for kernel's only comp any day if data is midsize! Helps us come-up with better solutions that could be used in a feasible setting TBH.\n\nSeeing you around forums after Riiid! "
        },
        {
          "id": 1340485,
          "postDate": "2021-06-08T02:59:14.637Z",
          "content": "<p>Plus they are often productionizable by host</p>",
          "rawMarkdown": "Plus they are often productionizable by host"
        }
      ]
    },
    {
      "id": 1306130,
      "postDate": "2021-05-13T16:13:23.293Z",
      "content": "<p>I think it's even worse: with ~36k test samples, hand-labeling is not out of the question for a dedicated group of people…</p>",
      "rawMarkdown": "I think it's even worse: with ~36k test samples, hand-labeling is not out of the question for a dedicated group of people...",
      "votes": 9,
      "replies": [
        {
          "id": 1306133,
          "postDate": "2021-05-13T16:15:39.910Z",
          "content": "<p>Agree, this is my biggest concern. A team of 5 only need to label 7000 each, over the space of 3 months.<br>\nThat's less than a hundred a day per person.<br>\nAnd most of those will be easily labelled as normal by a quick glance.</p>",
          "rawMarkdown": "Agree, this is my biggest concern. A team of 5 only need to label 7000 each, over the space of 3 months.\nThat's less than a hundred a day per person.\nAnd most of those will be easily labelled as normal by a quick glance.",
          "votes": 6
        },
        {
          "id": 1306171,
          "postDate": "2021-05-13T16:42:42.973Z",
          "content": "<p>But handlabelling the test set is not allowed, correct? If someone does it, he needs to use it in an indirect way (like select submission that is close to the handlabels)</p>",
          "rawMarkdown": "But handlabelling the test set is not allowed, correct? If someone does it, he needs to use it in an indirect way (like select submission that is close to the handlabels)",
          "votes": 4
        },
        {
          "id": 1306417,
          "postDate": "2021-05-13T18:46:15.093Z",
          "content": "<p>Exactly. It helps to select the \"best\" model.</p>",
          "rawMarkdown": "Exactly. It helps to select the \"best\" model.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1307133,
      "postDate": "2021-05-14T09:16:57.080Z",
      "content": "<p>I am not quite sure why emsembling or postprocessing would be considered a problem.</p>\n<p>The realisation that a consensus of multiple predictors is often stronger than a single model is an important insight that goes back to Galton's work more than a century ago. Recently it has been fashionable to call this the Wisdom of Crowds.</p>\n<p>Postprocessing is essentially sanity checking our predictions for compliance with what we think we know about the problem. It's what we teach students to do when answering an exam question. It has worked well in the Indoor Location competition and gives us confidence that our predictions are reasonable.</p>",
      "rawMarkdown": "I am not quite sure why emsembling or postprocessing would be considered a problem.\n\nThe realisation that a consensus of multiple predictors is often stronger than a single model is an important insight that goes back to Galton's work more than a century ago. Recently it has been fashionable to call this the Wisdom of Crowds.\n\nPostprocessing is essentially sanity checking our predictions for compliance with what we think we know about the problem. It's what we teach students to do when answering an exam question. It has worked well in the Indoor Location competition and gives us confidence that our predictions are reasonable.",
      "votes": 3,
      "replies": [
        {
          "id": 1307330,
          "postDate": "2021-05-14T11:32:47.670Z",
          "content": "<p>In constrained mode like kernel competitions, you will spend more time thinking about clever , optimized, and mostly efficient solution, trying to improve your solution etc. Which will be more useful for the organizers, particularly during test time inference.   While the other way around you may just keep ensembling as many models as possible.</p>\n<p>And you have still room to ensemble 5 or 6 models in kernel competitions not 100 or 200 </p>\n<p>By postprocessing  I mean trying to find and exploit every trick and leak on test data for the sake of score. Things irrelevent in real world examples. </p>",
          "rawMarkdown": "In constrained mode like kernel competitions, you will spend more time thinking about clever , optimized, and mostly efficient solution, trying to improve your solution etc. Which will be more useful for the organizers, particularly during test time inference.   While the other way around you may just keep ensembling as many models as possible.\n\nAnd you have still room to ensemble 5 or 6 models in kernel competitions not 100 or 200 \n\nBy postprocessing  I mean trying to find and exploit every trick and leak on test data for the sake of score. Things irrelevent in real world examples. \n\n",
          "votes": 13
        },
        {
          "id": 1309493,
          "postDate": "2021-05-16T02:57:57.570Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1320619,
      "postDate": "2021-05-24T07:43:46.617Z",
      "content": "<p>First of all, it is certain that code competition can reduce the number of ensemble models, but do you think that hundreds of models will work well together? I've tried that the non code competition ensemble's revenue is higher, but it's not what you think. When hundreds of models ensemble scores first increase and then decrease, there is such an optimal value. The optimal value is not that the more models, the better. It's a model with distinctive characteristics, and the greater the gap between model structures, The more effective the ensemble will be,</p>",
      "rawMarkdown": "First of all, it is certain that code competition can reduce the number of ensemble models, but do you think that hundreds of models will work well together? I've tried that the non code competition ensemble's revenue is higher, but it's not what you think. When hundreds of models ensemble scores first increase and then decrease, there is such an optimal value. The optimal value is not that the more models, the better. It's a model with distinctive characteristics, and the greater the gap between model structures, The more effective the ensemble will be,",
      "votes": 1,
      "replies": [
        {
          "id": 1320623,
          "postDate": "2021-05-24T07:47:16.400Z",
          "content": "<p>I used to use 12-18 models to achieve better results in the field of image segmentation, which is also innovative. I removed the smaller segmentation mask from the results of 12-18 models under a certain threshold, leaving only the larger segmentation mask. In this way, I got better results by taking the intersection, instead of directly performing the ensemble of 12-18 models without any technology</p>",
          "rawMarkdown": "I used to use 12-18 models to achieve better results in the field of image segmentation, which is also innovative. I removed the smaller segmentation mask from the results of 12-18 models under a certain threshold, leaving only the larger segmentation mask. In this way, I got better results by taking the intersection, instead of directly performing the ensemble of 12-18 models without any technology"
        }
      ]
    },
    {
      "id": 1320515,
      "postDate": "2021-05-24T06:25:38.007Z",
      "content": "<p>I think there are too many train data,we can easily get high score with simple model.Maybe we need more test data,otherwise we can only improve about 0.01~0.03 score in next two months.</p>",
      "rawMarkdown": "I think there are too many train data,we can easily get high score with simple model.Maybe we need more test data,otherwise we can only improve about 0.01~0.03 score in next two months.",
      "votes": 1
    },
    {
      "id": 1311299,
      "postDate": "2021-05-17T09:54:55.653Z",
      "content": "<p>Well since everyone will ensemble - only those that are having meaningful stuff in their ensemble will win. Honestly just adding together more garbage is not producing anything better. Not to mention that modelling on base models outputs/metadata is also fun - and important skill that needs to be exercised. It is like tabular competition on top of a DL competition. So perfectly fine (in my opinion) to still have a \"regular\" competitions…</p>",
      "rawMarkdown": "Well since everyone will ensemble - only those that are having meaningful stuff in their ensemble will win. Honestly just adding together more garbage is not producing anything better. Not to mention that modelling on base models outputs/metadata is also fun - and important skill that needs to be exercised. It is like tabular competition on top of a DL competition. So perfectly fine (in my opinion) to still have a \"regular\" competitions...",
      "votes": 1
    },
    {
      "id": 1308994,
      "postDate": "2021-05-15T15:21:41.947Z",
      "content": "<p>I do agree but I guess the risk of overfitting when ensembling is strong enough to demotivate people from building crazy ensemble , I predict there will be huge shakeup in the end as happened in SIIM ISIC Melanoma competition .</p>",
      "rawMarkdown": "I do agree but I guess the risk of overfitting when ensembling is strong enough to demotivate people from building crazy ensemble , I predict there will be huge shakeup in the end as happened in SIIM ISIC Melanoma competition .",
      "votes": 1,
      "replies": [
        {
          "id": 1311303,
          "postDate": "2021-05-17T09:59:56.083Z",
          "content": "<p>as the public score closer to 1, there may be a chance that overfitting the public data. So I agree with that there maybe a huge shake with the private board.</p>",
          "rawMarkdown": "as the public score closer to 1, there may be a chance that overfitting the public data. So I agree with that there maybe a huge shake with the private board."
        }
      ]
    },
    {
      "id": 1307420,
      "postDate": "2021-05-14T12:36:17.437Z",
      "content": "<p>You can contact the orgnizers and kaggle stuffs.</p>",
      "rawMarkdown": "You can contact the orgnizers and kaggle stuffs.",
      "votes": 1
    },
    {
      "id": 1308461,
      "postDate": "2021-05-15T08:20:58.033Z",
      "content": "<p>Agreen, the kernel competition actually limited the number of ensemble models, otherwise it would end up with ensembling game.</p>",
      "rawMarkdown": "Agreen, the kernel competition actually limited the number of ensemble models, otherwise it would end up with ensembling game.",
      "votes": 2
    },
    {
      "id": 1353231,
      "postDate": "2021-06-17T02:00:04.977Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1319084,
      "postDate": "2021-05-22T21:35:13.793Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1319114,
          "postDate": "2021-05-22T22:52:52.507Z",
          "content": "<p>You can work in a notebook.</p>",
          "rawMarkdown": "You can work in a notebook."
        },
        {
          "id": 1319116,
          "postDate": "2021-05-22T22:59:19.880Z",
          "content": "<p>Yep, I am aware. But kind of makes it hard if I am doing multiple competitions at once, with 30 hours of cap on GPU. </p>\n<p>Guess I should buy a CollabPro subscription. </p>",
          "rawMarkdown": "Yep, I am aware. But kind of makes it hard if I am doing multiple competitions at once, with 30 hours of cap on GPU. \n\nGuess I should buy a CollabPro subscription. "
        },
        {
          "id": 1319696,
          "postDate": "2021-05-23T12:44:18.593Z",
          "content": "<p>Ignorant of your financial situation, and you probably meant it as a joke, but a 128 GB USB is $8</p>",
          "rawMarkdown": "Ignorant of your financial situation, and you probably meant it as a joke, but a 128 GB USB is $8"
        },
        {
          "id": 1319772,
          "postDate": "2021-05-23T13:41:35.433Z",
          "content": "<p>I often advocate for Kernel Competitions. Its just plains the field for all participants. Generally, participants with more hardware access / from personal setup / professional sponsor have a greater edge. </p>\n<p>Also, I meant it as a joke in the slightest capacity 👍</p>",
          "rawMarkdown": "I often advocate for Kernel Competitions. Its just plains the field for all participants. Generally, participants with more hardware access / from personal setup / professional sponsor have a greater edge. \n\nAlso, I meant it as a joke in the slightest capacity 👍"
        },
        {
          "id": 1319894,
          "postDate": "2021-05-23T15:12:52.520Z",
          "content": "<p>Code competitions are great, but I wish we would not be constricted to Python or R. Open up for Mathematica or C#.</p>\n<p>By the way, I have 198 GB SSD, 8 GB RAM and 4x2.5 GHz, so I have to use Zen coding and feel your pain!</p>",
          "rawMarkdown": "Code competitions are great, but I wish we would not be constricted to Python or R. Open up for Mathematica or C#.\n\nBy the way, I have 198 GB SSD, 8 GB RAM and 4x2.5 GHz, so I have to use Zen coding and feel your pain!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1305631,
      "author_name": "Ulrich G.",
      "author_url": "",
      "post_date": "2021-05-13T11:58:04.563000",
      "content": "<p>Your point is right !!! But kernel competition is sometimes disappointing. Personally i almost lost love with Kaggle after my Hubmap experience. I got a 0.000 LB because of memory error behind the hood when kernels are rerun. Yet i put so much effort in it. This kind of competition is not necessarily the solution. Maybe Kaggle Community will have to think further.</p>",
      "votes": 17,
      "replies": [
        {
          "id": 1306480,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-05-13T19:35:16.407000",
          "content": "<p>HubMap was quite disappointing for me too. Especially after the public LB went completely crazy after the hand labeling incident.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1306103,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-05-13T16:01:03.190000",
      "content": "<p>I agree. Best solution might be around 0.99 very soon with hundred of models ensembled. And for the last 2 months people are going to fight for 0.001 digit. I may give it a try anyway to see if I can discover UFO.</p>",
      "votes": 18,
      "replies": [
        {
          "id": 1340447,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2021-06-08T00:33:39.350000",
          "content": "<p>Guess, you have discovered the UFO :) I will vote for kernel's only comp any day if data is midsize! Helps us come-up with better solutions that could be used in a feasible setting TBH.</p>\n<p>Seeing you around forums after Riiid! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1340485,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2021-06-08T02:59:14.637000",
          "content": "<p>Plus they are often productionizable by host</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1306130,
      "author_name": "Markus Frank",
      "author_url": "",
      "post_date": "2021-05-13T16:13:23.293000",
      "content": "<p>I think it's even worse: with ~36k test samples, hand-labeling is not out of the question for a dedicated group of people…</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1306133,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2021-05-13T16:15:39.910000",
          "content": "<p>Agree, this is my biggest concern. A team of 5 only need to label 7000 each, over the space of 3 months.<br>\nThat's less than a hundred a day per person.<br>\nAnd most of those will be easily labelled as normal by a quick glance.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1306171,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-05-13T16:42:42.973000",
          "content": "<p>But handlabelling the test set is not allowed, correct? If someone does it, he needs to use it in an indirect way (like select submission that is close to the handlabels)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1306417,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-05-13T18:46:15.093000",
          "content": "<p>Exactly. It helps to select the \"best\" model.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1307133,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2021-05-14T09:16:57.080000",
      "content": "<p>I am not quite sure why emsembling or postprocessing would be considered a problem.</p>\n<p>The realisation that a consensus of multiple predictors is often stronger than a single model is an important insight that goes back to Galton's work more than a century ago. Recently it has been fashionable to call this the Wisdom of Crowds.</p>\n<p>Postprocessing is essentially sanity checking our predictions for compliance with what we think we know about the problem. It's what we teach students to do when answering an exam question. It has worked well in the Indoor Location competition and gives us confidence that our predictions are reasonable.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1307330,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-05-14T11:32:47.670000",
          "content": "<p>In constrained mode like kernel competitions, you will spend more time thinking about clever , optimized, and mostly efficient solution, trying to improve your solution etc. Which will be more useful for the organizers, particularly during test time inference.   While the other way around you may just keep ensembling as many models as possible.</p>\n<p>And you have still room to ensemble 5 or 6 models in kernel competitions not 100 or 200 </p>\n<p>By postprocessing  I mean trying to find and exploit every trick and leak on test data for the sake of score. Things irrelevent in real world examples. </p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 1309493,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-16T02:57:57.570000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1320619,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-05-24T07:43:46.617000",
      "content": "<p>First of all, it is certain that code competition can reduce the number of ensemble models, but do you think that hundreds of models will work well together? I've tried that the non code competition ensemble's revenue is higher, but it's not what you think. When hundreds of models ensemble scores first increase and then decrease, there is such an optimal value. The optimal value is not that the more models, the better. It's a model with distinctive characteristics, and the greater the gap between model structures, The more effective the ensemble will be,</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1320623,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-05-24T07:47:16.400000",
          "content": "<p>I used to use 12-18 models to achieve better results in the field of image segmentation, which is also innovative. I removed the smaller segmentation mask from the results of 12-18 models under a certain threshold, leaving only the larger segmentation mask. In this way, I got better results by taking the intersection, instead of directly performing the ensemble of 12-18 models without any technology</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1320515,
      "author_name": "hujianxin",
      "author_url": "",
      "post_date": "2021-05-24T06:25:38.007000",
      "content": "<p>I think there are too many train data,we can easily get high score with simple model.Maybe we need more test data,otherwise we can only improve about 0.01~0.03 score in next two months.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1311299,
      "author_name": "Georgi Pamukov",
      "author_url": "",
      "post_date": "2021-05-17T09:54:55.653000",
      "content": "<p>Well since everyone will ensemble - only those that are having meaningful stuff in their ensemble will win. Honestly just adding together more garbage is not producing anything better. Not to mention that modelling on base models outputs/metadata is also fun - and important skill that needs to be exercised. It is like tabular competition on top of a DL competition. So perfectly fine (in my opinion) to still have a \"regular\" competitions…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1308994,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2021-05-15T15:21:41.947000",
      "content": "<p>I do agree but I guess the risk of overfitting when ensembling is strong enough to demotivate people from building crazy ensemble , I predict there will be huge shakeup in the end as happened in SIIM ISIC Melanoma competition .</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1311303,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2021-05-17T09:59:56.083000",
          "content": "<p>as the public score closer to 1, there may be a chance that overfitting the public data. So I agree with that there maybe a huge shake with the private board.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1307420,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-05-14T12:36:17.437000",
      "content": "<p>You can contact the orgnizers and kaggle stuffs.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1308461,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2021-05-15T08:20:58.033000",
      "content": "<p>Agreen, the kernel competition actually limited the number of ensemble models, otherwise it would end up with ensembling game.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1353231,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-17T02:00:04.977000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1319084,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-22T21:35:13.793000",
      "content": "",
      "votes": -2,
      "replies": [
        {
          "id": 1319114,
          "author_name": "glazed",
          "author_url": "",
          "post_date": "2021-05-22T22:52:52.507000",
          "content": "<p>You can work in a notebook.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319116,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2021-05-22T22:59:19.880000",
          "content": "<p>Yep, I am aware. But kind of makes it hard if I am doing multiple competitions at once, with 30 hours of cap on GPU. </p>\n<p>Guess I should buy a CollabPro subscription. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319696,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-05-23T12:44:18.593000",
          "content": "<p>Ignorant of your financial situation, and you probably meant it as a joke, but a 128 GB USB is $8</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319772,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2021-05-23T13:41:35.433000",
          "content": "<p>I often advocate for Kernel Competitions. Its just plains the field for all participants. Generally, participants with more hardware access / from personal setup / professional sponsor have a greater edge. </p>\n<p>Also, I meant it as a joke in the slightest capacity 👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319894,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-05-23T15:12:52.520000",
          "content": "<p>Code competitions are great, but I wish we would not be constricted to Python or R. Open up for Mathematica or C#.</p>\n<p>By the way, I have 198 GB SSD, 8 GB RAM and 4x2.5 GHz, so I have to use Zen coding and feel your pain!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1305508": "Given the moderate size of the data, the metric, and the objective. \n\nOtherwise I'm afraid it will end up with endless ensembling and postprocessing tricks.  ",
    "1305631": "Your point is right !!! But kernel competition is sometimes disappointing. Personally i almost lost love with Kaggle after my Hubmap experience. I got a 0.000 LB because of memory error behind the hood when kernels are rerun. Yet i put so much effort in it. This kind of competition is not necessarily the solution. Maybe Kaggle Community will have to think further.",
    "1306103": "I agree. Best solution might be around 0.99 very soon with hundred of models ensembled. And for the last 2 months people are going to fight for 0.001 digit. I may give it a try anyway to see if I can discover UFO.",
    "1306130": "I think it's even worse: with ~36k test samples, hand-labeling is not out of the question for a dedicated group of people...",
    "1307133": "I am not quite sure why emsembling or postprocessing would be considered a problem.\n\nThe realisation that a consensus of multiple predictors is often stronger than a single model is an important insight that goes back to Galton's work more than a century ago. Recently it has been fashionable to call this the Wisdom of Crowds.\n\nPostprocessing is essentially sanity checking our predictions for compliance with what we think we know about the problem. It's what we teach students to do when answering an exam question. It has worked well in the Indoor Location competition and gives us confidence that our predictions are reasonable.",
    "1320619": "First of all, it is certain that code competition can reduce the number of ensemble models, but do you think that hundreds of models will work well together? I've tried that the non code competition ensemble's revenue is higher, but it's not what you think. When hundreds of models ensemble scores first increase and then decrease, there is such an optimal value. The optimal value is not that the more models, the better. It's a model with distinctive characteristics, and the greater the gap between model structures, The more effective the ensemble will be,",
    "1320515": "I think there are too many train data,we can easily get high score with simple model.Maybe we need more test data,otherwise we can only improve about 0.01~0.03 score in next two months.",
    "1311299": "Well since everyone will ensemble - only those that are having meaningful stuff in their ensemble will win. Honestly just adding together more garbage is not producing anything better. Not to mention that modelling on base models outputs/metadata is also fun - and important skill that needs to be exercised. It is like tabular competition on top of a DL competition. So perfectly fine (in my opinion) to still have a \"regular\" competitions...",
    "1308994": "I do agree but I guess the risk of overfitting when ensembling is strong enough to demotivate people from building crazy ensemble , I predict there will be huge shakeup in the end as happened in SIIM ISIC Melanoma competition .",
    "1307420": "You can contact the orgnizers and kaggle stuffs.",
    "1308461": "Agreen, the kernel competition actually limited the number of ensemble models, otherwise it would end up with ensembling game.",
    "1353231": "",
    "1319084": ""
  }
}