{
  "id": 188716,
  "title": "Is overfitting the sad, but successful way to succeed in this competition?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/188716",
  "author_name": "from coffee import *",
  "post_date": "2020-10-04T20:39:48.656000",
  "votes": 21,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Dear fellow Kagglers,</p>\n<p>the top 15+ scored Notebooks are all derived from one or two Notebooks and have been re-published with minimal changes: \"optimized\" hyperparameters for additional overfitting.<br>\nThis mostly even happend without referencing to the original authors of the Notebooks, which is really bad practice in my opinion --&gt; plagiarism!</p>\n<h3>So, what will the private LB bring? Best overfitting wins the race, or one giant shakeup?</h3>\n<p>Check the top scoring Notebooks, dear fellow Kagglers:<br>\n<a href=\"https://ibb.co/rwMD43h\"><img src=\"https://i.ibb.co/DRk05YF/delete.png\" alt=\"delete\"></a></p>\n<p>-Comments deactivated, as the author doesn't want to hear any opinions of his fellow Kagglers.<br>\n-Notebook runs with tracebacks and bugs, but it's supposed to be like this, as \"the fixed version got worse LB\".<br>\n-No LR Schedulers, Checkpoint Savers, Val-Loss-Monitors, etc are used: Simply a fixed number of epochs (855 by trial and error) is chosen and the other hyperparameters (majorly dropout and some merging parameters) are tweaked/optimized for additional overfitting.</p>\n<p>Checking the Leaderboard supports our intuition, that many many more than the documented 30 submissions have been made with those notebooks.</p>\n<p>Sadly the community can't really learn a thing from those Notebooks.</p>\n<p>If someone upvoted those notebooks (except the \"original\" notebooks), please explain to my why you did. I would love to hear and understand.</p>\n<h3>What's your opinion? Giant shakeup for valid CV-backed solutions, or overfitters-winning?</h3>",
  "messages": [
    {
      "id": 1037323,
      "postDate": "2020-10-04T20:39:48.657Z",
      "content": "<p>Dear fellow Kagglers,</p>\n<p>the top 15+ scored Notebooks are all derived from one or two Notebooks and have been re-published with minimal changes: \"optimized\" hyperparameters for additional overfitting.<br>\nThis mostly even happend without referencing to the original authors of the Notebooks, which is really bad practice in my opinion --&gt; plagiarism!</p>\n<h3>So, what will the private LB bring? Best overfitting wins the race, or one giant shakeup?</h3>\n<p>Check the top scoring Notebooks, dear fellow Kagglers:<br>\n<a href=\"https://ibb.co/rwMD43h\"><img src=\"https://i.ibb.co/DRk05YF/delete.png\" alt=\"delete\"></a></p>\n<p>-Comments deactivated, as the author doesn't want to hear any opinions of his fellow Kagglers.<br>\n-Notebook runs with tracebacks and bugs, but it's supposed to be like this, as \"the fixed version got worse LB\".<br>\n-No LR Schedulers, Checkpoint Savers, Val-Loss-Monitors, etc are used: Simply a fixed number of epochs (855 by trial and error) is chosen and the other hyperparameters (majorly dropout and some merging parameters) are tweaked/optimized for additional overfitting.</p>\n<p>Checking the Leaderboard supports our intuition, that many many more than the documented 30 submissions have been made with those notebooks.</p>\n<p>Sadly the community can't really learn a thing from those Notebooks.</p>\n<p>If someone upvoted those notebooks (except the \"original\" notebooks), please explain to my why you did. I would love to hear and understand.</p>\n<h3>What's your opinion? Giant shakeup for valid CV-backed solutions, or overfitters-winning?</h3>",
      "rawMarkdown": "Dear fellow Kagglers,\n\nthe top 15+ scored Notebooks are all derived from one or two Notebooks and have been re-published with minimal changes: \"optimized\" hyperparameters for additional overfitting.\nThis mostly even happend without referencing to the original authors of the Notebooks, which is really bad practice in my opinion --> plagiarism!\n\n### So, what will the private LB bring? Best overfitting wins the race, or one giant shakeup?\n\nCheck the top scoring Notebooks, dear fellow Kagglers:\n<a href=\"https://ibb.co/rwMD43h\"><img src=\"https://i.ibb.co/DRk05YF/delete.png\" alt=\"delete\" border=\"0\"></a>\n\n-Comments deactivated, as the author doesn't want to hear any opinions of his fellow Kagglers.\n-Notebook runs with tracebacks and bugs, but it's supposed to be like this, as \"the fixed version got worse LB\".\n-No LR Schedulers, Checkpoint Savers, Val-Loss-Monitors, etc are used: Simply a fixed number of epochs (855 by trial and error) is chosen and the other hyperparameters (majorly dropout and some merging parameters) are tweaked/optimized for additional overfitting.\n\nChecking the Leaderboard supports our intuition, that many many more than the documented 30 submissions have been made with those notebooks.\n\nSadly the community can't really learn a thing from those Notebooks.\n\nIf someone upvoted those notebooks (except the \"original\" notebooks), please explain to my why you did. I would love to hear and understand.\n\n### What's your opinion? Giant shakeup for valid CV-backed solutions, or overfitters-winning?",
      "votes": 20
    },
    {
      "id": 1037965,
      "postDate": "2020-10-05T12:57:53.447Z",
      "content": "<p>I feel like publishing notebooks with code in an early phase sort of reduces the innovation brought by the community, as everyone is already biased by certain approaches. This was my first Kaggle challenge btw. (+ late joiner) - next time I'll start working on the problem by my own before I have a look at what others do. Also, I was a bit confused how people just copy-paste notebooks and change some values xD</p>",
      "rawMarkdown": "I feel like publishing notebooks with code in an early phase sort of reduces the innovation brought by the community, as everyone is already biased by certain approaches. This was my first Kaggle challenge btw. (+ late joiner) - next time I'll start working on the problem by my own before I have a look at what others do. Also, I was a bit confused how people just copy-paste notebooks and change some values xD",
      "votes": 6
    },
    {
      "id": 1041510,
      "postDate": "2020-10-07T19:32:32.273Z",
      "content": "<p>Thank you for this consideration , hope this will give us all way to avoid overfitting especially for new people competing in Kaggle. I absolutely agree with you, when there is a big corpus of work which is copied from another notebook, there should be a reference to the author.</p>",
      "rawMarkdown": "Thank you for this consideration , hope this will give us all way to avoid overfitting especially for new people competing in Kaggle. I absolutely agree with you, when there is a big corpus of work which is copied from another notebook, there should be a reference to the author.",
      "votes": 3
    },
    {
      "id": 1040537,
      "postDate": "2020-10-07T07:53:26.227Z",
      "content": "<p>Maybe it's a good thing? I mean, this competition is perfectly made to learn some good lessons: 1) work hard on your CV strategy and 2) try your own ideas. Not the classic kaggle competition, but I believe it will end up giving good models </p>",
      "rawMarkdown": "Maybe it's a good thing? I mean, this competition is perfectly made to learn some good lessons: 1) work hard on your CV strategy and 2) try your own ideas. Not the classic kaggle competition, but I believe it will end up giving good models ",
      "votes": 1
    },
    {
      "id": 1040022,
      "postDate": "2020-10-07T00:39:47.470Z",
      "content": "<p>Now, it is clear that overfitters lose…</p>",
      "rawMarkdown": "Now, it is clear that overfitters lose...",
      "votes": 1,
      "replies": [
        {
          "id": 1040464,
          "postDate": "2020-10-07T06:44:02.467Z",
          "content": "<p>Fully agree - sadly not only overfitters loose. Of course we followed our own rules and selected two CV-backed submissions and not our best LB submission.<br>\nSadly we had an imperfect CV-strategy, which then leads to bad results.</p>\n<p>I can only thank encouraging community members like <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and his post <br>\n<a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">Something is wrong…</a></p>",
          "rawMarkdown": "Fully agree - sadly not only overfitters loose. Of course we followed our own rules and selected two CV-backed submissions and not our best LB submission.\nSadly we had an imperfect CV-strategy, which then leads to bad results.\n\nI can only thank encouraging community members like @nischaydnk and his post \n[Something is wrong...](https://www.kaggle.com/nischaydnk)\n"
        }
      ]
    },
    {
      "id": 1038886,
      "postDate": "2020-10-06T06:24:20.310Z",
      "content": "<p>I think we should rely on our CV score rather than focusing on Public LB score. If we try overfitting then definitely there will be a giant shakeup on Private LB </p>",
      "rawMarkdown": "I think we should rely on our CV score rather than focusing on Public LB score. If we try overfitting then definitely there will be a giant shakeup on Private LB ",
      "votes": 1
    },
    {
      "id": 1038216,
      "postDate": "2020-10-05T16:46:26.267Z",
      "content": "<p>Titanic shakeup is impending given what i'm seeing, the top public notebooks (besides a few) are not creative besides \"mloss tuning\" and no innovation has been seen so far. Big shakeup coming.</p>",
      "rawMarkdown": "Titanic shakeup is impending given what i'm seeing, the top public notebooks (besides a few) are not creative besides \"mloss tuning\" and no innovation has been seen so far. Big shakeup coming.",
      "votes": 1
    },
    {
      "id": 1037602,
      "postDate": "2020-10-05T06:56:21.480Z",
      "content": "<p>Last month was spent trying to improve my CV rather than score in the public LB. <br>\nMoreover, I even consider not selecting my current 4th place as final submission. Let's see how that will go ;)</p>",
      "rawMarkdown": "Last month was spent trying to improve my CV rather than score in the public LB. \nMoreover, I even consider not selecting my current 4th place as final submission. Let's see how that will go ;)",
      "votes": 1
    },
    {
      "id": 1037524,
      "postDate": "2020-10-05T05:11:59.623Z",
      "content": "<p>those who are located from silver to bronze, a giant shakeup will be experienced<br>\nwith a small change in hyperparameters, scores change a lot!</p>",
      "rawMarkdown": "those who are located from silver to bronze, a giant shakeup will be experienced\nwith a small change in hyperparameters, scores change a lot!\n\n",
      "votes": 1
    },
    {
      "id": 1038498,
      "postDate": "2020-10-05T20:40:20.020Z",
      "content": "<p>This competition public leaderboard is a perfect trap</p>\n<ol>\n<li>Very small public test set</li>\n<li>Lot of well intentioned initial notebooks were forked, re-forked, , ensembled, re-ensembled  - so much so that it is difficult to find out the original source. The bugs just kept carrying over - the smart people obviously corrected them.</li>\n</ol>\n<p>I just don't trust my current score (88th ) and will probably not select that score. <br>\nBig shakeup coming up</p>",
      "rawMarkdown": "This competition public leaderboard is a perfect trap\n1. Very small public test set\n2. Lot of well intentioned initial notebooks were forked, re-forked, , ensembled, re-ensembled  - so much so that it is difficult to find out the original source. The bugs just kept carrying over - the smart people obviously corrected them.\n\nI just don't trust my current score (88th ) and will probably not select that score. \nBig shakeup coming up",
      "votes": 2
    },
    {
      "id": 1042421,
      "postDate": "2020-10-08T08:10:06.350Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1037965,
      "author_name": "Flogrammer",
      "author_url": "",
      "post_date": "2020-10-05T12:57:53.447000",
      "content": "<p>I feel like publishing notebooks with code in an early phase sort of reduces the innovation brought by the community, as everyone is already biased by certain approaches. This was my first Kaggle challenge btw. (+ late joiner) - next time I'll start working on the problem by my own before I have a look at what others do. Also, I was a bit confused how people just copy-paste notebooks and change some values xD</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1041510,
      "author_name": "Martin",
      "author_url": "",
      "post_date": "2020-10-07T19:32:32.273000",
      "content": "<p>Thank you for this consideration , hope this will give us all way to avoid overfitting especially for new people competing in Kaggle. I absolutely agree with you, when there is a big corpus of work which is copied from another notebook, there should be a reference to the author.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1040537,
      "author_name": "David Castillo",
      "author_url": "",
      "post_date": "2020-10-07T07:53:26.227000",
      "content": "<p>Maybe it's a good thing? I mean, this competition is perfectly made to learn some good lessons: 1) work hard on your CV strategy and 2) try your own ideas. Not the classic kaggle competition, but I believe it will end up giving good models </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1040022,
      "author_name": "wcysysu",
      "author_url": "",
      "post_date": "2020-10-07T00:39:47.470000",
      "content": "<p>Now, it is clear that overfitters lose…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1040464,
          "author_name": "from coffee import *",
          "author_url": "",
          "post_date": "2020-10-07T06:44:02.467000",
          "content": "<p>Fully agree - sadly not only overfitters loose. Of course we followed our own rules and selected two CV-backed submissions and not our best LB submission.<br>\nSadly we had an imperfect CV-strategy, which then leads to bad results.</p>\n<p>I can only thank encouraging community members like <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and his post <br>\n<a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">Something is wrong…</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1038886,
      "author_name": "VAIBHAV MATHUR",
      "author_url": "",
      "post_date": "2020-10-06T06:24:20.310000",
      "content": "<p>I think we should rely on our CV score rather than focusing on Public LB score. If we try overfitting then definitely there will be a giant shakeup on Private LB </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1038216,
      "author_name": "Trigram",
      "author_url": "",
      "post_date": "2020-10-05T16:46:26.267000",
      "content": "<p>Titanic shakeup is impending given what i'm seeing, the top public notebooks (besides a few) are not creative besides \"mloss tuning\" and no innovation has been seen so far. Big shakeup coming.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1037602,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2020-10-05T06:56:21.480000",
      "content": "<p>Last month was spent trying to improve my CV rather than score in the public LB. <br>\nMoreover, I even consider not selecting my current 4th place as final submission. Let's see how that will go ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1037524,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2020-10-05T05:11:59.623000",
      "content": "<p>those who are located from silver to bronze, a giant shakeup will be experienced<br>\nwith a small change in hyperparameters, scores change a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1038498,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-10-05T20:40:20.020000",
      "content": "<p>This competition public leaderboard is a perfect trap</p>\n<ol>\n<li>Very small public test set</li>\n<li>Lot of well intentioned initial notebooks were forked, re-forked, , ensembled, re-ensembled  - so much so that it is difficult to find out the original source. The bugs just kept carrying over - the smart people obviously corrected them.</li>\n</ol>\n<p>I just don't trust my current score (88th ) and will probably not select that score. <br>\nBig shakeup coming up</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1042421,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-08T08:10:06.350000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1037323": "Dear fellow Kagglers,\n\nthe top 15+ scored Notebooks are all derived from one or two Notebooks and have been re-published with minimal changes: \"optimized\" hyperparameters for additional overfitting.\nThis mostly even happend without referencing to the original authors of the Notebooks, which is really bad practice in my opinion --> plagiarism!\n\n### So, what will the private LB bring? Best overfitting wins the race, or one giant shakeup?\n\nCheck the top scoring Notebooks, dear fellow Kagglers:\n<a href=\"https://ibb.co/rwMD43h\"><img src=\"https://i.ibb.co/DRk05YF/delete.png\" alt=\"delete\" border=\"0\"></a>\n\n-Comments deactivated, as the author doesn't want to hear any opinions of his fellow Kagglers.\n-Notebook runs with tracebacks and bugs, but it's supposed to be like this, as \"the fixed version got worse LB\".\n-No LR Schedulers, Checkpoint Savers, Val-Loss-Monitors, etc are used: Simply a fixed number of epochs (855 by trial and error) is chosen and the other hyperparameters (majorly dropout and some merging parameters) are tweaked/optimized for additional overfitting.\n\nChecking the Leaderboard supports our intuition, that many many more than the documented 30 submissions have been made with those notebooks.\n\nSadly the community can't really learn a thing from those Notebooks.\n\nIf someone upvoted those notebooks (except the \"original\" notebooks), please explain to my why you did. I would love to hear and understand.\n\n### What's your opinion? Giant shakeup for valid CV-backed solutions, or overfitters-winning?",
    "1037965": "I feel like publishing notebooks with code in an early phase sort of reduces the innovation brought by the community, as everyone is already biased by certain approaches. This was my first Kaggle challenge btw. (+ late joiner) - next time I'll start working on the problem by my own before I have a look at what others do. Also, I was a bit confused how people just copy-paste notebooks and change some values xD",
    "1041510": "Thank you for this consideration , hope this will give us all way to avoid overfitting especially for new people competing in Kaggle. I absolutely agree with you, when there is a big corpus of work which is copied from another notebook, there should be a reference to the author.",
    "1040537": "Maybe it's a good thing? I mean, this competition is perfectly made to learn some good lessons: 1) work hard on your CV strategy and 2) try your own ideas. Not the classic kaggle competition, but I believe it will end up giving good models ",
    "1040022": "Now, it is clear that overfitters lose...",
    "1038886": "I think we should rely on our CV score rather than focusing on Public LB score. If we try overfitting then definitely there will be a giant shakeup on Private LB ",
    "1038216": "Titanic shakeup is impending given what i'm seeing, the top public notebooks (besides a few) are not creative besides \"mloss tuning\" and no innovation has been seen so far. Big shakeup coming.",
    "1037602": "Last month was spent trying to improve my CV rather than score in the public LB. \nMoreover, I even consider not selecting my current 4th place as final submission. Let's see how that will go ;)",
    "1037524": "those who are located from silver to bronze, a giant shakeup will be experienced\nwith a small change in hyperparameters, scores change a lot!\n\n",
    "1038498": "This competition public leaderboard is a perfect trap\n1. Very small public test set\n2. Lot of well intentioned initial notebooks were forked, re-forked, , ensembled, re-ensembled  - so much so that it is difficult to find out the original source. The bugs just kept carrying over - the smart people obviously corrected them.\n\nI just don't trust my current score (88th ) and will probably not select that score. \nBig shakeup coming up",
    "1042421": ""
  }
}