{
  "id": 189429,
  "title": "Post-mortem dissection & analysis",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/189429",
  "author_name": "",
  "post_date": "2020-10-07T15:18:27.029629500Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>The biggest question is what lead to the shakeup?</strong></p>\n<p>A. <strong><em>En masse</em> forking of public kernels without any real improvements</strong>: All the forking and high-scoring kernels (mostly all, some of the most brilliant kernels such as Ulrich's, Y. Nakama's and Tawara's were actually incredibly robust in shakeup). This in turn leads to LB becoming a house of cards, and LB being defined by who has the strongest CV strategy, one which the public kernels did not try to improve upon.</p>\n<p>B. <strong>The data itself:</strong> CSV metadata was <em>small, noisy and very similar to the Mercedes competition some time ago</em> (:p LB is also similar, with a similarly insane shake).</p>\n<p>C. <strong>The metric</strong>: Laplace log loss was a very good metric if this competition was FVC-only, but the confidence part was probably one of the most easily bypassed methods (one of my teammates set confidence to 275 for all values and predicted only on FVC, and got -6.9 on the public LB) so it was clear there was some sort of acceptable \"range\" for the confidence (and to a lesser extent the FVC, as upon examining preds you get to find out that even the FVC has a similar range albeit with some discrepancies (athletes?), so there was some potential for metric tuning). See <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189301\" target=\"_blank\">this post, which also displays the weaknesses of the metric</a>.</p>\n<p>Metric was prone to probing given that simply using constants could have worked magnificently, and <a href=\"https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data\" target=\"_blank\">it did work for currypurin to extrapolate number of patients</a>.</p>\n<p><strong>Modeling</strong></p>\n<p>So in this competition, it seems like less complex methods (<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189397\" target=\"_blank\">see llu's method of getting very high linear regression score</a>), so perhaps simplicity was the key? Our team also had relatively similar results - an ElasticNet regression would have given us a 22nd position on the leaderboard and same results have been reported [once more, this time by Jagadish Sivakumaran <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189242\" target=\"_blank\">see his solution description here</a>. </p>\n<p><strong>\"Trust in CV\"</strong></p>\n<p>This competition will become an example why the robustness of cross-validation strategy is always necessary - we spent a lot of time on our own validation (which benefited us a fair bit in the shake) - and also why you need to be very careful with the CV strat as a whole, because one chink and you might just tumble down 1000 places on LB.</p>\n<p>To quote 5th place, LukeReijnen:</p>\n<blockquote>\n  <p>Yes, we never looked at out leaderbord score :)</p>\n</blockquote>\n<p>which seems fairly apt to me, as the public LB was a very unreasonable indicator pointed out right from the very start by <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164837\" target=\"_blank\">Andrew Lukyanenko's discussion post, and pointing out that 85 percent of the test data is withheld</a> and we already had a prediction of the shakeup to come ;-) (\"Quite enough for a legendary shakeup 😑\")</p>\n<p>Our strat wasn't exactly trusting <strong>only in CV;</strong> we tried to get as good of a correlation between CV and leaderboard as was possible, and it worked pretty well seeing how we bounced up.</p>\n<p>So hopefully this brief post-mortem might be helpful to all those who come across it, and hopefully serve as a cautionary tale for any future competition which is in a similar scenario.</p>",
  "messages": [
    {
      "id": "1041127",
      "postDate": "10/07/2020 15:18:27",
      "content": "<p><strong>The biggest question is what lead to the shakeup?</strong></p>\n<p>A. <strong><em>En masse</em> forking of public kernels without any real improvements</strong>: All the forking and high-scoring kernels (mostly all, some of the most brilliant kernels such as Ulrich's, Y. Nakama's and Tawara's were actually incredibly robust in shakeup). This in turn leads to LB becoming a house of cards, and LB being defined by who has the strongest CV strategy, one which the public kernels did not try to improve upon.</p>\n<p>B. <strong>The data itself:</strong> CSV metadata was <em>small, noisy and very similar to the Mercedes competition some time ago</em> (:p LB is also similar, with a similarly insane shake).</p>\n<p>C. <strong>The metric</strong>: Laplace log loss was a very good metric if this competition was FVC-only, but the confidence part was probably one of the most easily bypassed methods (one of my teammates set confidence to 275 for all values and predicted only on FVC, and got -6.9 on the public LB) so it was clear there was some sort of acceptable \"range\" for the confidence (and to a lesser extent the FVC, as upon examining preds you get to find out that even the FVC has a similar range albeit with some discrepancies (athletes?), so there was some potential for metric tuning). See <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189301\" target=\"_blank\">this post, which also displays the weaknesses of the metric</a>.</p>\n<p>Metric was prone to probing given that simply using constants could have worked magnificently, and <a href=\"https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data\" target=\"_blank\">it did work for currypurin to extrapolate number of patients</a>.</p>\n<p><strong>Modeling</strong></p>\n<p>So in this competition, it seems like less complex methods (<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189397\" target=\"_blank\">see llu's method of getting very high linear regression score</a>), so perhaps simplicity was the key? Our team also had relatively similar results - an ElasticNet regression would have given us a 22nd position on the leaderboard and same results have been reported [once more, this time by Jagadish Sivakumaran <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189242\" target=\"_blank\">see his solution description here</a>. </p>\n<p><strong>\"Trust in CV\"</strong></p>\n<p>This competition will become an example why the robustness of cross-validation strategy is always necessary - we spent a lot of time on our own validation (which benefited us a fair bit in the shake) - and also why you need to be very careful with the CV strat as a whole, because one chink and you might just tumble down 1000 places on LB.</p>\n<p>To quote 5th place, LukeReijnen:</p>\n<blockquote>\n  <p>Yes, we never looked at out leaderbord score :)</p>\n</blockquote>\n<p>which seems fairly apt to me, as the public LB was a very unreasonable indicator pointed out right from the very start by <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164837\" target=\"_blank\">Andrew Lukyanenko's discussion post, and pointing out that 85 percent of the test data is withheld</a> and we already had a prediction of the shakeup to come ;-) (\"Quite enough for a legendary shakeup 😑\")</p>\n<p>Our strat wasn't exactly trusting <strong>only in CV;</strong> we tried to get as good of a correlation between CV and leaderboard as was possible, and it worked pretty well seeing how we bounced up.</p>\n<p>So hopefully this brief post-mortem might be helpful to all those who come across it, and hopefully serve as a cautionary tale for any future competition which is in a similar scenario.</p>",
      "rawMarkdown": "**The biggest question is what lead to the shakeup?**\n\nA. ***En masse* forking of public kernels without any real improvements**: All the forking and high-scoring kernels (mostly all, some of the most brilliant kernels such as Ulrich's, Y. Nakama's and Tawara's were actually incredibly robust in shakeup). This in turn leads to LB becoming a house of cards, and LB being defined by who has the strongest CV strategy, one which the public kernels did not try to improve upon.\n\nB. **The data itself:** CSV metadata was *small, noisy and very similar to the Mercedes competition some time ago* (:p LB is also similar, with a similarly insane shake).\n\nC. **The metric**: Laplace log loss was a very good metric if this competition was FVC-only, but the confidence part was probably one of the most easily bypassed methods (one of my teammates set confidence to 275 for all values and predicted only on FVC, and got -6.9 on the public LB) so it was clear there was some sort of acceptable \"range\" for the confidence (and to a lesser extent the FVC, as upon examining preds you get to find out that even the FVC has a similar range albeit with some discrepancies (athletes?), so there was some potential for metric tuning). See [this post, which also displays the weaknesses of the metric](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189301).\n\nMetric was prone to probing given that simply using constants could have worked magnificently, and [it did work for currypurin to extrapolate number of patients](https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data).\n\n**Modeling**\n\nSo in this competition, it seems like less complex methods ([see llu's method of getting very high linear regression score](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189397)), so perhaps simplicity was the key? Our team also had relatively similar results - an ElasticNet regression would have given us a 22nd position on the leaderboard and same results have been reported [once more, this time by Jagadish Sivakumaran [see his solution description here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189242). \n\n**\"Trust in CV\"**\n\nThis competition will become an example why the robustness of cross-validation strategy is always necessary - we spent a lot of time on our own validation (which benefited us a fair bit in the shake) - and also why you need to be very careful with the CV strat as a whole, because one chink and you might just tumble down 1000 places on LB.\n\nTo quote 5th place, LukeReijnen:\n> Yes, we never looked at out leaderbord score :)\n\nwhich seems fairly apt to me, as the public LB was a very unreasonable indicator pointed out right from the very start by [Andrew Lukyanenko's discussion post, and pointing out that 85 percent of the test data is withheld](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164837) and we already had a prediction of the shakeup to come ;-) (\"Quite enough for a legendary shakeup 😑\")\n\nOur strat wasn't exactly trusting **only in CV;** we tried to get as good of a correlation between CV and leaderboard as was possible, and it worked pretty well seeing how we bounced up.\n\nSo hopefully this brief post-mortem might be helpful to all those who come across it, and hopefully serve as a cautionary tale for any future competition which is in a similar scenario.",
      "votes": null
    },
    {
      "id": "1041140",
      "postDate": "10/07/2020 15:23:18",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nxrprime\" target=\"_blank\">@nxrprime</a> thanks for sharing this detailed info this will definitely be useful for future competitions similar to this as well </p>",
      "rawMarkdown": "Hey @nxrprime thanks for sharing this detailed info this will definitely be useful for future competitions similar to this as well",
      "votes": null
    },
    {
      "id": "1041472",
      "postDate": "10/07/2020 18:57:52",
      "content": "<p>We scored a lot better on CV all the time, so we thought that the LB dataset was not representative and couldn't be bothered to check models anymore</p>",
      "rawMarkdown": "We scored a lot better on CV all the time, so we thought that the LB dataset was not representative and couldn't be bothered to check models anymore",
      "votes": null
    },
    {
      "id": "1041473",
      "postDate": "10/07/2020 18:58:33",
      "content": "<p>And it's Reijnen, not Rejinen❤️</p>",
      "rawMarkdown": "And it's Reijnen, not Rejinen❤️",
      "votes": null
    },
    {
      "id": "1041974",
      "postDate": "10/08/2020 02:05:42",
      "content": "<p>Ah sorry for the typo.</p>",
      "rawMarkdown": "Ah sorry for the typo.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1041140,
      "author_name": "vaibhavmathur96",
      "author_url": "",
      "post_date": "10/07/2020 15:23:18",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nxrprime\" target=\"_blank\">@nxrprime</a> thanks for sharing this detailed info this will definitely be useful for future competitions similar to this as well </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1041472,
      "author_name": "lukereijnen",
      "author_url": "",
      "post_date": "10/07/2020 18:57:52",
      "content": "<p>We scored a lot better on CV all the time, so we thought that the LB dataset was not representative and couldn't be bothered to check models anymore</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041473,
          "author_name": "lukereijnen",
          "author_url": "",
          "post_date": "10/07/2020 18:58:33",
          "content": "<p>And it's Reijnen, not Rejinen❤️</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041974,
          "author_name": "nxrprime",
          "author_url": "",
          "post_date": "10/08/2020 02:05:42",
          "content": "<p>Ah sorry for the typo.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1041127": "**The biggest question is what lead to the shakeup?**\n\nA. ***En masse* forking of public kernels without any real improvements**: All the forking and high-scoring kernels (mostly all, some of the most brilliant kernels such as Ulrich's, Y. Nakama's and Tawara's were actually incredibly robust in shakeup). This in turn leads to LB becoming a house of cards, and LB being defined by who has the strongest CV strategy, one which the public kernels did not try to improve upon.\n\nB. **The data itself:** CSV metadata was *small, noisy and very similar to the Mercedes competition some time ago* (:p LB is also similar, with a similarly insane shake).\n\nC. **The metric**: Laplace log loss was a very good metric if this competition was FVC-only, but the confidence part was probably one of the most easily bypassed methods (one of my teammates set confidence to 275 for all values and predicted only on FVC, and got -6.9 on the public LB) so it was clear there was some sort of acceptable \"range\" for the confidence (and to a lesser extent the FVC, as upon examining preds you get to find out that even the FVC has a similar range albeit with some discrepancies (athletes?), so there was some potential for metric tuning). See [this post, which also displays the weaknesses of the metric](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189301).\n\nMetric was prone to probing given that simply using constants could have worked magnificently, and [it did work for currypurin to extrapolate number of patients](https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data).\n\n**Modeling**\n\nSo in this competition, it seems like less complex methods ([see llu's method of getting very high linear regression score](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189397)), so perhaps simplicity was the key? Our team also had relatively similar results - an ElasticNet regression would have given us a 22nd position on the leaderboard and same results have been reported [once more, this time by Jagadish Sivakumaran [see his solution description here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189242). \n\n**\"Trust in CV\"**\n\nThis competition will become an example why the robustness of cross-validation strategy is always necessary - we spent a lot of time on our own validation (which benefited us a fair bit in the shake) - and also why you need to be very careful with the CV strat as a whole, because one chink and you might just tumble down 1000 places on LB.\n\nTo quote 5th place, LukeReijnen:\n> Yes, we never looked at out leaderbord score :)\n\nwhich seems fairly apt to me, as the public LB was a very unreasonable indicator pointed out right from the very start by [Andrew Lukyanenko's discussion post, and pointing out that 85 percent of the test data is withheld](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164837) and we already had a prediction of the shakeup to come ;-) (\"Quite enough for a legendary shakeup 😑\")\n\nOur strat wasn't exactly trusting **only in CV;** we tried to get as good of a correlation between CV and leaderboard as was possible, and it worked pretty well seeing how we bounced up.\n\nSo hopefully this brief post-mortem might be helpful to all those who come across it, and hopefully serve as a cautionary tale for any future competition which is in a similar scenario.",
    "1041140": "Hey @nxrprime thanks for sharing this detailed info this will definitely be useful for future competitions similar to this as well",
    "1041472": "We scored a lot better on CV all the time, so we thought that the LB dataset was not representative and couldn't be bothered to check models anymore",
    "1041473": "And it's Reijnen, not Rejinen❤️",
    "1041974": "Ah sorry for the typo."
  },
  "source": "meta"
}