{
  "id": 416074,
  "title": "Lessons from dropping 600 ranks in the LB: Trust CV and Ignore LB",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/416074",
  "author_name": "",
  "post_date": "2023-06-09T12:02:32.544210700Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations to the winners!</p>\n<p>This has been the most educational challenge I’ve been a part of. 143 submissions later, let’s breakdown where we went wrong.</p>\n<p>The task here was to predict the occurrence of Parkinson’s FoG events in time series data. The data supplied was significantly imbalanced, which was to be expected, with the annotated datasets exhibiting an overwhelming number of cases where nothing occurs. Didn't have time for significant EDA, which was a mistake. I have a suspicion, based on discussion threads, that some intelligence on data source selection may have helped clear up my problems.</p>\n<p>I joined quite late, with a month left, and started by forking the public LGBM MultiOutputRegressor notebook (thanks <a href=\"https://www.kaggle.com/nicholasgray1\" target=\"_blank\">@nicholasgray1</a> ) which gave me a public LB score of 0.311. </p>\n<p>I was having significant problems getting the CV and public LB values to make sense, a problem was quite common.  I understand the mantra of \"Trust CV\", but never previously encountered a challenge where no correspondence occurred at all. More specifically, improvements in CV were not corresponding to improvements in LB. Ended up even checking if the correspondence between train and test (well, public LB) via t-tests as a sanity check. This applies for all of the following methods.</p>\n<pre><code>Undersampling\nStratification\nFeature Selection\nusing group means variances.\nVarious pandas-derived statistical features\nStep Rate features\nAlternative models\n</code></pre>\n<p>The only method that did work to improve public LB score (0.317, private 0.229 ) was the using a IsotonicRegressor and a hold-out dataset to calibrate the outputs into probabilities. I got that idea from this <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638\" target=\"_blank\">thread</a>. Inherently made sense, as the original regression notebook utilized clipping, and hence wasnt generating probabilities. I have a suspicion that post-calibration, CV values made more sense, but this needs further study.</p>\n<p>I also had a failure to get my TabNetRegressor submission to work with prediction calibration, it would run but timeout during submission. Unsure what the problem is there. Without calibration, the TabNetRegressor still exhibited superior performance out of the box in the private LB (0.278, 0.237).</p>\n<p>So what actually worked well? Looking at the private LB scores, I can conclude the following did help private LB performance:</p>\n<pre><code>     .\n   \n    \n     \n   (Nice work  )\n</code></pre>\n<p>To conclude, trust CV, discard the public LB, and spend more time performing EDA and gaining domain knowledge.</p>",
  "messages": [
    {
      "id": "2293685",
      "postDate": "06/09/2023 12:02:32",
      "content": "<p>Congratulations to the winners!</p>\n<p>This has been the most educational challenge I’ve been a part of. 143 submissions later, let’s breakdown where we went wrong.</p>\n<p>The task here was to predict the occurrence of Parkinson’s FoG events in time series data. The data supplied was significantly imbalanced, which was to be expected, with the annotated datasets exhibiting an overwhelming number of cases where nothing occurs. Didn't have time for significant EDA, which was a mistake. I have a suspicion, based on discussion threads, that some intelligence on data source selection may have helped clear up my problems.</p>\n<p>I joined quite late, with a month left, and started by forking the public LGBM MultiOutputRegressor notebook (thanks <a href=\"https://www.kaggle.com/nicholasgray1\" target=\"_blank\">@nicholasgray1</a> ) which gave me a public LB score of 0.311. </p>\n<p>I was having significant problems getting the CV and public LB values to make sense, a problem was quite common.  I understand the mantra of \"Trust CV\", but never previously encountered a challenge where no correspondence occurred at all. More specifically, improvements in CV were not corresponding to improvements in LB. Ended up even checking if the correspondence between train and test (well, public LB) via t-tests as a sanity check. This applies for all of the following methods.</p>\n<pre><code>Undersampling\nStratification\nFeature Selection\nusing group means variances.\nVarious pandas-derived statistical features\nStep Rate features\nAlternative models\n</code></pre>\n<p>The only method that did work to improve public LB score (0.317, private 0.229 ) was the using a IsotonicRegressor and a hold-out dataset to calibrate the outputs into probabilities. I got that idea from this <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638\" target=\"_blank\">thread</a>. Inherently made sense, as the original regression notebook utilized clipping, and hence wasnt generating probabilities. I have a suspicion that post-calibration, CV values made more sense, but this needs further study.</p>\n<p>I also had a failure to get my TabNetRegressor submission to work with prediction calibration, it would run but timeout during submission. Unsure what the problem is there. Without calibration, the TabNetRegressor still exhibited superior performance out of the box in the private LB (0.278, 0.237).</p>\n<p>So what actually worked well? Looking at the private LB scores, I can conclude the following did help private LB performance:</p>\n<pre><code>     .\n   \n    \n     \n   (Nice work  )\n</code></pre>\n<p>To conclude, trust CV, discard the public LB, and spend more time performing EDA and gaining domain knowledge.</p>",
      "rawMarkdown": "Congratulations to the winners!\n\nThis has been the most educational challenge I’ve been a part of. 143 submissions later, let’s breakdown where we went wrong.\n\nThe task here was to predict the occurrence of Parkinson’s FoG events in time series data. The data supplied was significantly imbalanced, which was to be expected, with the annotated datasets exhibiting an overwhelming number of cases where nothing occurs. Didn't have time for significant EDA, which was a mistake. I have a suspicion, based on discussion threads, that some intelligence on data source selection may have helped clear up my problems.\n\nI joined quite late, with a month left, and started by forking the public LGBM MultiOutputRegressor notebook (thanks @nicholasgray1 ) which gave me a public LB score of 0.311. \n\n\nI was having significant problems getting the CV and public LB values to make sense, a problem was quite common.  I understand the mantra of \"Trust CV\", but never previously encountered a challenge where no correspondence occurred at all. More specifically, improvements in CV were not corresponding to improvements in LB. Ended up even checking if the correspondence between train and test (well, public LB) via t-tests as a sanity check. This applies for all of the following methods.\n\n    Undersampling\n    Stratification\n    Feature Selection\n    Normalisation using different group means and variances.\n    Various pandas-derived and statistical features\n    Step Rate features\n    Alternative models\n\n\nThe only method that did work to improve public LB score (0.317, private 0.229 ) was the using a IsotonicRegressor and a hold-out dataset to calibrate the outputs into probabilities. I got that idea from this [thread](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638). Inherently made sense, as the original regression notebook utilized clipping, and hence wasnt generating probabilities. I have a suspicion that post-calibration, CV values made more sense, but this needs further study.\n\nI also had a failure to get my TabNetRegressor submission to work with prediction calibration, it would run but timeout during submission. Unsure what the problem is there. Without calibration, the TabNetRegressor still exhibited superior performance out of the box in the private LB (0.278, 0.237).\n    \nSo what actually worked well? Looking at the private LB scores, I can conclude the following did help private LB performance:\n\n    Calibration of Regressor Predictions to outputs.\n    Catboost and Tabnet models\n    Changes in active window size\n    Certain statistical features and feature permutations\n    Step rate feature (Nice work @vrbaryshev )\n\nTo conclude, trust CV, discard the public LB, and spend more time performing EDA and gaining domain knowledge.",
      "votes": null
    },
    {
      "id": "2293706",
      "postDate": "06/09/2023 12:22:21",
      "content": "<p>StratifiedGroupKFold with 5 folds and group = Subject worked nice for me (my place went down just by 5).</p>",
      "rawMarkdown": "StratifiedGroupKFold with 5 folds and group = Subject worked nice for me (my place went down just by 5).",
      "votes": null
    },
    {
      "id": "2293715",
      "postDate": "06/09/2023 12:28:52",
      "content": "<p>Nice going, yeah my stratified notebooks performed better in private. Did you observe the same CV-LB inconsistency?</p>",
      "rawMarkdown": "Nice going, yeah my stratified notebooks performed better in private. Did you observe the same CV-LB inconsistency?",
      "votes": null
    },
    {
      "id": "2293730",
      "postDate": "06/09/2023 12:42:29",
      "content": "<p>In fact, my CV was consistent with my LB (I mean that increase in my CV resulted in increase of my public LB standing). For example, my winning solution has CV of 0.297 with private LB of 0.328, and it had the best CV among all my submissions.</p>",
      "rawMarkdown": "In fact, my CV was consistent with my LB (I mean that increase in my CV resulted in increase of my public LB standing). For example, my winning solution has CV of 0.297 with private LB of 0.328, and it had the best CV among all my submissions.",
      "votes": null
    },
    {
      "id": "2293739",
      "postDate": "06/09/2023 12:49:22",
      "content": "<p>Did you use all three datasets?</p>",
      "rawMarkdown": "Did you use all three datasets?",
      "votes": null
    },
    {
      "id": "2293742",
      "postDate": "06/09/2023 12:52:43",
      "content": "<p>I used <code>defog</code> and <code>tdcsfog</code> together. I haven't used <code>unlabeled</code>.</p>",
      "rawMarkdown": "I used `defog` and `tdcsfog` together. I haven't used `unlabeled`.",
      "votes": null
    },
    {
      "id": "2293746",
      "postDate": "06/09/2023 12:58:04",
      "content": "<p>Would you mind sharing a notebook? Doesn't have to be your best solution, just want to check the CV vs LB</p>",
      "rawMarkdown": "Would you mind sharing a notebook? Doesn't have to be your best solution, just want to check the CV vs LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2293706,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "06/09/2023 12:22:21",
      "content": "<p>StratifiedGroupKFold with 5 folds and group = Subject worked nice for me (my place went down just by 5).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2293715,
          "author_name": "exjustice",
          "author_url": "",
          "post_date": "06/09/2023 12:28:52",
          "content": "<p>Nice going, yeah my stratified notebooks performed better in private. Did you observe the same CV-LB inconsistency?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2293730,
              "author_name": "atamazian",
              "author_url": "",
              "post_date": "06/09/2023 12:42:29",
              "content": "<p>In fact, my CV was consistent with my LB (I mean that increase in my CV resulted in increase of my public LB standing). For example, my winning solution has CV of 0.297 with private LB of 0.328, and it had the best CV among all my submissions.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2293739,
                  "author_name": "exjustice",
                  "author_url": "",
                  "post_date": "06/09/2023 12:49:22",
                  "content": "<p>Did you use all three datasets?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2293742,
                      "author_name": "atamazian",
                      "author_url": "",
                      "post_date": "06/09/2023 12:52:43",
                      "content": "<p>I used <code>defog</code> and <code>tdcsfog</code> together. I haven't used <code>unlabeled</code>.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2293746,
                          "author_name": "exjustice",
                          "author_url": "",
                          "post_date": "06/09/2023 12:58:04",
                          "content": "<p>Would you mind sharing a notebook? Doesn't have to be your best solution, just want to check the CV vs LB</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2293685": "Congratulations to the winners!\n\nThis has been the most educational challenge I’ve been a part of. 143 submissions later, let’s breakdown where we went wrong.\n\nThe task here was to predict the occurrence of Parkinson’s FoG events in time series data. The data supplied was significantly imbalanced, which was to be expected, with the annotated datasets exhibiting an overwhelming number of cases where nothing occurs. Didn't have time for significant EDA, which was a mistake. I have a suspicion, based on discussion threads, that some intelligence on data source selection may have helped clear up my problems.\n\nI joined quite late, with a month left, and started by forking the public LGBM MultiOutputRegressor notebook (thanks @nicholasgray1 ) which gave me a public LB score of 0.311. \n\n\nI was having significant problems getting the CV and public LB values to make sense, a problem was quite common.  I understand the mantra of \"Trust CV\", but never previously encountered a challenge where no correspondence occurred at all. More specifically, improvements in CV were not corresponding to improvements in LB. Ended up even checking if the correspondence between train and test (well, public LB) via t-tests as a sanity check. This applies for all of the following methods.\n\n    Undersampling\n    Stratification\n    Feature Selection\n    Normalisation using different group means and variances.\n    Various pandas-derived and statistical features\n    Step Rate features\n    Alternative models\n\n\nThe only method that did work to improve public LB score (0.317, private 0.229 ) was the using a IsotonicRegressor and a hold-out dataset to calibrate the outputs into probabilities. I got that idea from this [thread](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638). Inherently made sense, as the original regression notebook utilized clipping, and hence wasnt generating probabilities. I have a suspicion that post-calibration, CV values made more sense, but this needs further study.\n\nI also had a failure to get my TabNetRegressor submission to work with prediction calibration, it would run but timeout during submission. Unsure what the problem is there. Without calibration, the TabNetRegressor still exhibited superior performance out of the box in the private LB (0.278, 0.237).\n    \nSo what actually worked well? Looking at the private LB scores, I can conclude the following did help private LB performance:\n\n    Calibration of Regressor Predictions to outputs.\n    Catboost and Tabnet models\n    Changes in active window size\n    Certain statistical features and feature permutations\n    Step rate feature (Nice work @vrbaryshev )\n\nTo conclude, trust CV, discard the public LB, and spend more time performing EDA and gaining domain knowledge.",
    "2293706": "StratifiedGroupKFold with 5 folds and group = Subject worked nice for me (my place went down just by 5).",
    "2293715": "Nice going, yeah my stratified notebooks performed better in private. Did you observe the same CV-LB inconsistency?",
    "2293730": "In fact, my CV was consistent with my LB (I mean that increase in my CV resulted in increase of my public LB standing). For example, my winning solution has CV of 0.297 with private LB of 0.328, and it had the best CV among all my submissions.",
    "2293739": "Did you use all three datasets?",
    "2293742": "I used `defog` and `tdcsfog` together. I haven't used `unlabeled`.",
    "2293746": "Would you mind sharing a notebook? Doesn't have to be your best solution, just want to check the CV vs LB"
  },
  "source": "meta"
}