{
  "id": 399318,
  "title": "Stuck with 0.095 Score [Solved]",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/399318",
  "author_name": "",
  "post_date": "2023-04-03T15:11:56.279866100Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello fellow Kagglers!</p>\n<p>I have been stuck for a long while with this score. Isn't this the baseline score? I wonder if I am doing some fundamental mistake somewhere. Any tips, modifications would be immensely appreciated!</p>\n<p><a href=\"https://www.kaggle.com/code/sireesh/fog-lab1?kernelSessionId=124213568\" target=\"_blank\">Here</a> is the code! (Sorry for the long comment blocks).</p>\n<p>Since I haven't mentioned in the code, I would like to state here that I took the code for merging defog and tdcsfog data from <a href=\"https://www.kaggle.com/kimtaehun\" target=\"_blank\">@kimtaehun</a> mentioned <a href=\"https://www.kaggle.com/code/kimtaehun/simple-lgbm-multi-class-classification-baseline\" target=\"_blank\">here</a> and the submission part from <a href=\"https://www.kaggle.com/ammarnassanalhajali\" target=\"_blank\">@ammarnassanalhajali</a> </p>\n<p>Maybe I need to go for a long walk :')</p>\n<p>TIA and happy kaggling!</p>\n<p>Update:</p>\n<p>As pointed out by <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> below, I was committing a blunder with my test prediction loop with the dataframe concatenation. i just updated it by matching it with 'submission.csv' output code block by <a href=\"https://www.kaggle.com/ammarnassanalhajali\" target=\"_blank\">@ammarnassanalhajali</a> as mentioned  <a href=\"https://www.kaggle.com/code/ammarnassanalhajali/freezing-of-gait-prediction/notebook\" target=\"_blank\">here</a> and it works! Thanks for all the help and suggestions!</p>",
  "messages": [
    {
      "id": "2207670",
      "postDate": "04/03/2023 15:11:56",
      "content": "<p>Hello fellow Kagglers!</p>\n<p>I have been stuck for a long while with this score. Isn't this the baseline score? I wonder if I am doing some fundamental mistake somewhere. Any tips, modifications would be immensely appreciated!</p>\n<p><a href=\"https://www.kaggle.com/code/sireesh/fog-lab1?kernelSessionId=124213568\" target=\"_blank\">Here</a> is the code! (Sorry for the long comment blocks).</p>\n<p>Since I haven't mentioned in the code, I would like to state here that I took the code for merging defog and tdcsfog data from <a href=\"https://www.kaggle.com/kimtaehun\" target=\"_blank\">@kimtaehun</a> mentioned <a href=\"https://www.kaggle.com/code/kimtaehun/simple-lgbm-multi-class-classification-baseline\" target=\"_blank\">here</a> and the submission part from <a href=\"https://www.kaggle.com/ammarnassanalhajali\" target=\"_blank\">@ammarnassanalhajali</a> </p>\n<p>Maybe I need to go for a long walk :')</p>\n<p>TIA and happy kaggling!</p>\n<p>Update:</p>\n<p>As pointed out by <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> below, I was committing a blunder with my test prediction loop with the dataframe concatenation. i just updated it by matching it with 'submission.csv' output code block by <a href=\"https://www.kaggle.com/ammarnassanalhajali\" target=\"_blank\">@ammarnassanalhajali</a> as mentioned  <a href=\"https://www.kaggle.com/code/ammarnassanalhajali/freezing-of-gait-prediction/notebook\" target=\"_blank\">here</a> and it works! Thanks for all the help and suggestions!</p>",
      "rawMarkdown": "Hello fellow Kagglers!\n\nI have been stuck for a long while with this score. Isn't this the baseline score? I wonder if I am doing some fundamental mistake somewhere. Any tips, modifications would be immensely appreciated!\n\n[Here](https://www.kaggle.com/code/sireesh/fog-lab1?kernelSessionId=124213568) is the code! (Sorry for the long comment blocks).\n\nSince I haven't mentioned in the code, I would like to state here that I took the code for merging defog and tdcsfog data from @kimtaehun mentioned [here](https://www.kaggle.com/code/kimtaehun/simple-lgbm-multi-class-classification-baseline) and the submission part from @ammarnassanalhajali \n\nMaybe I need to go for a long walk :')\n\nTIA and happy kaggling!\n\n\nUpdate:\n\nAs pointed out by @jsday96 below, I was committing a blunder with my test prediction loop with the dataframe concatenation. i just updated it by matching it with 'submission.csv' output code block by @ammarnassanalhajali as mentioned  [here](https://www.kaggle.com/code/ammarnassanalhajali/freezing-of-gait-prediction/notebook) and it works! Thanks for all the help and suggestions!",
      "votes": null
    },
    {
      "id": "2209781",
      "postDate": "04/04/2023 22:48:53",
      "content": "<p>I haven't checked your code but if it's correct i would advise you to check confusion matrix of you predicted values. Your model might not predict any fogs because of data imbalance</p>",
      "rawMarkdown": "I haven't checked your code but if it's correct i would advise you to check confusion matrix of you predicted values. Your model might not predict any fogs because of data imbalance",
      "votes": null
    },
    {
      "id": "2209868",
      "postDate": "04/05/2023 02:06:35",
      "content": "<p>Thanks for the tip! Sounds like a legit problem with this dataset! I will check it! Did you find this issue in yours?</p>",
      "rawMarkdown": "Thanks for the tip! Sounds like a legit problem with this dataset! I will check it! Did you find this issue in yours?",
      "votes": null
    },
    {
      "id": "2209918",
      "postDate": "04/05/2023 03:12:50",
      "content": "<p>I just took a brief look at your code. It appears there's a bug it which it is hardcoded to form the \"final_pred\" dataframe while processing the second input file (\"k==1\"). That works fine on the small publicly visible sample test dataset but is likely to be a major problem on the secret test dataset used to score submissions, which is much larger (probably a hundred or so files). Normally I would expect a bug like that to cause a submission scoring error, but it appears you're doing something weird to merge the predictions dataframe with a larger dataframe that has all the correct rows (all the IDs but no predictions) and fill in the missing predictions with zeros (\".fillna(0.0)\"). So, the end result is probably a file that's mostly full of zeros, just like the sample submission. That explains why your score is effectively identical to the sample submission score.</p>\n<p>As a side note, I also found it strange that you're using MSE loss and relu activations in the final layer of your model. That's very non-standard for a classification problem like this. I doubt that's the main reason you're stuck at 0.095, but might be beneficial for you to experiment with alternatives. </p>",
      "rawMarkdown": "I just took a brief look at your code. It appears there's a bug it which it is hardcoded to form the \"final_pred\" dataframe while processing the second input file (\"k==1\"). That works fine on the small publicly visible sample test dataset but is likely to be a major problem on the secret test dataset used to score submissions, which is much larger (probably a hundred or so files). Normally I would expect a bug like that to cause a submission scoring error, but it appears you're doing something weird to merge the predictions dataframe with a larger dataframe that has all the correct rows (all the IDs but no predictions) and fill in the missing predictions with zeros (\".fillna(0.0)\"). So, the end result is probably a file that's mostly full of zeros, just like the sample submission. That explains why your score is effectively identical to the sample submission score.\n\nAs a side note, I also found it strange that you're using MSE loss and relu activations in the final layer of your model. That's very non-standard for a classification problem like this. I doubt that's the main reason you're stuck at 0.095, but might be beneficial for you to experiment with alternatives.",
      "votes": null
    },
    {
      "id": "2220206",
      "postDate": "04/13/2023 07:35:01",
      "content": "<p>Hey James! Thanks a lot for taking out time to go through the code! You were right, there was something wrong with my prediction merging and dataframe concat. I made changes and it is working now! </p>\n<p>Regarding the MSE loss and relu activations, i totally understand. I saw a lot of people using regression techniques instead of going for classification strategies for this comp which I wanted to try out as well. Turns out they can be at least better than sample predictions in the public LB as of right now. i guess it has more to do with the metric Avg Precision Score.</p>",
      "rawMarkdown": "Hey James! Thanks a lot for taking out time to go through the code! You were right, there was something wrong with my prediction merging and dataframe concat. I made changes and it is working now! \n\nRegarding the MSE loss and relu activations, i totally understand. I saw a lot of people using regression techniques instead of going for classification strategies for this comp which I wanted to try out as well. Turns out they can be at least better than sample predictions in the public LB as of right now. i guess it has more to do with the metric Avg Precision Score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2209781,
      "author_name": "julienlambertdia5t09",
      "author_url": "",
      "post_date": "04/04/2023 22:48:53",
      "content": "<p>I haven't checked your code but if it's correct i would advise you to check confusion matrix of you predicted values. Your model might not predict any fogs because of data imbalance</p>",
      "votes": null,
      "replies": [
        {
          "id": 2209868,
          "author_name": "sireesh",
          "author_url": "",
          "post_date": "04/05/2023 02:06:35",
          "content": "<p>Thanks for the tip! Sounds like a legit problem with this dataset! I will check it! Did you find this issue in yours?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2209918,
      "author_name": "jsday96",
      "author_url": "",
      "post_date": "04/05/2023 03:12:50",
      "content": "<p>I just took a brief look at your code. It appears there's a bug it which it is hardcoded to form the \"final_pred\" dataframe while processing the second input file (\"k==1\"). That works fine on the small publicly visible sample test dataset but is likely to be a major problem on the secret test dataset used to score submissions, which is much larger (probably a hundred or so files). Normally I would expect a bug like that to cause a submission scoring error, but it appears you're doing something weird to merge the predictions dataframe with a larger dataframe that has all the correct rows (all the IDs but no predictions) and fill in the missing predictions with zeros (\".fillna(0.0)\"). So, the end result is probably a file that's mostly full of zeros, just like the sample submission. That explains why your score is effectively identical to the sample submission score.</p>\n<p>As a side note, I also found it strange that you're using MSE loss and relu activations in the final layer of your model. That's very non-standard for a classification problem like this. I doubt that's the main reason you're stuck at 0.095, but might be beneficial for you to experiment with alternatives. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2220206,
          "author_name": "sireesh",
          "author_url": "",
          "post_date": "04/13/2023 07:35:01",
          "content": "<p>Hey James! Thanks a lot for taking out time to go through the code! You were right, there was something wrong with my prediction merging and dataframe concat. I made changes and it is working now! </p>\n<p>Regarding the MSE loss and relu activations, i totally understand. I saw a lot of people using regression techniques instead of going for classification strategies for this comp which I wanted to try out as well. Turns out they can be at least better than sample predictions in the public LB as of right now. i guess it has more to do with the metric Avg Precision Score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2207670": "Hello fellow Kagglers!\n\nI have been stuck for a long while with this score. Isn't this the baseline score? I wonder if I am doing some fundamental mistake somewhere. Any tips, modifications would be immensely appreciated!\n\n[Here](https://www.kaggle.com/code/sireesh/fog-lab1?kernelSessionId=124213568) is the code! (Sorry for the long comment blocks).\n\nSince I haven't mentioned in the code, I would like to state here that I took the code for merging defog and tdcsfog data from @kimtaehun mentioned [here](https://www.kaggle.com/code/kimtaehun/simple-lgbm-multi-class-classification-baseline) and the submission part from @ammarnassanalhajali \n\nMaybe I need to go for a long walk :')\n\nTIA and happy kaggling!\n\n\nUpdate:\n\nAs pointed out by @jsday96 below, I was committing a blunder with my test prediction loop with the dataframe concatenation. i just updated it by matching it with 'submission.csv' output code block by @ammarnassanalhajali as mentioned  [here](https://www.kaggle.com/code/ammarnassanalhajali/freezing-of-gait-prediction/notebook) and it works! Thanks for all the help and suggestions!",
    "2209781": "I haven't checked your code but if it's correct i would advise you to check confusion matrix of you predicted values. Your model might not predict any fogs because of data imbalance",
    "2209868": "Thanks for the tip! Sounds like a legit problem with this dataset! I will check it! Did you find this issue in yours?",
    "2209918": "I just took a brief look at your code. It appears there's a bug it which it is hardcoded to form the \"final_pred\" dataframe while processing the second input file (\"k==1\"). That works fine on the small publicly visible sample test dataset but is likely to be a major problem on the secret test dataset used to score submissions, which is much larger (probably a hundred or so files). Normally I would expect a bug like that to cause a submission scoring error, but it appears you're doing something weird to merge the predictions dataframe with a larger dataframe that has all the correct rows (all the IDs but no predictions) and fill in the missing predictions with zeros (\".fillna(0.0)\"). So, the end result is probably a file that's mostly full of zeros, just like the sample submission. That explains why your score is effectively identical to the sample submission score.\n\nAs a side note, I also found it strange that you're using MSE loss and relu activations in the final layer of your model. That's very non-standard for a classification problem like this. I doubt that's the main reason you're stuck at 0.095, but might be beneficial for you to experiment with alternatives.",
    "2220206": "Hey James! Thanks a lot for taking out time to go through the code! You were right, there was something wrong with my prediction merging and dataframe concat. I made changes and it is working now! \n\nRegarding the MSE loss and relu activations, i totally understand. I saw a lot of people using regression techniques instead of going for classification strategies for this comp which I wanted to try out as well. Turns out they can be at least better than sample predictions in the public LB as of right now. i guess it has more to do with the metric Avg Precision Score."
  },
  "source": "meta"
}