{
  "id": 398720,
  "title": "Yet another thread for 'Submission Scoring Error' [Solved]",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/398720",
  "author_name": "Sireesh Limbu",
  "post_date": "2023-03-31T12:22:50.234000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello fellow Kagglers!</p>\n<p>Sorry for another post about submission error and such. I have been stuck for a while trying to figure out why my notebook gives this error (Notebook Threw Exception). After going through a few posts on similar error, I made all the changes but I still couldn't point the error here. My output (head) looks like this:<br>\n.               Id            StartHesitation   Turn       Walking<br>\n0    003f117e14_0    0.0              0.002          0.0<br>\n1    003f117e14_1            0.0              0.002          0.0<br>\n2    003f117e14_2            0.0              0.002          0.0<br>\n3    003f117e14_3            0.0              0.002          0.0<br>\n4    003f117e14_4    0.0              0.002          0.0</p>\n<p>and tail like this</p>\n<p>.                          Id           StartHesitation   Turn        Walking<br>\n286365    02ab235146_281683   0.0             0.0      0.0<br>\n286366    02ab235146_281684   0.0             0.0      0.0<br>\n286367    02ab235146_281685   0.0               0.0        0.0<br>\n286368    02ab235146_281686   0.0             0.0      0.0<br>\n286369    02ab235146_281687   0.0             0.0      0.0</p>\n<p>Is it because the column types are float and not int because I thought that was no issue. Also the the decimal point limit is just 3 which worked in all cases in the posts I went through. None of the values exceed 0.9.</p>\n<p>Any sort of idea/tip would be greatly appreciated!</p>\n<p>TIA</p>\n<p>Update 1 : I made the test source more generic after <a href=\"https://www.kaggle.com/avivlevi815\" target=\"_blank\">@avivlevi815</a> pointed it out so that even if ID changes, they can be processed. Now I am getting 'Submission Scoring Error'. I have checked my number of rows and columns are correct.  And I rounded my probabilities to 3 after going through the post on 'Submission Scoring Error' although the metric has been updated to cater for this issue. The values are between 0 and 1. No row has more than one column with non zero value.</p>\n<p>Update 2 [Solved]:</p>\n<p>So few things that I did to rectify the error : </p>\n<ol>\n<li><p>The update 1 mentioned above to use a generic code that uses file with any name.csv inside the test folder that Kaggle tests from their end</p></li>\n<li><p>For some reason I had not made use of sample_submission.csv file to create the submission.csv. I just looked at the structure and tried to imitate it. I copied the small following code from <a href=\"https://www.kaggle.com/ammarnassanalhajali\" target=\"_blank\">@ammarnassanalhajali</a> to use submission_sample.csv while creating the submission file. As <a href=\"https://www.kaggle.com/avivlevi815\" target=\"_blank\">@avivlevi815</a> rightly mentioned that both sample_submission.csv and test files are change in Kaggles backend for scoring. Also, index = False.</p>\n<pre><code>sub = pd.read_csv(p+'sample_submission.csv')\nsub['t'] = 0\nsubmission = pd.merge(sub[['Id','t']], final_pred, how='left', on='Id').fillna(0.0)\nsubmission[['Id','StartHesitation', 'Turn' , 'Walking']].to_csv('submission.csv', index=False)\n</code></pre></li>\n</ol>\n<p>Hope this helps someone going through the same issue.<br>\nHappy Kaggling!</p>",
  "messages": [
    {
      "id": 2204219,
      "postDate": "2023-03-31T12:22:50.233Z",
      "content": "<p>Hello fellow Kagglers!</p>\n<p>Sorry for another post about submission error and such. I have been stuck for a while trying to figure out why my notebook gives this error (Notebook Threw Exception). After going through a few posts on similar error, I made all the changes but I still couldn't point the error here. My output (head) looks like this:<br>\n.               Id            StartHesitation   Turn       Walking<br>\n0    003f117e14_0    0.0              0.002          0.0<br>\n1    003f117e14_1            0.0              0.002          0.0<br>\n2    003f117e14_2            0.0              0.002          0.0<br>\n3    003f117e14_3            0.0              0.002          0.0<br>\n4    003f117e14_4    0.0              0.002          0.0</p>\n<p>and tail like this</p>\n<p>.                          Id           StartHesitation   Turn        Walking<br>\n286365    02ab235146_281683   0.0             0.0      0.0<br>\n286366    02ab235146_281684   0.0             0.0      0.0<br>\n286367    02ab235146_281685   0.0               0.0        0.0<br>\n286368    02ab235146_281686   0.0             0.0      0.0<br>\n286369    02ab235146_281687   0.0             0.0      0.0</p>\n<p>Is it because the column types are float and not int because I thought that was no issue. Also the the decimal point limit is just 3 which worked in all cases in the posts I went through. None of the values exceed 0.9.</p>\n<p>Any sort of idea/tip would be greatly appreciated!</p>\n<p>TIA</p>\n<p>Update 1 : I made the test source more generic after <a href=\"https://www.kaggle.com/avivlevi815\" target=\"_blank\">@avivlevi815</a> pointed it out so that even if ID changes, they can be processed. Now I am getting 'Submission Scoring Error'. I have checked my number of rows and columns are correct.  And I rounded my probabilities to 3 after going through the post on 'Submission Scoring Error' although the metric has been updated to cater for this issue. The values are between 0 and 1. No row has more than one column with non zero value.</p>\n<p>Update 2 [Solved]:</p>\n<p>So few things that I did to rectify the error : </p>\n<ol>\n<li><p>The update 1 mentioned above to use a generic code that uses file with any name.csv inside the test folder that Kaggle tests from their end</p></li>\n<li><p>For some reason I had not made use of sample_submission.csv file to create the submission.csv. I just looked at the structure and tried to imitate it. I copied the small following code from <a href=\"https://www.kaggle.com/ammarnassanalhajali\" target=\"_blank\">@ammarnassanalhajali</a> to use submission_sample.csv while creating the submission file. As <a href=\"https://www.kaggle.com/avivlevi815\" target=\"_blank\">@avivlevi815</a> rightly mentioned that both sample_submission.csv and test files are change in Kaggles backend for scoring. Also, index = False.</p>\n<pre><code>sub = pd.read_csv(p+'sample_submission.csv')\nsub['t'] = 0\nsubmission = pd.merge(sub[['Id','t']], final_pred, how='left', on='Id').fillna(0.0)\nsubmission[['Id','StartHesitation', 'Turn' , 'Walking']].to_csv('submission.csv', index=False)\n</code></pre></li>\n</ol>\n<p>Hope this helps someone going through the same issue.<br>\nHappy Kaggling!</p>",
      "rawMarkdown": "Hello fellow Kagglers!\n\nSorry for another post about submission error and such. I have been stuck for a while trying to figure out why my notebook gives this error (Notebook Threw Exception). After going through a few posts on similar error, I made all the changes but I still couldn't point the error here. My output (head) looks like this:\n.               Id\t        StartHesitation   Turn\t     Walking\n0\t003f117e14_0 \t0.0\t             0.002\t        0.0\n1\t003f117e14_1\t        0.0\t             0.002\t        0.0\n2\t003f117e14_2\t        0.0\t             0.002\t        0.0\n3\t003f117e14_3\t        0.0\t             0.002\t        0.0\n4\t003f117e14_4   \t0.0\t             0.002\t        0.0\n\nand tail like this\n\n.\t                      Id\t       StartHesitation   Turn\t     Walking\n286365\t02ab235146_281683\t0.0\t            0.0\t     0.0\n286366\t02ab235146_281684\t0.0\t            0.0\t     0.0\n286367\t02ab235146_281685\t0.0               0.0\t     0.0\n286368\t02ab235146_281686\t0.0\t            0.0\t     0.0\n286369\t02ab235146_281687\t0.0\t            0.0\t     0.0\n\nIs it because the column types are float and not int because I thought that was no issue. Also the the decimal point limit is just 3 which worked in all cases in the posts I went through. None of the values exceed 0.9.\n\nAny sort of idea/tip would be greatly appreciated!\n\nTIA\n\n\nUpdate 1 : I made the test source more generic after @avivlevi815 pointed it out so that even if ID changes, they can be processed. Now I am getting 'Submission Scoring Error'. I have checked my number of rows and columns are correct.  And I rounded my probabilities to 3 after going through the post on 'Submission Scoring Error' although the metric has been updated to cater for this issue. The values are between 0 and 1. No row has more than one column with non zero value.\n\nUpdate 2 [Solved]:\n\nSo few things that I did to rectify the error : \n1. The update 1 mentioned above to use a generic code that uses file with any name.csv inside the test folder that Kaggle tests from their end\n2. For some reason I had not made use of sample_submission.csv file to create the submission.csv. I just looked at the structure and tried to imitate it. I copied the small following code from @ammarnassanalhajali to use submission_sample.csv while creating the submission file. As @avivlevi815 rightly mentioned that both sample_submission.csv and test files are change in Kaggles backend for scoring. Also, index = False.\n\n        sub = pd.read_csv(p+'sample_submission.csv')\n        sub['t'] = 0\n        submission = pd.merge(sub[['Id','t']], final_pred, how='left', on='Id').fillna(0.0)\n        submission[['Id','StartHesitation', 'Turn' , 'Walking']].to_csv('submission.csv', index=False)\n\nHope this helps someone going through the same issue.\nHappy Kaggling!",
      "votes": 4
    },
    {
      "id": 2204380,
      "postDate": "2023-03-31T14:15:07.833Z",
      "content": "<p>Did you take into consideration that the test  &amp; sample submission files are replaced when run at Kaggles end? (after you submit)</p>",
      "rawMarkdown": "Did you take into consideration that the test  & sample submission files are replaced when run at Kaggles end? (after you submit)",
      "votes": 1,
      "replies": [
        {
          "id": 2204480,
          "postDate": "2023-03-31T15:45:30.857Z",
          "content": "<p>Hi again. I did read this but as I think of it now, I did not pay much heed to it. I mean we only take into account the test.csv right? At Kaggle's end they just replace this test.csv with another test.csv with different values? At least that is what I could make of it. <br>\nOr does it also mean that there will no longer be files named '02ab235146.csv' and '003f117e14.csv'? And that I should code it in a generic way that caters for whatever filename.csv it has? Sorry if this is a basic question I am new to this comp. </p>",
          "rawMarkdown": "Hi again. I did read this but as I think of it now, I did not pay much heed to it. I mean we only take into account the test.csv right? At Kaggle's end they just replace this test.csv with another test.csv with different values? At least that is what I could make of it. \nOr does it also mean that there will no longer be files named '02ab235146.csv' and '003f117e14.csv'? And that I should code it in a generic way that caters for whatever filename.csv it has? Sorry if this is a basic question I am new to this comp. \n",
          "replies": [
            {
              "id": 2204743,
              "postDate": "2023-03-31T22:00:22.220Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2204806,
              "postDate": "2023-04-01T01:17:54.013Z",
              "content": "<p>Yes thanks! I did this update! </p>",
              "rawMarkdown": "Yes thanks! I did this update! "
            }
          ]
        }
      ]
    },
    {
      "id": 2204742,
      "postDate": "2023-03-31T21:56:37.667Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2204805,
          "postDate": "2023-04-01T01:17:23.590Z",
          "content": "<p>Hey! Thanks for the tip! After the generic code update for test.csv I am now getting a \"submission scoring error\" instead. Which means something in the output file is amiss! My output still looks the same as above. I have done all the basic checks. I'm really not sure what the issue is?</p>",
          "rawMarkdown": "Hey! Thanks for the tip! After the generic code update for test.csv I am now getting a \"submission scoring error\" instead. Which means something in the output file is amiss! My output still looks the same as above. I have done all the basic checks. I'm really not sure what the issue is?",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2204380,
      "author_name": "Aviv Levi",
      "author_url": "",
      "post_date": "2023-03-31T14:15:07.833000",
      "content": "<p>Did you take into consideration that the test  &amp; sample submission files are replaced when run at Kaggles end? (after you submit)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2204480,
          "author_name": "Sireesh Limbu",
          "author_url": "",
          "post_date": "2023-03-31T15:45:30.857000",
          "content": "<p>Hi again. I did read this but as I think of it now, I did not pay much heed to it. I mean we only take into account the test.csv right? At Kaggle's end they just replace this test.csv with another test.csv with different values? At least that is what I could make of it. <br>\nOr does it also mean that there will no longer be files named '02ab235146.csv' and '003f117e14.csv'? And that I should code it in a generic way that caters for whatever filename.csv it has? Sorry if this is a basic question I am new to this comp. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2204743,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-31T22:00:22.220000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2204806,
              "author_name": "Sireesh Limbu",
              "author_url": "",
              "post_date": "2023-04-01T01:17:54.013000",
              "content": "<p>Yes thanks! I did this update! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2204742,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-31T21:56:37.667000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2204805,
          "author_name": "Sireesh Limbu",
          "author_url": "",
          "post_date": "2023-04-01T01:17:23.590000",
          "content": "<p>Hey! Thanks for the tip! After the generic code update for test.csv I am now getting a \"submission scoring error\" instead. Which means something in the output file is amiss! My output still looks the same as above. I have done all the basic checks. I'm really not sure what the issue is?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2204219": "Hello fellow Kagglers!\n\nSorry for another post about submission error and such. I have been stuck for a while trying to figure out why my notebook gives this error (Notebook Threw Exception). After going through a few posts on similar error, I made all the changes but I still couldn't point the error here. My output (head) looks like this:\n.               Id\t        StartHesitation   Turn\t     Walking\n0\t003f117e14_0 \t0.0\t             0.002\t        0.0\n1\t003f117e14_1\t        0.0\t             0.002\t        0.0\n2\t003f117e14_2\t        0.0\t             0.002\t        0.0\n3\t003f117e14_3\t        0.0\t             0.002\t        0.0\n4\t003f117e14_4   \t0.0\t             0.002\t        0.0\n\nand tail like this\n\n.\t                      Id\t       StartHesitation   Turn\t     Walking\n286365\t02ab235146_281683\t0.0\t            0.0\t     0.0\n286366\t02ab235146_281684\t0.0\t            0.0\t     0.0\n286367\t02ab235146_281685\t0.0               0.0\t     0.0\n286368\t02ab235146_281686\t0.0\t            0.0\t     0.0\n286369\t02ab235146_281687\t0.0\t            0.0\t     0.0\n\nIs it because the column types are float and not int because I thought that was no issue. Also the the decimal point limit is just 3 which worked in all cases in the posts I went through. None of the values exceed 0.9.\n\nAny sort of idea/tip would be greatly appreciated!\n\nTIA\n\n\nUpdate 1 : I made the test source more generic after @avivlevi815 pointed it out so that even if ID changes, they can be processed. Now I am getting 'Submission Scoring Error'. I have checked my number of rows and columns are correct.  And I rounded my probabilities to 3 after going through the post on 'Submission Scoring Error' although the metric has been updated to cater for this issue. The values are between 0 and 1. No row has more than one column with non zero value.\n\nUpdate 2 [Solved]:\n\nSo few things that I did to rectify the error : \n1. The update 1 mentioned above to use a generic code that uses file with any name.csv inside the test folder that Kaggle tests from their end\n2. For some reason I had not made use of sample_submission.csv file to create the submission.csv. I just looked at the structure and tried to imitate it. I copied the small following code from @ammarnassanalhajali to use submission_sample.csv while creating the submission file. As @avivlevi815 rightly mentioned that both sample_submission.csv and test files are change in Kaggles backend for scoring. Also, index = False.\n\n        sub = pd.read_csv(p+'sample_submission.csv')\n        sub['t'] = 0\n        submission = pd.merge(sub[['Id','t']], final_pred, how='left', on='Id').fillna(0.0)\n        submission[['Id','StartHesitation', 'Turn' , 'Walking']].to_csv('submission.csv', index=False)\n\nHope this helps someone going through the same issue.\nHappy Kaggling!",
    "2204380": "Did you take into consideration that the test  & sample submission files are replaced when run at Kaggles end? (after you submit)",
    "2204742": ""
  }
}