{
  "id": 406700,
  "title": "Data Update and Rescore",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406700",
  "author_name": "Ryan Holbrook",
  "post_date": "2023-05-03T13:22:24.531000",
  "votes": 29,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>As observed in another thread, some of the event annotations for the DeFOG series were misaligned with the true timestamps. We've identified the source of the error and will be updating the dataset with corrected event annotations shortly.</p>\n<p>As this also affects the test set, we'll be updating the leaderboard with new scores. Submissions may be disabled for a bit during this process.</p>\n<p>Note that as part of the realignment process some consecutive events were merged into a single event. While the total event time is approximately the same, the number of events listed in the <code>events.csv</code> file has decreased.</p>\n<p>I'll update this post when the data update and rescore is complete.</p>\n<p><strong>UPDATE:</strong> The data update is complete. </p>\n<p><strong>UPDATE:</strong> The rescore is complete. Please let us know if you have any questions or concerns!</p>",
  "messages": [
    {
      "id": 2244147,
      "postDate": "2023-05-03T13:22:24.530Z",
      "content": "<p>Hi everyone, </p>\n<p>As observed in another thread, some of the event annotations for the DeFOG series were misaligned with the true timestamps. We've identified the source of the error and will be updating the dataset with corrected event annotations shortly.</p>\n<p>As this also affects the test set, we'll be updating the leaderboard with new scores. Submissions may be disabled for a bit during this process.</p>\n<p>Note that as part of the realignment process some consecutive events were merged into a single event. While the total event time is approximately the same, the number of events listed in the <code>events.csv</code> file has decreased.</p>\n<p>I'll update this post when the data update and rescore is complete.</p>\n<p><strong>UPDATE:</strong> The data update is complete. </p>\n<p><strong>UPDATE:</strong> The rescore is complete. Please let us know if you have any questions or concerns!</p>",
      "rawMarkdown": "Hi everyone, \n\nAs observed in another thread, some of the event annotations for the DeFOG series were misaligned with the true timestamps. We've identified the source of the error and will be updating the dataset with corrected event annotations shortly.\n\nAs this also affects the test set, we'll be updating the leaderboard with new scores. Submissions may be disabled for a bit during this process.\n\nNote that as part of the realignment process some consecutive events were merged into a single event. While the total event time is approximately the same, the number of events listed in the `events.csv` file has decreased.\n\nI'll update this post when the data update and rescore is complete.\n\n**UPDATE:** The data update is complete. ~~The rescore is pending.~~\n\n**UPDATE:** The rescore is complete. Please let us know if you have any questions or concerns!",
      "votes": 29
    },
    {
      "id": 2269383,
      "postDate": "2023-05-22T12:45:56.277Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hi! Sorry to bother you with questions, but here's another one:</p>\n<p>It looks like the 'Init' and 'Completion' columns from 'events.csv' do not accurately convert to timesteps. Just a little example:</p>\n<p>pd.read_csv('events.csv').iloc[1205] would give this:<br>\nId: 02ea782681<br>\nInit: 1377.175<br>\nCompletion: 1378.089<br>\nType: Turn<br>\nKinetic: 1.0</p>\n<p>This is the file from the 'defog' dataset, which has 100Hz frequency, which means that neither the 'Init' nor 'Completion' column could have a non-zero third digit after the floating point. Otherwise, it converts to a non-integer step number.</p>\n<p>Is there any way to fix this on your side and reload the correct file? I can write a few lines of code to fix this too, of course, but given the 9-hour time limit, it would be nice to save the computing time for training the model rather than cleaning the data.</p>\n<p>Thanks!</p>\n<p>Upd: there's also a negative 'Init' and 'Completion' value at index = 2300.</p>",
      "rawMarkdown": "@ryanholbrook Hi! Sorry to bother you with questions, but here's another one:\n\nIt looks like the 'Init' and 'Completion' columns from 'events.csv' do not accurately convert to timesteps. Just a little example:\n\npd.read_csv('events.csv').iloc[1205] would give this:\nId: 02ea782681\nInit: 1377.175\nCompletion: 1378.089\nType: Turn\nKinetic: 1.0\n\nThis is the file from the 'defog' dataset, which has 100Hz frequency, which means that neither the 'Init' nor 'Completion' column could have a non-zero third digit after the floating point. Otherwise, it converts to a non-integer step number.\n\nIs there any way to fix this on your side and reload the correct file? I can write a few lines of code to fix this too, of course, but given the 9-hour time limit, it would be nice to save the computing time for training the model rather than cleaning the data.\n\nThanks!\n\nUpd: there's also a negative 'Init' and 'Completion' value at index = 2300.",
      "votes": 1
    },
    {
      "id": 2263369,
      "postDate": "2023-05-17T14:35:15.350Z",
      "content": "<p>I have a question that is similar to <a href=\"https://www.kaggle.com/abandura\" target=\"_blank\">@abandura</a>.  Since this data issue affects training, I would like to ask whether the re-score process will re-run all submitted notebooks for the private score. Or had the private score also been already assessed at the time of the submission before the data was updated?</p>\n<p>In addition, I would like to ask whether the old data is still available for the training purpose. In fact, I do not get as good a score with the updated data as before with the old data. </p>",
      "rawMarkdown": "I have a question that is similar to @abandura.  Since this data issue affects training, I would like to ask whether the re-score process will re-run all submitted notebooks for the private score. Or had the private score also been already assessed at the time of the submission before the data was updated?\n\nIn addition, I would like to ask whether the old data is still available for the training purpose. In fact, I do not get as good a score with the updated data as before with the old data. ",
      "votes": 1
    },
    {
      "id": 2245051,
      "postDate": "2023-05-04T06:06:04.940Z",
      "content": "<p>At present, the entire dataset is still unable to be downloaded, displaying as 404. May I ask how to solve such a problem?</p>",
      "rawMarkdown": "At present, the entire dataset is still unable to be downloaded, displaying as 404. May I ask how to solve such a problem?",
      "votes": 1
    },
    {
      "id": 2244935,
      "postDate": "2023-05-04T03:06:44.660Z",
      "content": "<p>404 error while downloading the dataset using kaggle api.</p>",
      "rawMarkdown": "404 error while downloading the dataset using kaggle api.",
      "votes": 1
    },
    {
      "id": 2244758,
      "postDate": "2023-05-03T21:34:06.207Z",
      "content": "<p>Is notype dataset also updated?</p>",
      "rawMarkdown": "Is notype dataset also updated?",
      "votes": 1,
      "replies": [
        {
          "id": 2244771,
          "postDate": "2023-05-03T21:48:03.660Z",
          "content": "<p>Yes, it should be. Does there appear to be an error?</p>",
          "rawMarkdown": "Yes, it should be. Does there appear to be an error?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2244636,
      "postDate": "2023-05-03T19:12:32.687Z",
      "content": "<p>404 error while downloading the dataset.</p>",
      "rawMarkdown": "404 error while downloading the dataset.",
      "votes": 1
    },
    {
      "id": 2244593,
      "postDate": "2023-05-03T18:28:13.907Z",
      "content": "<p>Thanks for the updates! Since this data issue affects training, I wanted to verify that the rescore process will re-run all submitted notebooks (as opposed to just re-evaluating the previously generated output files with the new ground truth)?</p>",
      "rawMarkdown": "Thanks for the updates! Since this data issue affects training, I wanted to verify that the rescore process will re-run all submitted notebooks (as opposed to just re-evaluating the previously generated output files with the new ground truth)?",
      "votes": 1,
      "replies": [
        {
          "id": 2244639,
          "postDate": "2023-05-03T19:14:18.307Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/abandura\" target=\"_blank\">@abandura</a>, I'm only re-evaluating the already generated output files. Since the update didn't affect the test set inputs (only the target), submissions that only do test set inference will have the same outputs either way. If you also do training during submission, you may want to resubmit.</p>",
          "rawMarkdown": "Hi @abandura, I'm only re-evaluating the already generated output files. Since the update didn't affect the test set inputs (only the target), submissions that only do test set inference will have the same outputs either way. If you also do training during submission, you may want to resubmit.",
          "votes": 1,
          "replies": [
            {
              "id": 2244706,
              "postDate": "2023-05-03T20:34:33.213Z",
              "content": "<p>Sounds good, thanks for clarifying!</p>",
              "rawMarkdown": "Sounds good, thanks for clarifying!",
              "votes": 1
            },
            {
              "id": 2245728,
              "postDate": "2023-05-04T14:46:44.203Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2244412,
      "postDate": "2023-05-03T15:53:26.383Z",
      "content": "<p>Download seem to be broken now: <br>\n<code>kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction</code><br>\nFails with:<br>\n<code>404 - Not Found</code></p>\n<p>Trying to download directly from the webpage also encounters '404' error code.</p>",
      "rawMarkdown": "Download seem to be broken now: \n`kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction`\nFails with:\n`404 - Not Found`\n\nTrying to download directly from the webpage also encounters '404' error code.",
      "votes": 1,
      "replies": [
        {
          "id": 2244479,
          "postDate": "2023-05-03T16:44:54.413Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/eliransinvani\" target=\"_blank\">@eliransinvani</a>, thanks for the heads up. It may be a caching issue due to the large dataset size. I'll give it another hour or so and if it hasn't resolved, I'll look into it more.</p>",
          "rawMarkdown": "Hi @eliransinvani, thanks for the heads up. It may be a caching issue due to the large dataset size. I'll give it another hour or so and if it hasn't resolved, I'll look into it more.",
          "votes": 3,
          "replies": [
            {
              "id": 2244876,
              "postDate": "2023-05-04T00:51:47.643Z",
              "content": "<p>I'm getting the same 404 message with the kaggle api, and the download link in the browser also leads to a 404 page:</p>\n<p><a href=\"https://www.kaggle.com/competitions/41880/download-all\" target=\"_blank\">https://www.kaggle.com/competitions/41880/download-all</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2Fa4d0c0cdda544f39a90c29dfd6cec440%2FScreenshot%20from%202023-05-03%2020-50-11.png?generation=1683161478202181&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "I'm getting the same 404 message with the kaggle api, and the download link in the browser also leads to a 404 page:\n\nhttps://www.kaggle.com/competitions/41880/download-all\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2Fa4d0c0cdda544f39a90c29dfd6cec440%2FScreenshot%20from%202023-05-03%2020-50-11.png?generation=1683161478202181&alt=media)\n\n",
              "votes": 1
            },
            {
              "id": 2245586,
              "postDate": "2023-05-04T13:04:03.517Z",
              "content": "<p>Still getting 404-Not Found while downloading using API</p>",
              "rawMarkdown": "Still getting 404-Not Found while downloading using API"
            }
          ]
        }
      ]
    },
    {
      "id": 2248314,
      "postDate": "2023-05-06T17:58:59.027Z",
      "content": "<p>Hi! Thank you for the update. I should say that the data still looks a little bit strange. I understand that the nature of defog and tdcsfog datasets is different and they come from somewhat different distributions, but… Below is the distribution of 'annotated' (rows where both \"Valid\" and \"Task\" flags == True for defog dataset and all rows for tdcsfog dataset). between 4 classes (StartHesitation, Turn, Walking, NoHesitation). StartHesitation accounts only for 500 out of 4+ million 'annotated' rows for defog, while in tdcsfog it has 224 times higher frequency… <br>\ndefog: [1.22233549e-04 1.43460383e-01 2.40844096e-02 8.32332974e-01]<br>\ntscsfog: [0.04315506 0.23769786 0.02942767 0.68971941]<br>\ntotal: [0.02737241 0.20313548 0.02746799 0.74202413]</p>",
      "rawMarkdown": "Hi! Thank you for the update. I should say that the data still looks a little bit strange. I understand that the nature of defog and tdcsfog datasets is different and they come from somewhat different distributions, but... Below is the distribution of 'annotated' (rows where both \"Valid\" and \"Task\" flags == True for defog dataset and all rows for tdcsfog dataset). between 4 classes (StartHesitation, Turn, Walking, NoHesitation). StartHesitation accounts only for 500 out of 4+ million 'annotated' rows for defog, while in tdcsfog it has 224 times higher frequency... \ndefog: [1.22233549e-04 1.43460383e-01 2.40844096e-02 8.32332974e-01]\ntscsfog: [0.04315506 0.23769786 0.02942767 0.68971941]\ntotal: [0.02737241 0.20313548 0.02746799 0.74202413]",
      "votes": 2
    },
    {
      "id": 2285748,
      "postDate": "2023-06-03T00:59:37.543Z",
      "content": "<p>Hi, The publicly posted notebooks with top scores (0.31 is top 6%)  uses regression models that predicts degree of fog events for each reading. Are these notebooks eligible for the competition, considering the stated evaluation criteria , which requires identification and classification between the three events, and not detecting degree of each event?  (refer my other posts in this regard) and also considering that regression scores (without being converted into confidence scores) cannot be presented to average_precision_score in any meaningful way?</p>\n<p><a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638#2285421\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638#2285421</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/414165\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/414165</a></p>",
      "rawMarkdown": "Hi, The publicly posted notebooks with top scores (0.31 is top 6%)  uses regression models that predicts degree of fog events for each reading. Are these notebooks eligible for the competition, considering the stated evaluation criteria , which requires identification and classification between the three events, and not detecting degree of each event?  (refer my other posts in this regard) and also considering that regression scores (without being converted into confidence scores) cannot be presented to average_precision_score in any meaningful way?\n\nhttps://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638#2285421\n\nhttps://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/414165"
    },
    {
      "id": 2255881,
      "postDate": "2023-05-12T04:32:23.233Z",
      "content": "<p>contributor</p>",
      "rawMarkdown": "contributor",
      "replies": [
        {
          "id": 2268752,
          "postDate": "2023-05-22T01:28:02.580Z",
          "content": "<p>Me too Contributor!</p>",
          "rawMarkdown": "Me too Contributor!"
        }
      ]
    },
    {
      "id": 2249611,
      "postDate": "2023-05-08T00:28:39.303Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> - I'm not sure if it was already covered - you have the same file - '003f117e14.csv' present both train and test sets. The only difference - is target indicators, which, obviously enough, are present in the copy of that file train set and are omitted in the test set.<br>\nWas it some malfunction? Just wanted to ensure that no duplicates would be across the train/test sets by the deadline and suggest deleting the file from both sets as this kind of leakage could seriously affect preliminary ranks. Should I have missed or misunderstood something - would appreciate your feedback.</p>",
      "rawMarkdown": "@ryanholbrook - I'm not sure if it was already covered - you have the same file - '003f117e14.csv' present both train and test sets. The only difference - is target indicators, which, obviously enough, are present in the copy of that file train set and are omitted in the test set.\nWas it some malfunction? Just wanted to ensure that no duplicates would be across the train/test sets by the deadline and suggest deleting the file from both sets as this kind of leakage could seriously affect preliminary ranks. Should I have missed or misunderstood something - would appreciate your feedback.",
      "replies": [
        {
          "id": 2250239,
          "postDate": "2023-05-08T12:41:16.793Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , The publicly available test data is just an example drawn from the training set. It's not the real test data. The hidden test data (which your submission is run on) doesn't have any series from the training data.</p>",
          "rawMarkdown": "Hi @victorshlepov , The publicly available test data is just an example drawn from the training set. It's not the real test data. The hidden test data (which your submission is run on) doesn't have any series from the training data.",
          "votes": 2,
          "replies": [
            {
              "id": 2287063,
              "postDate": "2023-06-04T05:24:11.297Z",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Can we join the subjects data with the hidden test data? I think I can, but this is just a question to make sure.</p>",
              "rawMarkdown": "@ryanholbrook Can we join the subjects data with the hidden test data? I think I can, but this is just a question to make sure."
            }
          ]
        }
      ]
    },
    {
      "id": 2245831,
      "postDate": "2023-05-04T16:13:20.543Z",
      "content": "<p>The data download appears to be working now!</p>",
      "rawMarkdown": "The data download appears to be working now!",
      "replies": [
        {
          "id": 2245854,
          "postDate": "2023-05-04T16:28:10.737Z",
          "content": "<p>I'm still getting a 404 - Not Found error using the kaggle api download.</p>",
          "rawMarkdown": "I'm still getting a 404 - Not Found error using the kaggle api download.",
          "votes": 1,
          "replies": [
            {
              "id": 2245943,
              "postDate": "2023-05-04T17:22:47.290Z",
              "content": "<p>Hmm. That's odd. Download from the webpage is working. I'll look into it.</p>",
              "rawMarkdown": "Hmm. That's odd. Download from the webpage is working. I'll look into it.",
              "votes": 1
            },
            {
              "id": 2246019,
              "postDate": "2023-05-04T18:48:40.920Z",
              "content": "<p><a href=\"https://www.kaggle.com/dynamic24\" target=\"_blank\">@dynamic24</a> The kaggle api download is working for me now. Would you be able to try it again?</p>",
              "rawMarkdown": "@dynamic24 The kaggle api download is working for me now. Would you be able to try it again?",
              "votes": 1
            },
            {
              "id": 2246022,
              "postDate": "2023-05-04T18:54:17.967Z",
              "content": "<p>It seems to be working now.  Thank you!</p>",
              "rawMarkdown": "It seems to be working now.  Thank you!",
              "votes": 1
            },
            {
              "id": 2268755,
              "postDate": "2023-05-22T01:32:06.333Z",
              "content": "<p>Thank god it’s resolved.!! Just another contributor.</p>",
              "rawMarkdown": "Thank god it’s resolved.!! Just another contributor."
            }
          ]
        }
      ]
    },
    {
      "id": 2245731,
      "postDate": "2023-05-04T14:50:49.043Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Is there a way for us to just download DEFOG folder and the changed files? It seems re-downloading 70GB is not a good option? (as well, 404 error is still there)</p>",
      "rawMarkdown": "@ryanholbrook Is there a way for us to just download DEFOG folder and the changed files? It seems re-downloading 70GB is not a good option? (as well, 404 error is still there)",
      "replies": [
        {
          "id": 2245739,
          "postDate": "2023-05-04T14:57:18.960Z",
          "content": "<p>The files seem to be accessible through notebooks, so you could maybe copy the defog and notype folders and <code>events.csv</code> to the working directory of the notebook editor and then download the outputs.</p>\n<p>We're still investigating why the download archive is failing.</p>",
          "rawMarkdown": "The files seem to be accessible through notebooks, so you could maybe copy the defog and notype folders and `events.csv` to the working directory of the notebook editor and then download the outputs.\n\nWe're still investigating why the download archive is failing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2245611,
      "postDate": "2023-05-04T13:24:14.803Z",
      "content": "<p>It appears the archive creation failed, causing the downloads to 404. I'm going to try recreating.</p>",
      "rawMarkdown": "It appears the archive creation failed, causing the downloads to 404. I'm going to try recreating."
    }
  ],
  "comments": [
    {
      "id": 2269383,
      "author_name": "Victor Shlepov",
      "author_url": "",
      "post_date": "2023-05-22T12:45:56.277000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Hi! Sorry to bother you with questions, but here's another one:</p>\n<p>It looks like the 'Init' and 'Completion' columns from 'events.csv' do not accurately convert to timesteps. Just a little example:</p>\n<p>pd.read_csv('events.csv').iloc[1205] would give this:<br>\nId: 02ea782681<br>\nInit: 1377.175<br>\nCompletion: 1378.089<br>\nType: Turn<br>\nKinetic: 1.0</p>\n<p>This is the file from the 'defog' dataset, which has 100Hz frequency, which means that neither the 'Init' nor 'Completion' column could have a non-zero third digit after the floating point. Otherwise, it converts to a non-integer step number.</p>\n<p>Is there any way to fix this on your side and reload the correct file? I can write a few lines of code to fix this too, of course, but given the 9-hour time limit, it would be nice to save the computing time for training the model rather than cleaning the data.</p>\n<p>Thanks!</p>\n<p>Upd: there's also a negative 'Init' and 'Completion' value at index = 2300.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2263369,
      "author_name": "Goki Fujiya",
      "author_url": "",
      "post_date": "2023-05-17T14:35:15.350000",
      "content": "<p>I have a question that is similar to <a href=\"https://www.kaggle.com/abandura\" target=\"_blank\">@abandura</a>.  Since this data issue affects training, I would like to ask whether the re-score process will re-run all submitted notebooks for the private score. Or had the private score also been already assessed at the time of the submission before the data was updated?</p>\n<p>In addition, I would like to ask whether the old data is still available for the training purpose. In fact, I do not get as good a score with the updated data as before with the old data. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2245051,
      "author_name": "Lingduo Wang",
      "author_url": "",
      "post_date": "2023-05-04T06:06:04.940000",
      "content": "<p>At present, the entire dataset is still unable to be downloaded, displaying as 404. May I ask how to solve such a problem?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2244935,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2023-05-04T03:06:44.660000",
      "content": "<p>404 error while downloading the dataset using kaggle api.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2244758,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2023-05-03T21:34:06.207000",
      "content": "<p>Is notype dataset also updated?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244771,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-05-03T21:48:03.660000",
          "content": "<p>Yes, it should be. Does there appear to be an error?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2244636,
      "author_name": "Mayank Jain",
      "author_url": "",
      "post_date": "2023-05-03T19:12:32.687000",
      "content": "<p>404 error while downloading the dataset.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2244593,
      "author_name": "abandura",
      "author_url": "",
      "post_date": "2023-05-03T18:28:13.907000",
      "content": "<p>Thanks for the updates! Since this data issue affects training, I wanted to verify that the rescore process will re-run all submitted notebooks (as opposed to just re-evaluating the previously generated output files with the new ground truth)?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244639,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-05-03T19:14:18.307000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/abandura\" target=\"_blank\">@abandura</a>, I'm only re-evaluating the already generated output files. Since the update didn't affect the test set inputs (only the target), submissions that only do test set inference will have the same outputs either way. If you also do training during submission, you may want to resubmit.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2244706,
              "author_name": "abandura",
              "author_url": "",
              "post_date": "2023-05-03T20:34:33.213000",
              "content": "<p>Sounds good, thanks for clarifying!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2245728,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-05-04T14:46:44.203000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2244412,
      "author_name": "Eliran Sinvani",
      "author_url": "",
      "post_date": "2023-05-03T15:53:26.383000",
      "content": "<p>Download seem to be broken now: <br>\n<code>kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction</code><br>\nFails with:<br>\n<code>404 - Not Found</code></p>\n<p>Trying to download directly from the webpage also encounters '404' error code.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2244479,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-05-03T16:44:54.413000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/eliransinvani\" target=\"_blank\">@eliransinvani</a>, thanks for the heads up. It may be a caching issue due to the large dataset size. I'll give it another hour or so and if it hasn't resolved, I'll look into it more.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2244876,
              "author_name": "Rob Mulla",
              "author_url": "",
              "post_date": "2023-05-04T00:51:47.643000",
              "content": "<p>I'm getting the same 404 message with the kaggle api, and the download link in the browser also leads to a 404 page:</p>\n<p><a href=\"https://www.kaggle.com/competitions/41880/download-all\" target=\"_blank\">https://www.kaggle.com/competitions/41880/download-all</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F644036%2Fa4d0c0cdda544f39a90c29dfd6cec440%2FScreenshot%20from%202023-05-03%2020-50-11.png?generation=1683161478202181&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2245586,
              "author_name": "Sdilshad",
              "author_url": "",
              "post_date": "2023-05-04T13:04:03.517000",
              "content": "<p>Still getting 404-Not Found while downloading using API</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2248314,
      "author_name": "Victor Shlepov",
      "author_url": "",
      "post_date": "2023-05-06T17:58:59.027000",
      "content": "<p>Hi! Thank you for the update. I should say that the data still looks a little bit strange. I understand that the nature of defog and tdcsfog datasets is different and they come from somewhat different distributions, but… Below is the distribution of 'annotated' (rows where both \"Valid\" and \"Task\" flags == True for defog dataset and all rows for tdcsfog dataset). between 4 classes (StartHesitation, Turn, Walking, NoHesitation). StartHesitation accounts only for 500 out of 4+ million 'annotated' rows for defog, while in tdcsfog it has 224 times higher frequency… <br>\ndefog: [1.22233549e-04 1.43460383e-01 2.40844096e-02 8.32332974e-01]<br>\ntscsfog: [0.04315506 0.23769786 0.02942767 0.68971941]<br>\ntotal: [0.02737241 0.20313548 0.02746799 0.74202413]</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2285748,
      "author_name": "Murugesan Narayanaswamy",
      "author_url": "",
      "post_date": "2023-06-03T00:59:37.543000",
      "content": "<p>Hi, The publicly posted notebooks with top scores (0.31 is top 6%)  uses regression models that predicts degree of fog events for each reading. Are these notebooks eligible for the competition, considering the stated evaluation criteria , which requires identification and classification between the three events, and not detecting degree of each event?  (refer my other posts in this regard) and also considering that regression scores (without being converted into confidence scores) cannot be presented to average_precision_score in any meaningful way?</p>\n<p><a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638#2285421\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638#2285421</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/414165\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/414165</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2255881,
      "author_name": "efsssfe112",
      "author_url": "",
      "post_date": "2023-05-12T04:32:23.233000",
      "content": "<p>contributor</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2268752,
          "author_name": "Mani Sai Kamal Darla",
          "author_url": "",
          "post_date": "2023-05-22T01:28:02.580000",
          "content": "<p>Me too Contributor!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2249611,
      "author_name": "Victor Shlepov",
      "author_url": "",
      "post_date": "2023-05-08T00:28:39.303000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> - I'm not sure if it was already covered - you have the same file - '003f117e14.csv' present both train and test sets. The only difference - is target indicators, which, obviously enough, are present in the copy of that file train set and are omitted in the test set.<br>\nWas it some malfunction? Just wanted to ensure that no duplicates would be across the train/test sets by the deadline and suggest deleting the file from both sets as this kind of leakage could seriously affect preliminary ranks. Should I have missed or misunderstood something - would appreciate your feedback.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2250239,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-05-08T12:41:16.793000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> , The publicly available test data is just an example drawn from the training set. It's not the real test data. The hidden test data (which your submission is run on) doesn't have any series from the training data.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2287063,
              "author_name": "YT",
              "author_url": "",
              "post_date": "2023-06-04T05:24:11.297000",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Can we join the subjects data with the hidden test data? I think I can, but this is just a question to make sure.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2245831,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2023-05-04T16:13:20.543000",
      "content": "<p>The data download appears to be working now!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2245854,
          "author_name": "dynamic24",
          "author_url": "",
          "post_date": "2023-05-04T16:28:10.737000",
          "content": "<p>I'm still getting a 404 - Not Found error using the kaggle api download.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2245943,
              "author_name": "Ryan Holbrook",
              "author_url": "",
              "post_date": "2023-05-04T17:22:47.290000",
              "content": "<p>Hmm. That's odd. Download from the webpage is working. I'll look into it.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2246019,
              "author_name": "Ryan Holbrook",
              "author_url": "",
              "post_date": "2023-05-04T18:48:40.920000",
              "content": "<p><a href=\"https://www.kaggle.com/dynamic24\" target=\"_blank\">@dynamic24</a> The kaggle api download is working for me now. Would you be able to try it again?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2246022,
              "author_name": "dynamic24",
              "author_url": "",
              "post_date": "2023-05-04T18:54:17.967000",
              "content": "<p>It seems to be working now.  Thank you!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2268755,
              "author_name": "Krishna Sumanth Yenugula",
              "author_url": "",
              "post_date": "2023-05-22T01:32:06.333000",
              "content": "<p>Thank god it’s resolved.!! Just another contributor.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2245731,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2023-05-04T14:50:49.043000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Is there a way for us to just download DEFOG folder and the changed files? It seems re-downloading 70GB is not a good option? (as well, 404 error is still there)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2245739,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-05-04T14:57:18.960000",
          "content": "<p>The files seem to be accessible through notebooks, so you could maybe copy the defog and notype folders and <code>events.csv</code> to the working directory of the notebook editor and then download the outputs.</p>\n<p>We're still investigating why the download archive is failing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2245611,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2023-05-04T13:24:14.803000",
      "content": "<p>It appears the archive creation failed, causing the downloads to 404. I'm going to try recreating.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2244147": "Hi everyone, \n\nAs observed in another thread, some of the event annotations for the DeFOG series were misaligned with the true timestamps. We've identified the source of the error and will be updating the dataset with corrected event annotations shortly.\n\nAs this also affects the test set, we'll be updating the leaderboard with new scores. Submissions may be disabled for a bit during this process.\n\nNote that as part of the realignment process some consecutive events were merged into a single event. While the total event time is approximately the same, the number of events listed in the `events.csv` file has decreased.\n\nI'll update this post when the data update and rescore is complete.\n\n**UPDATE:** The data update is complete. ~~The rescore is pending.~~\n\n**UPDATE:** The rescore is complete. Please let us know if you have any questions or concerns!",
    "2269383": "@ryanholbrook Hi! Sorry to bother you with questions, but here's another one:\n\nIt looks like the 'Init' and 'Completion' columns from 'events.csv' do not accurately convert to timesteps. Just a little example:\n\npd.read_csv('events.csv').iloc[1205] would give this:\nId: 02ea782681\nInit: 1377.175\nCompletion: 1378.089\nType: Turn\nKinetic: 1.0\n\nThis is the file from the 'defog' dataset, which has 100Hz frequency, which means that neither the 'Init' nor 'Completion' column could have a non-zero third digit after the floating point. Otherwise, it converts to a non-integer step number.\n\nIs there any way to fix this on your side and reload the correct file? I can write a few lines of code to fix this too, of course, but given the 9-hour time limit, it would be nice to save the computing time for training the model rather than cleaning the data.\n\nThanks!\n\nUpd: there's also a negative 'Init' and 'Completion' value at index = 2300.",
    "2263369": "I have a question that is similar to @abandura.  Since this data issue affects training, I would like to ask whether the re-score process will re-run all submitted notebooks for the private score. Or had the private score also been already assessed at the time of the submission before the data was updated?\n\nIn addition, I would like to ask whether the old data is still available for the training purpose. In fact, I do not get as good a score with the updated data as before with the old data. ",
    "2245051": "At present, the entire dataset is still unable to be downloaded, displaying as 404. May I ask how to solve such a problem?",
    "2244935": "404 error while downloading the dataset using kaggle api.",
    "2244758": "Is notype dataset also updated?",
    "2244636": "404 error while downloading the dataset.",
    "2244593": "Thanks for the updates! Since this data issue affects training, I wanted to verify that the rescore process will re-run all submitted notebooks (as opposed to just re-evaluating the previously generated output files with the new ground truth)?",
    "2244412": "Download seem to be broken now: \n`kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction`\nFails with:\n`404 - Not Found`\n\nTrying to download directly from the webpage also encounters '404' error code.",
    "2248314": "Hi! Thank you for the update. I should say that the data still looks a little bit strange. I understand that the nature of defog and tdcsfog datasets is different and they come from somewhat different distributions, but... Below is the distribution of 'annotated' (rows where both \"Valid\" and \"Task\" flags == True for defog dataset and all rows for tdcsfog dataset). between 4 classes (StartHesitation, Turn, Walking, NoHesitation). StartHesitation accounts only for 500 out of 4+ million 'annotated' rows for defog, while in tdcsfog it has 224 times higher frequency... \ndefog: [1.22233549e-04 1.43460383e-01 2.40844096e-02 8.32332974e-01]\ntscsfog: [0.04315506 0.23769786 0.02942767 0.68971941]\ntotal: [0.02737241 0.20313548 0.02746799 0.74202413]",
    "2285748": "Hi, The publicly posted notebooks with top scores (0.31 is top 6%)  uses regression models that predicts degree of fog events for each reading. Are these notebooks eligible for the competition, considering the stated evaluation criteria , which requires identification and classification between the three events, and not detecting degree of each event?  (refer my other posts in this regard) and also considering that regression scores (without being converted into confidence scores) cannot be presented to average_precision_score in any meaningful way?\n\nhttps://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/413638#2285421\n\nhttps://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/414165",
    "2255881": "contributor",
    "2249611": "@ryanholbrook - I'm not sure if it was already covered - you have the same file - '003f117e14.csv' present both train and test sets. The only difference - is target indicators, which, obviously enough, are present in the copy of that file train set and are omitted in the test set.\nWas it some malfunction? Just wanted to ensure that no duplicates would be across the train/test sets by the deadline and suggest deleting the file from both sets as this kind of leakage could seriously affect preliminary ranks. Should I have missed or misunderstood something - would appreciate your feedback.",
    "2245831": "The data download appears to be working now!",
    "2245731": "@ryanholbrook Is there a way for us to just download DEFOG folder and the changed files? It seems re-downloading 70GB is not a good option? (as well, 404 error is still there)",
    "2245611": "It appears the archive creation failed, causing the downloads to 404. I'm going to try recreating."
  }
}