{
  "id": 397982,
  "title": "Incorrect notype event annotations?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/397982",
  "author_name": "",
  "post_date": "2023-03-28T04:12:49.955533Z",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The entries in the \"events.csv\" file appear to be inconsistent with the \"Event\" column values in the notype training data files. It appears there may have been a unit conversion bug in the utility that populated the \"Event\" column of the train/notype/*.csv files.</p>\n<p>For example, based on the row \"1e8d55d48d,1021.904,1022.598,,\" in events.csv, I would expect train/notype/1e8d55d48d.csv to contain an event in the time range 1021.904 - 1022.598 seconds. The data tab seems to indicate that the notype files use a 100 Hz sample rate, so I would expect that event to impact timesteps 102190 - 102260. However, when I go to those rows I see a ton of rows with no event…</p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>102190</td>\n<td>-0.993268230764918</td>\n<td>-0.116984274771297</td>\n<td>-0.0413640801272263</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102191</td>\n<td>-0.993131260519859</td>\n<td>-0.117569239874438</td>\n<td>-0.0385470373678197</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102192</td>\n<td>-0.996553741019819</td>\n<td>-0.119481812857987</td>\n<td>-0.0380859375</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td></td>\n</tr>\n<tr>\n<td>102217</td>\n<td>-0.990922505974983</td>\n<td>-0.155768805534884</td>\n<td>-0.138887650788746</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102218</td>\n<td>-0.9871714142369</td>\n<td>-0.154445185733108</td>\n<td>-0.146245019082986</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102219</td>\n<td>-0.987959928058283</td>\n<td>-0.14827756563152</td>\n<td>-0.154395251826217</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td></td>\n</tr>\n<tr>\n<td>102258</td>\n<td>-0.869133339258985</td>\n<td>-0.136572456059369</td>\n<td>-0.34188932850263</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102259</td>\n<td>-0.861150736311346</td>\n<td>-0.142233667312971</td>\n<td>-0.353702279450288</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102260</td>\n<td>-0.85640032408877</td>\n<td>-0.153244516476253</td>\n<td>-0.357986408573962</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n</tbody>\n</table>\n<p></p>However, if I scroll way up to time sample 1022, I see an extremely short event…<p></p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1021</td>\n<td>-0.878437304100535</td>\n<td>-0.0897442834092885</td>\n<td>0.484561134573394</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>1022</td>\n<td>-0.878560270712976</td>\n<td>-0.0903227291237997</td>\n<td>0.484079230285104</td>\n<td>1</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>1023</td>\n<td>-0.876633455000578</td>\n<td>-0.0898869469325225</td>\n<td>0.486403429901475</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n</tbody>\n</table>\n<p></p>If I'm understanding the dataset correctly, time sample 1022 occurred (1022 samples) / (100 samples per second) = 10.22 seconds into the time series, not during the time range 1021.904 - 1022.598 seconds indicated by the events.csv file. Additionally, the \"events.csv\" file does not indicate that an event occurred around the time 10.22 seconds in series 1e8d55d48d. So, it appears it appears to me that the \"Event\" column was populated as if the sample rate was 1 Hz, rather than 100 Hz.<p></p>\n<p>The example above is not an outlier. I've observed many similar incidents where the events.csv file and train/notype/*.csv files appear to disagree with one another. In general, this causes the train/notype/*.csv files to incidate that there were a lot of unrealistically short events early on in each time series, with no events in the last 99% of each file.</p>\n<p>I am aware of the fact there was a problem early on in the competition in which the \"Event\" column was completely unpopulated (set to 0), so I redownloaded the dataset 2 days ago. The original issue appears to have been fixed, but the problem described above is still present in what I believe to be the most recent version.</p>\n<p>Am I misunderstanding something or is there something wrong with the latest version of the notype data?</p>\n<h1>Edit: The fix introduced a new problem</h1>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Uploaded an updated dataset in which the issue described above has been fixed. However, it appears there is a new problem in which the \"Time\" column contains some incorrect values.</p>\n<p>For example, in the previous version of the dataset train/notype/affdf8553f.csv contained the following rows</p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>208770</td>\n<td>-1.01594105348519</td>\n<td>-0.182570542573927</td>\n<td>0.168195947263265</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208771</td>\n<td>-1.00286768203618</td>\n<td>-0.138492169717451</td>\n<td>0.160291244375521</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208772</td>\n<td>-0.995033477474611</td>\n<td>-0.0994876348160082</td>\n<td>0.135104737617867</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208773</td>\n<td>-0.993005946294685</td>\n<td>-0.0848971999139799</td>\n<td>0.106476057956496</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208774</td>\n<td>-0.998477668146109</td>\n<td>-0.0886234945345456</td>\n<td>0.0923842235424117</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208775</td>\n<td>-1.00423168766911</td>\n<td>-0.0936964872820936</td>\n<td>0.0891615731894944</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n</tbody>\n</table>\n<p></p> In the latest version, those rows of train/notype/affdf8553f.csv instead contain<p></p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>208769</td>\n<td>-1.01594105348519</td>\n<td>-0.182570542573927</td>\n<td>0.168195947263265</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208771</td>\n<td>-1.00286768203618</td>\n<td>-0.138492169717451</td>\n<td>0.160291244375521</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208771</td>\n<td>-0.995033477474611</td>\n<td>-0.0994876348160082</td>\n<td>0.135104737617867</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208773</td>\n<td>-0.993005946294685</td>\n<td>-0.0848971999139799</td>\n<td>0.106476057956496</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208773</td>\n<td>-0.998477668146109</td>\n<td>-0.0886234945345456</td>\n<td>0.0923842235424117</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208775</td>\n<td>-1.00423168766911</td>\n<td>-0.0936964872820936</td>\n<td>0.0891615731894944</td>\n<td>1</td>\n<td>true</td>\n<td>true</td>\n</tr>\n</tbody>\n</table>\n<p></p> Notice that there are some rows with duplicate \"Time\" values that don't align with the original. Once again, this isn't a completely isolated anomaly, but its a bit rarer than the issue I originally created this post for. I haven't validated the dataset thoroughly enough to be sure of this, but it seems to mostly impact events that occur relatively late in the files, so I have a suspicion it might have been caused by some sort of rounding bug that causes intermittent off by one errors for large time values. <p></p>\n<h1>Edit 2: It appears everything has been fixed</h1>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Uploaded another iteration of the dataset. It appears all of the issues described above have been fixed. The latest dataset passes my sanity checks :-)</p>",
  "messages": [
    {
      "id": "2199812",
      "postDate": "03/28/2023 04:12:49",
      "content": "<p>The entries in the \"events.csv\" file appear to be inconsistent with the \"Event\" column values in the notype training data files. It appears there may have been a unit conversion bug in the utility that populated the \"Event\" column of the train/notype/*.csv files.</p>\n<p>For example, based on the row \"1e8d55d48d,1021.904,1022.598,,\" in events.csv, I would expect train/notype/1e8d55d48d.csv to contain an event in the time range 1021.904 - 1022.598 seconds. The data tab seems to indicate that the notype files use a 100 Hz sample rate, so I would expect that event to impact timesteps 102190 - 102260. However, when I go to those rows I see a ton of rows with no event…</p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>102190</td>\n<td>-0.993268230764918</td>\n<td>-0.116984274771297</td>\n<td>-0.0413640801272263</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102191</td>\n<td>-0.993131260519859</td>\n<td>-0.117569239874438</td>\n<td>-0.0385470373678197</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102192</td>\n<td>-0.996553741019819</td>\n<td>-0.119481812857987</td>\n<td>-0.0380859375</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td></td>\n</tr>\n<tr>\n<td>102217</td>\n<td>-0.990922505974983</td>\n<td>-0.155768805534884</td>\n<td>-0.138887650788746</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102218</td>\n<td>-0.9871714142369</td>\n<td>-0.154445185733108</td>\n<td>-0.146245019082986</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102219</td>\n<td>-0.987959928058283</td>\n<td>-0.14827756563152</td>\n<td>-0.154395251826217</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td>…</td>\n<td></td>\n</tr>\n<tr>\n<td>102258</td>\n<td>-0.869133339258985</td>\n<td>-0.136572456059369</td>\n<td>-0.34188932850263</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102259</td>\n<td>-0.861150736311346</td>\n<td>-0.142233667312971</td>\n<td>-0.353702279450288</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n<tr>\n<td>102260</td>\n<td>-0.85640032408877</td>\n<td>-0.153244516476253</td>\n<td>-0.357986408573962</td>\n<td>0</td>\n<td>false</td>\n<td>false</td>\n</tr>\n</tbody>\n</table>\n<p></p>However, if I scroll way up to time sample 1022, I see an extremely short event…<p></p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1021</td>\n<td>-0.878437304100535</td>\n<td>-0.0897442834092885</td>\n<td>0.484561134573394</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>1022</td>\n<td>-0.878560270712976</td>\n<td>-0.0903227291237997</td>\n<td>0.484079230285104</td>\n<td>1</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>1023</td>\n<td>-0.876633455000578</td>\n<td>-0.0898869469325225</td>\n<td>0.486403429901475</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n</tbody>\n</table>\n<p></p>If I'm understanding the dataset correctly, time sample 1022 occurred (1022 samples) / (100 samples per second) = 10.22 seconds into the time series, not during the time range 1021.904 - 1022.598 seconds indicated by the events.csv file. Additionally, the \"events.csv\" file does not indicate that an event occurred around the time 10.22 seconds in series 1e8d55d48d. So, it appears it appears to me that the \"Event\" column was populated as if the sample rate was 1 Hz, rather than 100 Hz.<p></p>\n<p>The example above is not an outlier. I've observed many similar incidents where the events.csv file and train/notype/*.csv files appear to disagree with one another. In general, this causes the train/notype/*.csv files to incidate that there were a lot of unrealistically short events early on in each time series, with no events in the last 99% of each file.</p>\n<p>I am aware of the fact there was a problem early on in the competition in which the \"Event\" column was completely unpopulated (set to 0), so I redownloaded the dataset 2 days ago. The original issue appears to have been fixed, but the problem described above is still present in what I believe to be the most recent version.</p>\n<p>Am I misunderstanding something or is there something wrong with the latest version of the notype data?</p>\n<h1>Edit: The fix introduced a new problem</h1>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Uploaded an updated dataset in which the issue described above has been fixed. However, it appears there is a new problem in which the \"Time\" column contains some incorrect values.</p>\n<p>For example, in the previous version of the dataset train/notype/affdf8553f.csv contained the following rows</p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>208770</td>\n<td>-1.01594105348519</td>\n<td>-0.182570542573927</td>\n<td>0.168195947263265</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208771</td>\n<td>-1.00286768203618</td>\n<td>-0.138492169717451</td>\n<td>0.160291244375521</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208772</td>\n<td>-0.995033477474611</td>\n<td>-0.0994876348160082</td>\n<td>0.135104737617867</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208773</td>\n<td>-0.993005946294685</td>\n<td>-0.0848971999139799</td>\n<td>0.106476057956496</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208774</td>\n<td>-0.998477668146109</td>\n<td>-0.0886234945345456</td>\n<td>0.0923842235424117</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208775</td>\n<td>-1.00423168766911</td>\n<td>-0.0936964872820936</td>\n<td>0.0891615731894944</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n</tbody>\n</table>\n<p></p> In the latest version, those rows of train/notype/affdf8553f.csv instead contain<p></p>\n<table>\n<thead>\n<tr>\n<th>Time</th>\n<th>AccV</th>\n<th>AccML</th>\n<th>AccAP</th>\n<th>Event</th>\n<th>Valid</th>\n<th>Task</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>208769</td>\n<td>-1.01594105348519</td>\n<td>-0.182570542573927</td>\n<td>0.168195947263265</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208771</td>\n<td>-1.00286768203618</td>\n<td>-0.138492169717451</td>\n<td>0.160291244375521</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208771</td>\n<td>-0.995033477474611</td>\n<td>-0.0994876348160082</td>\n<td>0.135104737617867</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208773</td>\n<td>-0.993005946294685</td>\n<td>-0.0848971999139799</td>\n<td>0.106476057956496</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208773</td>\n<td>-0.998477668146109</td>\n<td>-0.0886234945345456</td>\n<td>0.0923842235424117</td>\n<td>0</td>\n<td>true</td>\n<td>true</td>\n</tr>\n<tr>\n<td>208775</td>\n<td>-1.00423168766911</td>\n<td>-0.0936964872820936</td>\n<td>0.0891615731894944</td>\n<td>1</td>\n<td>true</td>\n<td>true</td>\n</tr>\n</tbody>\n</table>\n<p></p> Notice that there are some rows with duplicate \"Time\" values that don't align with the original. Once again, this isn't a completely isolated anomaly, but its a bit rarer than the issue I originally created this post for. I haven't validated the dataset thoroughly enough to be sure of this, but it seems to mostly impact events that occur relatively late in the files, so I have a suspicion it might have been caused by some sort of rounding bug that causes intermittent off by one errors for large time values. <p></p>\n<h1>Edit 2: It appears everything has been fixed</h1>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Uploaded another iteration of the dataset. It appears all of the issues described above have been fixed. The latest dataset passes my sanity checks :-)</p>",
      "rawMarkdown": "The entries in the \"events.csv\" file appear to be inconsistent with the \"Event\" column values in the notype training data files. It appears there may have been a unit conversion bug in the utility that populated the \"Event\" column of the train/notype/\\*.csv files.\n\nFor example, based on the row \"1e8d55d48d,1021.904,1022.598,,\" in events.csv, I would expect train/notype/1e8d55d48d.csv to contain an event in the time range 1021.904 - 1022.598 seconds. The data tab seems to indicate that the notype files use a 100 Hz sample rate, so I would expect that event to impact timesteps 102190 - 102260. However, when I go to those rows I see a ton of rows with no event...\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 102190 | -0.993268230764918 | -0.116984274771297 | -0.0413640801272263 | 0 | false | false |\n| 102191 | -0.993131260519859 | -0.117569239874438 | -0.0385470373678197 | 0 | false | false |\n| 102192 | -0.996553741019819 | -0.119481812857987 | -0.0380859375 | 0 | false | false |\n| ... | ... | ... | ... | ... | ... |\n| 102217 | -0.990922505974983 | -0.155768805534884 | -0.138887650788746 | 0 | false | false |\n| 102218 | -0.9871714142369 | -0.154445185733108 | -0.146245019082986 | 0 | false | false |\n| 102219 | -0.987959928058283 | -0.14827756563152 | -0.154395251826217 | 0 | false | false |\n| ... | ... | ... | ... | ... | ... |\n| 102258 | -0.869133339258985 | -0.136572456059369 | -0.34188932850263 | 0 | false | false |\n| 102259 | -0.861150736311346 | -0.142233667312971 | -0.353702279450288 | 0 | false | false |\n| 102260 | -0.85640032408877 | -0.153244516476253 | -0.357986408573962 | 0 | false | false |\n\n</p>However, if I scroll way up to time sample 1022, I see an extremely short event...\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 1021 | -0.878437304100535 | -0.0897442834092885 | 0.484561134573394 | 0 | true | true |\n| 1022 | -0.878560270712976 | -0.0903227291237997 | 0.484079230285104 | 1 | true | true |\n| 1023 | -0.876633455000578 | -0.0898869469325225 | 0.486403429901475 | 0 | true | true |\n\n</p>If I'm understanding the dataset correctly, time sample 1022 occurred (1022 samples) / (100 samples per second) = 10.22 seconds into the time series, not during the time range 1021.904 - 1022.598 seconds indicated by the events.csv file. Additionally, the \"events.csv\" file does not indicate that an event occurred around the time 10.22 seconds in series 1e8d55d48d. So, it appears it appears to me that the \"Event\" column was populated as if the sample rate was 1 Hz, rather than 100 Hz.\n\nThe example above is not an outlier. I've observed many similar incidents where the events.csv file and train/notype/\\*.csv files appear to disagree with one another. In general, this causes the train/notype/*.csv files to incidate that there were a lot of unrealistically short events early on in each time series, with no events in the last 99% of each file.\n\nI am aware of the fact there was a problem early on in the competition in which the \"Event\" column was completely unpopulated (set to 0), so I redownloaded the dataset 2 days ago. The original issue appears to have been fixed, but the problem described above is still present in what I believe to be the most recent version.\n\nAm I misunderstanding something or is there something wrong with the latest version of the notype data?\n\n# Edit: The fix introduced a new problem\n@ryanholbrook Uploaded an updated dataset in which the issue described above has been fixed. However, it appears there is a new problem in which the \"Time\" column contains some incorrect values.\n\nFor example, in the previous version of the dataset train/notype/affdf8553f.csv contained the following rows\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 208770 | -1.01594105348519 | -0.182570542573927 | 0.168195947263265 | 0 | true | true |\n| 208771 | -1.00286768203618 | -0.138492169717451 | 0.160291244375521 | 0 | true | true |\n| 208772 | -0.995033477474611 | -0.0994876348160082 | 0.135104737617867 | 0 | true | true |\n| 208773 | -0.993005946294685 | -0.0848971999139799 | 0.106476057956496 | 0 | true | true |\n| 208774 | -0.998477668146109 | -0.0886234945345456 | 0.0923842235424117 | 0 | true | true |\n| 208775 | -1.00423168766911 | -0.0936964872820936 | 0.0891615731894944 | 0 | true | true |\n\n</p> In the latest version, those rows of train/notype/affdf8553f.csv instead contain\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 208769 | -1.01594105348519 | -0.182570542573927 | 0.168195947263265 | 0 | true | true |\n| 208771 | -1.00286768203618 | -0.138492169717451 | 0.160291244375521 | 0 | true | true |\n| 208771 | -0.995033477474611 | -0.0994876348160082 | 0.135104737617867 | 0 | true | true |\n| 208773 | -0.993005946294685 | -0.0848971999139799 | 0.106476057956496 | 0 | true | true |\n| 208773 | -0.998477668146109 | -0.0886234945345456 | 0.0923842235424117 | 0 | true | true |\n| 208775 | -1.00423168766911 | -0.0936964872820936 | 0.0891615731894944 | 1 | true | true |\n\n</p> Notice that there are some rows with duplicate \"Time\" values that don't align with the original. Once again, this isn't a completely isolated anomaly, but its a bit rarer than the issue I originally created this post for. I haven't validated the dataset thoroughly enough to be sure of this, but it seems to mostly impact events that occur relatively late in the files, so I have a suspicion it might have been caused by some sort of rounding bug that causes intermittent off by one errors for large time values. \n\n# Edit 2: It appears everything has been fixed\n@ryanholbrook Uploaded another iteration of the dataset. It appears all of the issues described above have been fixed. The latest dataset passes my sanity checks :-)",
      "votes": null
    },
    {
      "id": "2200004",
      "postDate": "03/28/2023 07:46:39",
      "content": "<p>Hey, James I am also confused about the dataset because I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly. On which values can a human tell a particular subject has FoG? If you have an idea we can have an online meeting. Thanks</p>",
      "rawMarkdown": "Hey, James I am also confused about the dataset because I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly. On which values can a human tell a particular subject has FoG? If you have an idea we can have an online meeting. Thanks",
      "votes": null
    },
    {
      "id": "2200263",
      "postDate": "03/28/2023 12:24:12",
      "content": "<p>\"I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly\" - They are vertical, mediolateral, and anteroposterior accelerations. A diagram illustrating what that means was posted at <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/393852\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/393852</a>.</p>\n<p>\"On which values can a human tell a particular subject has FoG?\"</p>\n<ul>\n<li>The AccV , AccML and AccAP values can be used as features that are fed to a classification model or used to calculate more sophisticated features (like the frequency domain ones described in <a href=\"https://www.mdpi.com/1424-8220/19/23/5141\" target=\"_blank\">https://www.mdpi.com/1424-8220/19/23/5141</a> ).</li>\n<li>The StartHesitation, Turn, and Walking values in the defog and tdcsfog training datasets are useful labels for training classification models.</li>\n<li>The Valid and Task values in the defog training dataset are useful for filtering the training data to exclude timesteps that weren't annotated or may have been annotated inaccurately.</li>\n</ul>\n<p>This is all very tangential to the problem described in my original post. I think the documentation at <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data</a> is pretty clear but suspect the event labels in the notype training dataset are inaccurate.</p>",
      "rawMarkdown": "\"I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly\" - They are vertical, mediolateral, and anteroposterior accelerations. A diagram illustrating what that means was posted at https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/393852.\n\n\"On which values can a human tell a particular subject has FoG?\"\n- The AccV , AccML and AccAP values can be used as features that are fed to a classification model or used to calculate more sophisticated features (like the frequency domain ones described in https://www.mdpi.com/1424-8220/19/23/5141 ).\n- The StartHesitation, Turn, and Walking values in the defog and tdcsfog training datasets are useful labels for training classification models.\n- The Valid and Task values in the defog training dataset are useful for filtering the training data to exclude timesteps that weren't annotated or may have been annotated inaccurately.\n\nThis is all very tangential to the problem described in my original post. I think the documentation at https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data is pretty clear but suspect the event labels in the notype training dataset are inaccurate.",
      "votes": null
    },
    {
      "id": "2208198",
      "postDate": "04/03/2023 21:45:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> , Thank you for your feedback. I updated the dataset with what should be the correct annotations for the <code>notype</code> series. Please let me know if anything else comes up.</p>",
      "rawMarkdown": "Hi @jsday96 , Thank you for your feedback. I updated the dataset with what should be the correct annotations for the `notype` series. Please let me know if anything else comes up.",
      "votes": null
    },
    {
      "id": "2208241",
      "postDate": "04/03/2023 22:45:44",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I'm currently getting \"404 - Not Found\" errors when I attempt to download the updated dataset. This happens both when I attempt to download it via the kaggle CLI (\"kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction\") and when I attempt to download it via the download button at <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data</a>.</p>\n<p>Edit: The 404 errors seem to have been fixed. I'm redownloading the dataset now. Should be able to validate it within the next hour.<br>\nEdit 2: Just redownloaded the dataset and did a little sanity checking. It appears the original problem has been fixed but a new issue was introduced. I edited the main post to describe it.</p>",
      "rawMarkdown": "ryanholbrook I'm currently getting \"404 - Not Found\" errors when I attempt to download the updated dataset. This happens both when I attempt to download it via the kaggle CLI (\"kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction\") and when I attempt to download it via the download button at [https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data).\n\nEdit: The 404 errors seem to have been fixed. I'm redownloading the dataset now. Should be able to validate it within the next hour.\nEdit 2: Just redownloaded the dataset and did a little sanity checking. It appears the original problem has been fixed but a new issue was introduced. I edited the main post to describe it.",
      "votes": null
    },
    {
      "id": "2208742",
      "postDate": "04/04/2023 09:25:39",
      "content": "<p>I also found that there are disagreements when it comes to Init time and first timestamp (and vice versa for event end time) in tcdsfog for example \"003f117e14\". Have you found similar cases?</p>",
      "rawMarkdown": "I also found that there are disagreements when it comes to Init time and first timestamp (and vice versa for event end time) in tcdsfog for example \"003f117e14\". Have you found similar cases?",
      "votes": null
    },
    {
      "id": "2208924",
      "postDate": "04/04/2023 11:21:24",
      "content": "<p>003f117e14 looks fine to me. It has a single event which supposedly starts at 8.61312 seconds and a 128 Hz sample rate, so in theory that event should start at timestep 8.61312 * 128 = 1102.47. It actually starts at 1103, which seems close enough to me.</p>\n<p>Perhaps your math was thrown off by the fact tcdsfog has a 128 Hz sample rate, rather than 100 Hz? </p>",
      "rawMarkdown": "003f117e14 looks fine to me. It has a single event which supposedly starts at 8.61312 seconds and a 128 Hz sample rate, so in theory that event should start at timestep 8.61312 * 128 = 1102.47. It actually starts at 1103, which seems close enough to me.\n\nPerhaps your math was thrown off by the fact tcdsfog has a 128 Hz sample rate, rather than 100 Hz?",
      "votes": null
    },
    {
      "id": "2209644",
      "postDate": "04/04/2023 19:42:47",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a>, and apologies for the inconvenience. (It's been a hectic week.) I have a new update that I think should fix everything.</p>",
      "rawMarkdown": "Thanks, @jsday96, and apologies for the inconvenience. (It's been a hectic week.) I have a new update that I think should fix everything.",
      "votes": null
    },
    {
      "id": "2209844",
      "postDate": "04/05/2023 01:14:47",
      "content": "<p>Thanks! This version of the dataset looks good to me. It appears everything has been fixed.</p>",
      "rawMarkdown": "Thanks! This version of the dataset looks good to me. It appears everything has been fixed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2200004,
      "author_name": "zahidali06",
      "author_url": "",
      "post_date": "03/28/2023 07:46:39",
      "content": "<p>Hey, James I am also confused about the dataset because I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly. On which values can a human tell a particular subject has FoG? If you have an idea we can have an online meeting. Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 2200263,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "03/28/2023 12:24:12",
          "content": "<p>\"I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly\" - They are vertical, mediolateral, and anteroposterior accelerations. A diagram illustrating what that means was posted at <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/393852\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/393852</a>.</p>\n<p>\"On which values can a human tell a particular subject has FoG?\"</p>\n<ul>\n<li>The AccV , AccML and AccAP values can be used as features that are fed to a classification model or used to calculate more sophisticated features (like the frequency domain ones described in <a href=\"https://www.mdpi.com/1424-8220/19/23/5141\" target=\"_blank\">https://www.mdpi.com/1424-8220/19/23/5141</a> ).</li>\n<li>The StartHesitation, Turn, and Walking values in the defog and tdcsfog training datasets are useful labels for training classification models.</li>\n<li>The Valid and Task values in the defog training dataset are useful for filtering the training data to exclude timesteps that weren't annotated or may have been annotated inaccurately.</li>\n</ul>\n<p>This is all very tangential to the problem described in my original post. I think the documentation at <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data</a> is pretty clear but suspect the event labels in the notype training dataset are inaccurate.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2208198,
              "author_name": "ryanholbrook",
              "author_url": "",
              "post_date": "04/03/2023 21:45:49",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> , Thank you for your feedback. I updated the dataset with what should be the correct annotations for the <code>notype</code> series. Please let me know if anything else comes up.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2208241,
                  "author_name": "jsday96",
                  "author_url": "",
                  "post_date": "04/03/2023 22:45:44",
                  "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I'm currently getting \"404 - Not Found\" errors when I attempt to download the updated dataset. This happens both when I attempt to download it via the kaggle CLI (\"kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction\") and when I attempt to download it via the download button at <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data</a>.</p>\n<p>Edit: The 404 errors seem to have been fixed. I'm redownloading the dataset now. Should be able to validate it within the next hour.<br>\nEdit 2: Just redownloaded the dataset and did a little sanity checking. It appears the original problem has been fixed but a new issue was introduced. I edited the main post to describe it.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2209644,
                      "author_name": "ryanholbrook",
                      "author_url": "",
                      "post_date": "04/04/2023 19:42:47",
                      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a>, and apologies for the inconvenience. (It's been a hectic week.) I have a new update that I think should fix everything.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2209844,
                          "author_name": "jsday96",
                          "author_url": "",
                          "post_date": "04/05/2023 01:14:47",
                          "content": "<p>Thanks! This version of the dataset looks good to me. It appears everything has been fixed.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2208742,
      "author_name": "phantnguyen",
      "author_url": "",
      "post_date": "04/04/2023 09:25:39",
      "content": "<p>I also found that there are disagreements when it comes to Init time and first timestamp (and vice versa for event end time) in tcdsfog for example \"003f117e14\". Have you found similar cases?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2208924,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "04/04/2023 11:21:24",
          "content": "<p>003f117e14 looks fine to me. It has a single event which supposedly starts at 8.61312 seconds and a 128 Hz sample rate, so in theory that event should start at timestep 8.61312 * 128 = 1102.47. It actually starts at 1103, which seems close enough to me.</p>\n<p>Perhaps your math was thrown off by the fact tcdsfog has a 128 Hz sample rate, rather than 100 Hz? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2199812": "The entries in the \"events.csv\" file appear to be inconsistent with the \"Event\" column values in the notype training data files. It appears there may have been a unit conversion bug in the utility that populated the \"Event\" column of the train/notype/\\*.csv files.\n\nFor example, based on the row \"1e8d55d48d,1021.904,1022.598,,\" in events.csv, I would expect train/notype/1e8d55d48d.csv to contain an event in the time range 1021.904 - 1022.598 seconds. The data tab seems to indicate that the notype files use a 100 Hz sample rate, so I would expect that event to impact timesteps 102190 - 102260. However, when I go to those rows I see a ton of rows with no event...\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 102190 | -0.993268230764918 | -0.116984274771297 | -0.0413640801272263 | 0 | false | false |\n| 102191 | -0.993131260519859 | -0.117569239874438 | -0.0385470373678197 | 0 | false | false |\n| 102192 | -0.996553741019819 | -0.119481812857987 | -0.0380859375 | 0 | false | false |\n| ... | ... | ... | ... | ... | ... |\n| 102217 | -0.990922505974983 | -0.155768805534884 | -0.138887650788746 | 0 | false | false |\n| 102218 | -0.9871714142369 | -0.154445185733108 | -0.146245019082986 | 0 | false | false |\n| 102219 | -0.987959928058283 | -0.14827756563152 | -0.154395251826217 | 0 | false | false |\n| ... | ... | ... | ... | ... | ... |\n| 102258 | -0.869133339258985 | -0.136572456059369 | -0.34188932850263 | 0 | false | false |\n| 102259 | -0.861150736311346 | -0.142233667312971 | -0.353702279450288 | 0 | false | false |\n| 102260 | -0.85640032408877 | -0.153244516476253 | -0.357986408573962 | 0 | false | false |\n\n</p>However, if I scroll way up to time sample 1022, I see an extremely short event...\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 1021 | -0.878437304100535 | -0.0897442834092885 | 0.484561134573394 | 0 | true | true |\n| 1022 | -0.878560270712976 | -0.0903227291237997 | 0.484079230285104 | 1 | true | true |\n| 1023 | -0.876633455000578 | -0.0898869469325225 | 0.486403429901475 | 0 | true | true |\n\n</p>If I'm understanding the dataset correctly, time sample 1022 occurred (1022 samples) / (100 samples per second) = 10.22 seconds into the time series, not during the time range 1021.904 - 1022.598 seconds indicated by the events.csv file. Additionally, the \"events.csv\" file does not indicate that an event occurred around the time 10.22 seconds in series 1e8d55d48d. So, it appears it appears to me that the \"Event\" column was populated as if the sample rate was 1 Hz, rather than 100 Hz.\n\nThe example above is not an outlier. I've observed many similar incidents where the events.csv file and train/notype/\\*.csv files appear to disagree with one another. In general, this causes the train/notype/*.csv files to incidate that there were a lot of unrealistically short events early on in each time series, with no events in the last 99% of each file.\n\nI am aware of the fact there was a problem early on in the competition in which the \"Event\" column was completely unpopulated (set to 0), so I redownloaded the dataset 2 days ago. The original issue appears to have been fixed, but the problem described above is still present in what I believe to be the most recent version.\n\nAm I misunderstanding something or is there something wrong with the latest version of the notype data?\n\n# Edit: The fix introduced a new problem\n@ryanholbrook Uploaded an updated dataset in which the issue described above has been fixed. However, it appears there is a new problem in which the \"Time\" column contains some incorrect values.\n\nFor example, in the previous version of the dataset train/notype/affdf8553f.csv contained the following rows\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 208770 | -1.01594105348519 | -0.182570542573927 | 0.168195947263265 | 0 | true | true |\n| 208771 | -1.00286768203618 | -0.138492169717451 | 0.160291244375521 | 0 | true | true |\n| 208772 | -0.995033477474611 | -0.0994876348160082 | 0.135104737617867 | 0 | true | true |\n| 208773 | -0.993005946294685 | -0.0848971999139799 | 0.106476057956496 | 0 | true | true |\n| 208774 | -0.998477668146109 | -0.0886234945345456 | 0.0923842235424117 | 0 | true | true |\n| 208775 | -1.00423168766911 | -0.0936964872820936 | 0.0891615731894944 | 0 | true | true |\n\n</p> In the latest version, those rows of train/notype/affdf8553f.csv instead contain\n\n| Time | AccV | AccML | AccAP | Event | Valid | Task |\n| --- | --- | --- | --- | --- | --- |\n| 208769 | -1.01594105348519 | -0.182570542573927 | 0.168195947263265 | 0 | true | true |\n| 208771 | -1.00286768203618 | -0.138492169717451 | 0.160291244375521 | 0 | true | true |\n| 208771 | -0.995033477474611 | -0.0994876348160082 | 0.135104737617867 | 0 | true | true |\n| 208773 | -0.993005946294685 | -0.0848971999139799 | 0.106476057956496 | 0 | true | true |\n| 208773 | -0.998477668146109 | -0.0886234945345456 | 0.0923842235424117 | 0 | true | true |\n| 208775 | -1.00423168766911 | -0.0936964872820936 | 0.0891615731894944 | 1 | true | true |\n\n</p> Notice that there are some rows with duplicate \"Time\" values that don't align with the original. Once again, this isn't a completely isolated anomaly, but its a bit rarer than the issue I originally created this post for. I haven't validated the dataset thoroughly enough to be sure of this, but it seems to mostly impact events that occur relatively late in the files, so I have a suspicion it might have been caused by some sort of rounding bug that causes intermittent off by one errors for large time values. \n\n# Edit 2: It appears everything has been fixed\n@ryanholbrook Uploaded another iteration of the dataset. It appears all of the issues described above have been fixed. The latest dataset passes my sanity checks :-)",
    "2200004": "Hey, James I am also confused about the dataset because I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly. On which values can a human tell a particular subject has FoG? If you have an idea we can have an online meeting. Thanks",
    "2200263": "\"I am unable to understand the data and what the values of AccV , AccML and AccAP show exactly\" - They are vertical, mediolateral, and anteroposterior accelerations. A diagram illustrating what that means was posted at https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/393852.\n\n\"On which values can a human tell a particular subject has FoG?\"\n- The AccV , AccML and AccAP values can be used as features that are fed to a classification model or used to calculate more sophisticated features (like the frequency domain ones described in https://www.mdpi.com/1424-8220/19/23/5141 ).\n- The StartHesitation, Turn, and Walking values in the defog and tdcsfog training datasets are useful labels for training classification models.\n- The Valid and Task values in the defog training dataset are useful for filtering the training data to exclude timesteps that weren't annotated or may have been annotated inaccurately.\n\nThis is all very tangential to the problem described in my original post. I think the documentation at https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data is pretty clear but suspect the event labels in the notype training dataset are inaccurate.",
    "2208198": "Hi @jsday96 , Thank you for your feedback. I updated the dataset with what should be the correct annotations for the `notype` series. Please let me know if anything else comes up.",
    "2208241": "ryanholbrook I'm currently getting \"404 - Not Found\" errors when I attempt to download the updated dataset. This happens both when I attempt to download it via the kaggle CLI (\"kaggle competitions download -c tlvmc-parkinsons-freezing-gait-prediction\") and when I attempt to download it via the download button at [https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data).\n\nEdit: The 404 errors seem to have been fixed. I'm redownloading the dataset now. Should be able to validate it within the next hour.\nEdit 2: Just redownloaded the dataset and did a little sanity checking. It appears the original problem has been fixed but a new issue was introduced. I edited the main post to describe it.",
    "2208742": "I also found that there are disagreements when it comes to Init time and first timestamp (and vice versa for event end time) in tcdsfog for example \"003f117e14\". Have you found similar cases?",
    "2208924": "003f117e14 looks fine to me. It has a single event which supposedly starts at 8.61312 seconds and a 128 Hz sample rate, so in theory that event should start at timestep 8.61312 * 128 = 1102.47. It actually starts at 1103, which seems close enough to me.\n\nPerhaps your math was thrown off by the fact tcdsfog has a 128 Hz sample rate, rather than 100 Hz?",
    "2209644": "Thanks, @jsday96, and apologies for the inconvenience. (It's been a hectic week.) I have a new update that I think should fix everything.",
    "2209844": "Thanks! This version of the dataset looks good to me. It appears everything has been fixed."
  },
  "source": "meta"
}