{
  "id": 403470,
  "title": "Is the competition dataset SERIOUSLY WRONG?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/403470",
  "author_name": "Kha Vo",
  "post_date": "2023-04-23T10:04:33.255000",
  "votes": 62,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I see a common pattern in the defog dataset that events are present in a series, but both Task and Valid flag of those parts are False. That means these events won't take any effect at all during scoring. It is understood that where Task and Valid are both False, those parts are not annotated. </p>\n<p>So if they're not annotated, why are FOG events being labeled there? And what is the meaning of these labels?</p>\n<p>Another interesting pattern is that ALL StartHesitation events in defog dataset have Task and Valid as False (that means there is no StartHesitation in defog and should we just predict all 0?)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F3e57ddeb2288dd2af9a7664355416291%2Fdefogsample.png?generation=1682244053699840&amp;alt=media\" alt=\"\"></p>\n<p>EDIT:<br>\nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?</p>\n<p>Code to generate above plot:</p>\n<p>x = pd.read_csv('data/train/defog/e069a57511.csv')<br>\nplt.plot(x.AccV, label='AccV')<br>\nplt.plot(x.StartHesitation, label='StartHe')<br>\nplt.plot(x.Turn<em>2, label='Turn')\nplt.plot(x.Walking</em>3, label='Walking')<br>\nplt.plot(x.Task<em>-3, label='Task')\nplt.plot(x.Task</em>-1.5, label='Valid')<br>\nplt.title('defog sample e069a57511')<br>\nplt.legend();</p>\n<p><a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>",
  "messages": [
    {
      "id": 2231440,
      "postDate": "2023-04-23T10:04:33.257Z",
      "content": "<p>I see a common pattern in the defog dataset that events are present in a series, but both Task and Valid flag of those parts are False. That means these events won't take any effect at all during scoring. It is understood that where Task and Valid are both False, those parts are not annotated. </p>\n<p>So if they're not annotated, why are FOG events being labeled there? And what is the meaning of these labels?</p>\n<p>Another interesting pattern is that ALL StartHesitation events in defog dataset have Task and Valid as False (that means there is no StartHesitation in defog and should we just predict all 0?)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F3e57ddeb2288dd2af9a7664355416291%2Fdefogsample.png?generation=1682244053699840&amp;alt=media\" alt=\"\"></p>\n<p>EDIT:<br>\nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?</p>\n<p>Code to generate above plot:</p>\n<p>x = pd.read_csv('data/train/defog/e069a57511.csv')<br>\nplt.plot(x.AccV, label='AccV')<br>\nplt.plot(x.StartHesitation, label='StartHe')<br>\nplt.plot(x.Turn<em>2, label='Turn')\nplt.plot(x.Walking</em>3, label='Walking')<br>\nplt.plot(x.Task<em>-3, label='Task')\nplt.plot(x.Task</em>-1.5, label='Valid')<br>\nplt.title('defog sample e069a57511')<br>\nplt.legend();</p>\n<p><a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>",
      "rawMarkdown": "I see a common pattern in the defog dataset that events are present in a series, but both Task and Valid flag of those parts are False. That means these events won't take any effect at all during scoring. It is understood that where Task and Valid are both False, those parts are not annotated. \n\nSo if they're not annotated, why are FOG events being labeled there? And what is the meaning of these labels?\n\nAnother interesting pattern is that ALL StartHesitation events in defog dataset have Task and Valid as False (that means there is no StartHesitation in defog and should we just predict all 0?)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F3e57ddeb2288dd2af9a7664355416291%2Fdefogsample.png?generation=1682244053699840&alt=media)\n\nEDIT:\nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?\n\nCode to generate above plot:\n\nx = pd.read_csv('data/train/defog/e069a57511.csv')\nplt.plot(x.AccV, label='AccV')\nplt.plot(x.StartHesitation, label='StartHe')\nplt.plot(x.Turn*2, label='Turn')\nplt.plot(x.Walking*3, label='Walking')\nplt.plot(x.Task*-3, label='Task')\nplt.plot(x.Task*-1.5, label='Valid')\nplt.title('defog sample e069a57511')\nplt.legend();\n\n@jeffhausdorff @ryanholbrook ",
      "votes": 61
    },
    {
      "id": 2234757,
      "postDate": "2023-04-25T13:22:07.187Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a>,</p>\n<p>Thank you for the heads up. We are looking into it now.</p>",
      "rawMarkdown": "Hi @khahuras,\n\nThank you for the heads up. We are looking into it now.",
      "votes": 12,
      "replies": [
        {
          "id": 2235653,
          "postDate": "2023-04-26T08:25:45.913Z",
          "content": "<p>I think this means that there is a time lag between the sensor data and the correct label.</p>\n<p>It is likely that the start time of the sensor and the start time of the recorded data that the annotator uses for labeling are different.</p>\n<p>Other defog data also has a time lag.</p>",
          "rawMarkdown": "I think this means that there is a time lag between the sensor data and the correct label.\n\nIt is likely that the start time of the sensor and the start time of the recorded data that the annotator uses for labeling are different.\n\nOther defog data also has a time lag.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2244400,
      "postDate": "2023-05-03T15:48:41.090Z",
      "content": "<p>Data Update and Rescore by Kaggle Team<br>\n<a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406700\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406700</a></p>",
      "rawMarkdown": "Data Update and Rescore by Kaggle Team\nhttps://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406700",
      "votes": 7,
      "replies": [
        {
          "id": 2244590,
          "postDate": "2023-05-03T18:26:32.140Z",
          "content": "<p>Thanks for creating this topic <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> !</p>",
          "rawMarkdown": "Thanks for creating this topic @khahuras !",
          "votes": 2
        }
      ]
    },
    {
      "id": 2239041,
      "postDate": "2023-04-29T05:43:03.477Z",
      "content": "<p>I created <a href=\"https://www.kaggle.com/code/hiroakifukuse/pfogp-correct-defog-data-discrepancies\" target=\"_blank\">a notebook</a> to correct for the misalignment.<br>\nWhen I trained the model with that corrected data, loss decreased, but LB worsened.<br>\nAssuming a correlation between LB and CV, it may be possible to assume that the test data is also misaligned.</p>",
      "rawMarkdown": "I created [a notebook](https://www.kaggle.com/code/hiroakifukuse/pfogp-correct-defog-data-discrepancies) to correct for the misalignment.\nWhen I trained the model with that corrected data, loss decreased, but LB worsened.\nAssuming a correlation between LB and CV, it may be possible to assume that the test data is also misaligned.",
      "votes": 8,
      "replies": [
        {
          "id": 2242907,
          "postDate": "2023-05-02T15:14:19.217Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hiroakifukuse\" target=\"_blank\">@hiroakifukuse</a>, I think there is a way to try your method wisely. Let's train on your corrected data. Then predict on test data, and shift the predictions to misalign deliberately. Maybe you will have a very good score on LB!</p>",
          "rawMarkdown": "Hi @hiroakifukuse, I think there is a way to try your method wisely. Let's train on your corrected data. Then predict on test data, and shift the predictions to misalign deliberately. Maybe you will have a very good score on LB!",
          "votes": 3
        },
        {
          "id": 2244099,
          "postDate": "2023-05-03T12:48:53.363Z",
          "content": "<p>Considering the test set is just a superset of what we have, it is a likely prediction that it is misaligned as well (imagining that they had an unaccounted offset between video annotation and the recordings).</p>",
          "rawMarkdown": "Considering the test set is just a superset of what we have, it is a likely prediction that it is misaligned as well (imagining that they had an unaccounted offset between video annotation and the recordings)."
        }
      ]
    },
    {
      "id": 2238583,
      "postDate": "2023-04-28T16:40:18.033Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Can we get an update on this as soon as possible, or at least some information on whether this is confirmed or not? I know that validating something like that and communicating with the host might take a bit of time but I (and likely many others) find it hard to keep working on this competition not knowing whether the data is corrupt or not.</p>",
      "rawMarkdown": "@ryanholbrook Can we get an update on this as soon as possible, or at least some information on whether this is confirmed or not? I know that validating something like that and communicating with the host might take a bit of time but I (and likely many others) find it hard to keep working on this competition not knowing whether the data is corrupt or not.",
      "votes": 8,
      "replies": [
        {
          "id": 2238613,
          "postDate": "2023-04-28T17:07:05.723Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a>, I believe the problem has been diagnosed. We should have a fix up sometime next week. </p>",
          "rawMarkdown": "Hi @fritzcremer, I believe the problem has been diagnosed. We should have a fix up sometime next week. ",
          "votes": 17
        }
      ]
    },
    {
      "id": 2234420,
      "postDate": "2023-04-25T07:30:47.020Z",
      "content": "<p>I changed this topic's name to get more attention, as it is affecting all modelling steps, as well as scoring (average precision will be affected if this problem isn't answered or resolved)</p>\n<p><a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>",
      "rawMarkdown": "I changed this topic's name to get more attention, as it is affecting all modelling steps, as well as scoring (average precision will be affected if this problem isn't answered or resolved)\n\n@jeffhausdorff @ryanholbrook ",
      "votes": 3
    },
    {
      "id": 2234019,
      "postDate": "2023-04-24T19:23:03.783Z",
      "content": "<p>Is this alignment problem still there if one uses events.csv to get the labels?</p>",
      "rawMarkdown": "Is this alignment problem still there if one uses events.csv to get the labels?",
      "votes": 3
    },
    {
      "id": 2232704,
      "postDate": "2023-04-24T14:48:53.893Z",
      "content": "<p><a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?</p>",
      "rawMarkdown": "@jeffhausdorff @ryanholbrook \nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?",
      "votes": 3
    },
    {
      "id": 2241353,
      "postDate": "2023-05-01T12:17:13.077Z",
      "content": "<p>your assertion that: \"<em>ALL StartHesitation events in defog dataset have Task and Valid as False</em>\" is <strong>false</strong>. </p>\n<p>Time series 0d7ab3a9f9 in the defog train set do have a StartHesitation event with 88 timestamps where Valid an Task are both True (from Time 18150 -&gt; 18237)</p>",
      "rawMarkdown": "your assertion that: \"*ALL StartHesitation events in defog dataset have Task and Valid as False*\" is **false**. \n\nTime series 0d7ab3a9f9 in the defog train set do have a StartHesitation event with 88 timestamps where Valid an Task are both True (from Time 18150 -> 18237)",
      "votes": 1
    },
    {
      "id": 2238691,
      "postDate": "2023-04-28T18:30:45.380Z",
      "content": "<p>I do not see any mistake, your start hesitation color is blurring the pattern but still both Task and Valid are aligned with Turn green color.</p>",
      "rawMarkdown": "I do not see any mistake, your start hesitation color is blurring the pattern but still both Task and Valid are aligned with Turn green color.",
      "votes": 1
    },
    {
      "id": 2242078,
      "postDate": "2023-05-02T03:11:42.237Z",
      "content": "<p>Thank you for sharing! I also have the same question. After I set the label labeled as unavailable data to 0, LB significantly decreased.</p>\n<p>On the contrary, the code that ignores the Valid and Task tags has a higher LB.</p>",
      "rawMarkdown": "Thank you for sharing! I also have the same question. After I set the label labeled as unavailable data to 0, LB significantly decreased.\n\nOn the contrary, the code that ignores the Valid and Task tags has a higher LB.",
      "votes": 2
    },
    {
      "id": 2238519,
      "postDate": "2023-04-28T15:33:44.370Z",
      "content": "<p>Just started looking at this competition and noticed the same thing. Eg., looked at series <code>6041cad8ec</code> and events vs annotations line up much better if I shift the events by approximately 20000 samples which looks similar to your plot:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Fa9aef1a60229c38ea13e59296fa3ed29%2FUnshifted.png?generation=1682695860449698&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2F6d92efa11f168cd28d07777aecea0503%2FShifted.png?generation=1682695882283507&amp;alt=media\" alt=\"\"></p>\n<p>(<code>Event</code> is any event, and <code>Annotated</code> is <code>Valid</code> or <code>Task</code>.)</p>",
      "rawMarkdown": "Just started looking at this competition and noticed the same thing. Eg., looked at series `6041cad8ec` and events vs annotations line up much better if I shift the events by approximately 20000 samples which looks similar to your plot:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Fa9aef1a60229c38ea13e59296fa3ed29%2FUnshifted.png?generation=1682695860449698&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2F6d92efa11f168cd28d07777aecea0503%2FShifted.png?generation=1682695882283507&alt=media)\n\n(`Event` is any event, and `Annotated` is `Valid` or `Task`.)",
      "votes": 2,
      "replies": [
        {
          "id": 2242593,
          "postDate": "2023-05-02T11:41:02.273Z",
          "content": "<p>Interesting! Do you think that's the same for all samples or does it just happen to be 20k for this example? <br>\nHope this issue gets resolved quickly.</p>",
          "rawMarkdown": "Interesting! Do you think that's the same for all samples or does it just happen to be 20k for this example? \nHope this issue gets resolved quickly."
        }
      ]
    },
    {
      "id": 2237688,
      "postDate": "2023-04-27T21:28:29.117Z",
      "content": "<p>is there a typo? <br>\n<code>plt.plot(x.Task-1.5, label='Valid')</code>  x.Task --&gt; x.Valid</p>\n<p>Thanks, for pointing that issue btw</p>",
      "rawMarkdown": "is there a typo? \n`plt.plot(x.Task-1.5, label='Valid')`  x.Task --> x.Valid\n\nThanks, for pointing that issue btw",
      "votes": 2
    },
    {
      "id": 2294044,
      "postDate": "2023-06-09T17:18:44.913Z",
      "content": "<p>Thanks for pointing out this alignment problem <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a>!💪</p>",
      "rawMarkdown": "Thanks for pointing out this alignment problem @khahuras!💪"
    },
    {
      "id": 2254485,
      "postDate": "2023-05-11T03:33:21.783Z",
      "content": "<p>thx for sharing😁</p>",
      "rawMarkdown": "thx for sharing😁"
    },
    {
      "id": 2237714,
      "postDate": "2023-04-27T22:23:46.600Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2234757,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2023-04-25T13:22:07.187000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a>,</p>\n<p>Thank you for the heads up. We are looking into it now.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2235653,
          "author_name": "hisashi_h",
          "author_url": "",
          "post_date": "2023-04-26T08:25:45.913000",
          "content": "<p>I think this means that there is a time lag between the sensor data and the correct label.</p>\n<p>It is likely that the start time of the sensor and the start time of the recorded data that the annotator uses for labeling are different.</p>\n<p>Other defog data also has a time lag.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2244400,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2023-05-03T15:48:41.090000",
      "content": "<p>Data Update and Rescore by Kaggle Team<br>\n<a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406700\" target=\"_blank\">https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406700</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 2244590,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2023-05-03T18:26:32.140000",
          "content": "<p>Thanks for creating this topic <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a> !</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2239041,
      "author_name": "Hi F",
      "author_url": "",
      "post_date": "2023-04-29T05:43:03.477000",
      "content": "<p>I created <a href=\"https://www.kaggle.com/code/hiroakifukuse/pfogp-correct-defog-data-discrepancies\" target=\"_blank\">a notebook</a> to correct for the misalignment.<br>\nWhen I trained the model with that corrected data, loss decreased, but LB worsened.<br>\nAssuming a correlation between LB and CV, it may be possible to assume that the test data is also misaligned.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2242907,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2023-05-02T15:14:19.217000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hiroakifukuse\" target=\"_blank\">@hiroakifukuse</a>, I think there is a way to try your method wisely. Let's train on your corrected data. Then predict on test data, and shift the predictions to misalign deliberately. Maybe you will have a very good score on LB!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2244099,
          "author_name": "Kate",
          "author_url": "",
          "post_date": "2023-05-03T12:48:53.363000",
          "content": "<p>Considering the test set is just a superset of what we have, it is a likely prediction that it is misaligned as well (imagining that they had an unaccounted offset between video annotation and the recordings).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2238583,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2023-04-28T16:40:18.033000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Can we get an update on this as soon as possible, or at least some information on whether this is confirmed or not? I know that validating something like that and communicating with the host might take a bit of time but I (and likely many others) find it hard to keep working on this competition not knowing whether the data is corrupt or not.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2238613,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-04-28T17:07:05.723000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a>, I believe the problem has been diagnosed. We should have a fix up sometime next week. </p>",
          "votes": 17,
          "replies": []
        }
      ]
    },
    {
      "id": 2234420,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2023-04-25T07:30:47.020000",
      "content": "<p>I changed this topic's name to get more attention, as it is affecting all modelling steps, as well as scoring (average precision will be affected if this problem isn't answered or resolved)</p>\n<p><a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2234019,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2023-04-24T19:23:03.783000",
      "content": "<p>Is this alignment problem still there if one uses events.csv to get the labels?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2232704,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2023-04-24T14:48:53.893000",
      "content": "<p><a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2241353,
      "author_name": "Alberto Annoni",
      "author_url": "",
      "post_date": "2023-05-01T12:17:13.077000",
      "content": "<p>your assertion that: \"<em>ALL StartHesitation events in defog dataset have Task and Valid as False</em>\" is <strong>false</strong>. </p>\n<p>Time series 0d7ab3a9f9 in the defog train set do have a StartHesitation event with 88 timestamps where Valid an Task are both True (from Time 18150 -&gt; 18237)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2238691,
      "author_name": "Omar Nurhusien",
      "author_url": "",
      "post_date": "2023-04-28T18:30:45.380000",
      "content": "<p>I do not see any mistake, your start hesitation color is blurring the pattern but still both Task and Valid are aligned with Turn green color.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2242078,
      "author_name": "Lingduo Wang",
      "author_url": "",
      "post_date": "2023-05-02T03:11:42.237000",
      "content": "<p>Thank you for sharing! I also have the same question. After I set the label labeled as unavailable data to 0, LB significantly decreased.</p>\n<p>On the contrary, the code that ignores the Valid and Task tags has a higher LB.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2238519,
      "author_name": "Christoffer Karlsson",
      "author_url": "",
      "post_date": "2023-04-28T15:33:44.370000",
      "content": "<p>Just started looking at this competition and noticed the same thing. Eg., looked at series <code>6041cad8ec</code> and events vs annotations line up much better if I shift the events by approximately 20000 samples which looks similar to your plot:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Fa9aef1a60229c38ea13e59296fa3ed29%2FUnshifted.png?generation=1682695860449698&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2F6d92efa11f168cd28d07777aecea0503%2FShifted.png?generation=1682695882283507&amp;alt=media\" alt=\"\"></p>\n<p>(<code>Event</code> is any event, and <code>Annotated</code> is <code>Valid</code> or <code>Task</code>.)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2242593,
          "author_name": "Tim Wiesner",
          "author_url": "",
          "post_date": "2023-05-02T11:41:02.273000",
          "content": "<p>Interesting! Do you think that's the same for all samples or does it just happen to be 20k for this example? <br>\nHope this issue gets resolved quickly.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2237688,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2023-04-27T21:28:29.117000",
      "content": "<p>is there a typo? <br>\n<code>plt.plot(x.Task-1.5, label='Valid')</code>  x.Task --&gt; x.Valid</p>\n<p>Thanks, for pointing that issue btw</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2294044,
      "author_name": "Vladimir Simões da Luz Junior",
      "author_url": "",
      "post_date": "2023-06-09T17:18:44.913000",
      "content": "<p>Thanks for pointing out this alignment problem <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a>!💪</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2254485,
      "author_name": "马甲线来来来",
      "author_url": "",
      "post_date": "2023-05-11T03:33:21.783000",
      "content": "<p>thx for sharing😁</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2237714,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-27T22:23:46.600000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2231440": "I see a common pattern in the defog dataset that events are present in a series, but both Task and Valid flag of those parts are False. That means these events won't take any effect at all during scoring. It is understood that where Task and Valid are both False, those parts are not annotated. \n\nSo if they're not annotated, why are FOG events being labeled there? And what is the meaning of these labels?\n\nAnother interesting pattern is that ALL StartHesitation events in defog dataset have Task and Valid as False (that means there is no StartHesitation in defog and should we just predict all 0?)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1829450%2F3e57ddeb2288dd2af9a7664355416291%2Fdefogsample.png?generation=1682244053699840&alt=media)\n\nEDIT:\nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?\n\nCode to generate above plot:\n\nx = pd.read_csv('data/train/defog/e069a57511.csv')\nplt.plot(x.AccV, label='AccV')\nplt.plot(x.StartHesitation, label='StartHe')\nplt.plot(x.Turn*2, label='Turn')\nplt.plot(x.Walking*3, label='Walking')\nplt.plot(x.Task*-3, label='Task')\nplt.plot(x.Task*-1.5, label='Valid')\nplt.title('defog sample e069a57511')\nplt.legend();\n\n@jeffhausdorff @ryanholbrook ",
    "2234757": "Hi @khahuras,\n\nThank you for the heads up. We are looking into it now.",
    "2244400": "Data Update and Rescore by Kaggle Team\nhttps://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406700",
    "2239041": "I created [a notebook](https://www.kaggle.com/code/hiroakifukuse/pfogp-correct-defog-data-discrepancies) to correct for the misalignment.\nWhen I trained the model with that corrected data, loss decreased, but LB worsened.\nAssuming a correlation between LB and CV, it may be possible to assume that the test data is also misaligned.",
    "2238583": "@ryanholbrook Can we get an update on this as soon as possible, or at least some information on whether this is confirmed or not? I know that validating something like that and communicating with the host might take a bit of time but I (and likely many others) find it hard to keep working on this competition not knowing whether the data is corrupt or not.",
    "2234420": "I changed this topic's name to get more attention, as it is affecting all modelling steps, as well as scoring (average precision will be affected if this problem isn't answered or resolved)\n\n@jeffhausdorff @ryanholbrook ",
    "2234019": "Is this alignment problem still there if one uses events.csv to get the labels?",
    "2232704": "@jeffhausdorff @ryanholbrook \nIf you look at the green clusters on the plot, you'll see they should align (overlap) with the 2 clusters where Valid=Task=True just before them, instead of not-overlap. So is there any data mistake with this defog dataset?",
    "2241353": "your assertion that: \"*ALL StartHesitation events in defog dataset have Task and Valid as False*\" is **false**. \n\nTime series 0d7ab3a9f9 in the defog train set do have a StartHesitation event with 88 timestamps where Valid an Task are both True (from Time 18150 -> 18237)",
    "2238691": "I do not see any mistake, your start hesitation color is blurring the pattern but still both Task and Valid are aligned with Turn green color.",
    "2242078": "Thank you for sharing! I also have the same question. After I set the label labeled as unavailable data to 0, LB significantly decreased.\n\nOn the contrary, the code that ignores the Valid and Task tags has a higher LB.",
    "2238519": "Just started looking at this competition and noticed the same thing. Eg., looked at series `6041cad8ec` and events vs annotations line up much better if I shift the events by approximately 20000 samples which looks similar to your plot:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2Fa9aef1a60229c38ea13e59296fa3ed29%2FUnshifted.png?generation=1682695860449698&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F94532%2F6d92efa11f168cd28d07777aecea0503%2FShifted.png?generation=1682695882283507&alt=media)\n\n(`Event` is any event, and `Annotated` is `Valid` or `Task`.)",
    "2237688": "is there a typo? \n`plt.plot(x.Task-1.5, label='Valid')`  x.Task --> x.Valid\n\nThanks, for pointing that issue btw",
    "2294044": "Thanks for pointing out this alignment problem @khahuras!💪",
    "2254485": "thx for sharing😁",
    "2237714": ""
  }
}