{
  "id": 393837,
  "title": "Unusable Y variable?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/393837",
  "author_name": "samsupertaco",
  "post_date": "2023-03-11T00:01:40.870000",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I may be missing something, but why is every \"Event\" datapoint in the \"notype\" directory '0'? I thought they were supposed to be nonzero when any event occurred, but for every single file it is 0 all the way down. That makes the whole directory unusable, since that is the dependent variable…</p>\n<p>Am I wrong?</p>",
  "messages": [
    {
      "id": 2176821,
      "postDate": "2023-03-11T00:01:40.870Z",
      "content": "<p>I may be missing something, but why is every \"Event\" datapoint in the \"notype\" directory '0'? I thought they were supposed to be nonzero when any event occurred, but for every single file it is 0 all the way down. That makes the whole directory unusable, since that is the dependent variable…</p>\n<p>Am I wrong?</p>",
      "rawMarkdown": "I may be missing something, but why is every \"Event\" datapoint in the \"notype\" directory '0'? I thought they were supposed to be nonzero when any event occurred, but for every single file it is 0 all the way down. That makes the whole directory unusable, since that is the dependent variable...\n\nAm I wrong?",
      "votes": 4
    },
    {
      "id": 2176851,
      "postDate": "2023-03-11T01:11:08.020Z",
      "content": "<p>I confirm same - all 'Event' are 0 : <a href=\"https://www.kaggle.com/code/kretes/eda-task-valid-columns-in-tdcsfog?scriptVersionId=121722015\" target=\"_blank\">https://www.kaggle.com/code/kretes/eda-task-valid-columns-in-tdcsfog?scriptVersionId=121722015</a> </p>",
      "rawMarkdown": "I confirm same - all 'Event' are 0 : https://www.kaggle.com/code/kretes/eda-task-valid-columns-in-tdcsfog?scriptVersionId=121722015 ",
      "votes": 1
    },
    {
      "id": 2177642,
      "postDate": "2023-03-11T16:39:02.300Z",
      "content": "<p>After reading the entire documentation of this competition, I thought 'notype' should be treated as non-annotated data. It's use will only make sense for unsupervised or self-supervised techniques. The same goes for all lines in the CSV files in the other two folders if the 'Valid' or the 'Task' column are False.</p>",
      "rawMarkdown": "After reading the entire documentation of this competition, I thought 'notype' should be treated as non-annotated data. It's use will only make sense for unsupervised or self-supervised techniques. The same goes for all lines in the CSV files in the other two folders if the 'Valid' or the 'Task' column are False.",
      "votes": 2,
      "replies": [
        {
          "id": 2177686,
          "postDate": "2023-03-11T17:18:26.717Z",
          "content": "<p>Yes, but I thought it required those techniques to predict the specific event; in other words, that we still needed to know if an event occurred, and the unsupervised learning could decide which event it is, since that is what the goal of the competition is. Even if intentional, why is there a column in our dataset that is all zeroes? Seems like they just would’ve omitted it.</p>",
          "rawMarkdown": "Yes, but I thought it required those techniques to predict the specific event; in other words, that we still needed to know if an event occurred, and the unsupervised learning could decide which event it is, since that is what the goal of the competition is. Even if intentional, why is there a column in our dataset that is all zeroes? Seems like they just would’ve omitted it.",
          "votes": 1
        },
        {
          "id": 2178289,
          "postDate": "2023-03-12T09:43:15.063Z",
          "content": "<p>non-annotated data, that might be useful for non-supervised training is in <code>unlabeled</code> directory. This dataset was described as <code>daily</code>. <code>notype</code> supposed to be a dataset with just a binary annotation for whether any of (startHesitation, Turn, Walking) happened.</p>",
          "rawMarkdown": "non-annotated data, that might be useful for non-supervised training is in `unlabeled` directory. This dataset was described as `daily`. `notype` supposed to be a dataset with just a binary annotation for whether any of (startHesitation, Turn, Walking) happened.",
          "votes": 2,
          "replies": [
            {
              "id": 2178296,
              "postDate": "2023-03-12T09:51:49.783Z",
              "content": "<p>I had misunderstood myself.  Your explanation seems correct.</p>",
              "rawMarkdown": "I had misunderstood myself.  Your explanation seems correct."
            },
            {
              "id": 2194639,
              "postDate": "2023-03-24T04:32:00.300Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2177494,
      "postDate": "2023-03-11T14:16:30.040Z",
      "content": "<p>Thanks for the heads up. Let me look into it and I'll report back.</p>",
      "rawMarkdown": "Thanks for the heads up. Let me look into it and I'll report back.",
      "votes": 2,
      "replies": [
        {
          "id": 2179228,
          "postDate": "2023-03-13T03:26:40.543Z",
          "content": "<p>Any update?</p>",
          "rawMarkdown": "Any update?",
          "replies": [
            {
              "id": 2179984,
              "postDate": "2023-03-13T14:06:58.427Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/samsupertaco\" target=\"_blank\">@samsupertaco</a>, Looks like the Event label got dropped somehow during data prep for those series. I'll try to get this patched sometime soon. In the meantime, the <code>events.csv</code> file also has this information if you want to create it yourself.</p>",
              "rawMarkdown": "Hi @samsupertaco, Looks like the Event label got dropped somehow during data prep for those series. I'll try to get this patched sometime soon. In the meantime, the `events.csv` file also has this information if you want to create it yourself.",
              "votes": 1
            },
            {
              "id": 2181313,
              "postDate": "2023-03-14T13:06:05.460Z",
              "content": "<p>Ah, makes sense. Thanks for the help. </p>",
              "rawMarkdown": "Ah, makes sense. Thanks for the help. ",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2176851,
      "author_name": "Tomasz Bartczak",
      "author_url": "",
      "post_date": "2023-03-11T01:11:08.020000",
      "content": "<p>I confirm same - all 'Event' are 0 : <a href=\"https://www.kaggle.com/code/kretes/eda-task-valid-columns-in-tdcsfog?scriptVersionId=121722015\" target=\"_blank\">https://www.kaggle.com/code/kretes/eda-task-valid-columns-in-tdcsfog?scriptVersionId=121722015</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2177642,
      "author_name": "FelipeKitamura, MD, PhD",
      "author_url": "",
      "post_date": "2023-03-11T16:39:02.300000",
      "content": "<p>After reading the entire documentation of this competition, I thought 'notype' should be treated as non-annotated data. It's use will only make sense for unsupervised or self-supervised techniques. The same goes for all lines in the CSV files in the other two folders if the 'Valid' or the 'Task' column are False.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2177686,
          "author_name": "samsupertaco",
          "author_url": "",
          "post_date": "2023-03-11T17:18:26.717000",
          "content": "<p>Yes, but I thought it required those techniques to predict the specific event; in other words, that we still needed to know if an event occurred, and the unsupervised learning could decide which event it is, since that is what the goal of the competition is. Even if intentional, why is there a column in our dataset that is all zeroes? Seems like they just would’ve omitted it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2178289,
          "author_name": "Tomasz Bartczak",
          "author_url": "",
          "post_date": "2023-03-12T09:43:15.063000",
          "content": "<p>non-annotated data, that might be useful for non-supervised training is in <code>unlabeled</code> directory. This dataset was described as <code>daily</code>. <code>notype</code> supposed to be a dataset with just a binary annotation for whether any of (startHesitation, Turn, Walking) happened.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2178296,
              "author_name": "FelipeKitamura, MD, PhD",
              "author_url": "",
              "post_date": "2023-03-12T09:51:49.783000",
              "content": "<p>I had misunderstood myself.  Your explanation seems correct.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2194639,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-24T04:32:00.300000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2177494,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2023-03-11T14:16:30.040000",
      "content": "<p>Thanks for the heads up. Let me look into it and I'll report back.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2179228,
          "author_name": "samsupertaco",
          "author_url": "",
          "post_date": "2023-03-13T03:26:40.543000",
          "content": "<p>Any update?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2179984,
              "author_name": "Ryan Holbrook",
              "author_url": "",
              "post_date": "2023-03-13T14:06:58.427000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/samsupertaco\" target=\"_blank\">@samsupertaco</a>, Looks like the Event label got dropped somehow during data prep for those series. I'll try to get this patched sometime soon. In the meantime, the <code>events.csv</code> file also has this information if you want to create it yourself.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2181313,
              "author_name": "samsupertaco",
              "author_url": "",
              "post_date": "2023-03-14T13:06:05.460000",
              "content": "<p>Ah, makes sense. Thanks for the help. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2176821": "I may be missing something, but why is every \"Event\" datapoint in the \"notype\" directory '0'? I thought they were supposed to be nonzero when any event occurred, but for every single file it is 0 all the way down. That makes the whole directory unusable, since that is the dependent variable...\n\nAm I wrong?",
    "2176851": "I confirm same - all 'Event' are 0 : https://www.kaggle.com/code/kretes/eda-task-valid-columns-in-tdcsfog?scriptVersionId=121722015 ",
    "2177642": "After reading the entire documentation of this competition, I thought 'notype' should be treated as non-annotated data. It's use will only make sense for unsupervised or self-supervised techniques. The same goes for all lines in the CSV files in the other two folders if the 'Valid' or the 'Task' column are False.",
    "2177494": "Thanks for the heads up. Let me look into it and I'll report back."
  }
}