{
  "id": 402189,
  "title": "Why is the test/defog data ('02ab235146.csv') included in the train/notype dataset?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/402189",
  "author_name": "",
  "post_date": "2023-04-17T09:38:43.102805600Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>There is a file with the same name '02ab235146.csv' as in the test/defog directory also present in the train/notype directory. Although the recorded accelerometer values are the same in both files, the file in the train/notype directory contains additional information regarding 'Event', 'Valid', and 'Task'. I'm currently unaware of how to utilize this information for prediction, but is it possible to use these details for prediction?</p>",
  "messages": [
    {
      "id": "2224420",
      "postDate": "04/17/2023 09:38:43",
      "content": "<p>There is a file with the same name '02ab235146.csv' as in the test/defog directory also present in the train/notype directory. Although the recorded accelerometer values are the same in both files, the file in the train/notype directory contains additional information regarding 'Event', 'Valid', and 'Task'. I'm currently unaware of how to utilize this information for prediction, but is it possible to use these details for prediction?</p>",
      "rawMarkdown": "There is a file with the same name '02ab235146.csv' as in the test/defog directory also present in the train/notype directory. Although the recorded accelerometer values are the same in both files, the file in the train/notype directory contains additional information regarding 'Event', 'Valid', and 'Task'. I'm currently unaware of how to utilize this information for prediction, but is it possible to use these details for prediction?",
      "votes": null
    },
    {
      "id": "2225277",
      "postDate": "04/18/2023 03:58:02",
      "content": "<p>There are indeed only three cols in \"testdata files\", but the rest cols are still useful. in <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/evaluation\" target=\"_blank\">evaluation</a>, you can view </p>\n<blockquote>\n  <p>Note that data series in the DeFOG dataset are annotated with Valid and Task labels (in addition to the event labels). Only the portions of the series where both are true should you consider to be annotated. Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).</p>\n</blockquote>\n<p>To another question, I'm not sure either. But I think after you read this notebook: <a href=\"https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets\" target=\"_blank\">Simple EDA on Time for targets</a>, you can notice something. the point is function <code>merge</code>, \"tasks\" is from \"tasks.csv\", the notebook writer merge it with file from test.👀</p>\n<pre><code>\n\nsub[] = \nsubmission = []\n f  test:\n    df = pd.read_csv(f)\n    df.set_index(, drop=, inplace=)\n\n    df[] = f.split()[-].split()[]\n\n    df[]=(df.index/df.index.()).values\n    df = pd.merge(df, tasks[[,]], how=, on=).fillna(-)\n\n    df = pd.merge(df, metadata_complex[[,]+[,,,]], how=, on=).fillna(-)\n    df_feats = fc.calculate(df, return_df=, include_final_window=, approve_sparsity=, window_idx=)\n    df = df.merge(df_feats, how=, left_index=, right_index=)\n    df.fillna(method=, inplace=)\n\n\n    res_vals=[]\n     i_fold  (N_FOLDS):\n        res_val=np.(regs[i_fold].predict(df[cols]).clip(,),)\n        res_vals.append(np.expand_dims(res_val,axis=))\n    res_vals=np.mean(np.concatenate(res_vals,axis=),axis=)\n    res = pd.DataFrame(res_vals, columns=pcols)\n\n    df = pd.concat([df,res], axis=)\n    df[] = df[].astype() +  + df.index.astype()\n    submission.append(df[scols])\nsubmission = pd.concat(submission)\nsubmission = pd.merge(sub[[]], submission, how=, on=).fillna()\nsubmission[scols].to_csv(, index=)\n</code></pre>\n<p>hope it's useful😄</p>",
      "rawMarkdown": "There are indeed only three cols in \"testdata files\", but the rest cols are still useful. in [evaluation](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/evaluation), you can view \n> Note that data series in the DeFOG dataset are annotated with Valid and Task labels (in addition to the event labels). Only the portions of the series where both are true should you consider to be annotated. Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).\n\nTo another question, I'm not sure either. But I think after you read this notebook: [Simple EDA on Time for targets](https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets), you can notice something. the point is function `merge`, \"tasks\" is from \"tasks.csv\", the notebook writer merge it with file from test.👀\n\n```python\n# this code is copied from the very end of the notebook mentioned before\n\nsub['t'] = 0\nsubmission = []\nfor f in test:\n    df = pd.read_csv(f)\n    df.set_index('Time', drop=True, inplace=True)\n\n    df['Id'] = f.split('/')[-1].split('.')[0]\n\n    df['Time_frac']=(df.index/df.index.max()).values#currently the index of data is actually \"Time\"\n    df = pd.merge(df, tasks[['Id','t_kmeans']], how='left', on='Id').fillna(-1)\n\n    df = pd.merge(df, metadata_complex[['Id','Subject']+['Visit','Test','Medication','s_kmeans']], how='left', on='Id').fillna(-1)\n    df_feats = fc.calculate(df, return_df=True, include_final_window=True, approve_sparsity=True, window_idx=\"begin\")\n    df = df.merge(df_feats, how=\"left\", left_index=True, right_index=True)\n    df.fillna(method=\"ffill\", inplace=True)\n\n    \n    res_vals=[]\n    for i_fold in range(N_FOLDS):\n        res_val=np.round(regs[i_fold].predict(df[cols]).clip(0.0,1.0),3)\n        res_vals.append(np.expand_dims(res_val,axis=2))\n    res_vals=np.mean(np.concatenate(res_vals,axis=2),axis=2)\n    res = pd.DataFrame(res_vals, columns=pcols)\n    \n    df = pd.concat([df,res], axis=1)\n    df['Id'] = df['Id'].astype(str) + '_' + df.index.astype(str)\n    submission.append(df[scols])\nsubmission = pd.concat(submission)\nsubmission = pd.merge(sub[['Id']], submission, how='left', on='Id').fillna(0.0)\nsubmission[scols].to_csv('submission.csv', index=False)\n```\n\nhope it's useful😄",
      "votes": null
    },
    {
      "id": "2226430",
      "postDate": "04/18/2023 22:53:22",
      "content": "<p>Thank you very much for your comment! I completely missed the sentence you pointed out… It was very helpful, thank you.</p>",
      "rawMarkdown": "Thank you very much for your comment! I completely missed the sentence you pointed out... It was very helpful, thank you.",
      "votes": null
    },
    {
      "id": "2229222",
      "postDate": "04/21/2023 07:00:20",
      "content": "<p>Hi Roger,</p>\n<p>I might not understand well this sentence \" Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).\". Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? Still I imagine that the input test data (the hidden data) for DeFOG will not have tha Valid and Tasks columns but for submission they know them and they will take into account only them?</p>",
      "rawMarkdown": "Hi Roger,\n\nI might not understand well this sentence \" Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).\". Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? Still I imagine that the input test data (the hidden data) for DeFOG will not have tha Valid and Tasks columns but for submission they know them and they will take into account only them?",
      "votes": null
    },
    {
      "id": "2231031",
      "postDate": "04/23/2023 03:07:11",
      "content": "<blockquote>\n  <p>Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? </p>\n</blockquote>\n<p>I don't think so👀, whether you \"fliter\" data by Valid and Tasks or not base on your model performance, however when you evaluation your model, you should to the \"fliter\" process as the host will do that😄</p>",
      "rawMarkdown": "> Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? \n\nI don't think so👀, whether you \"fliter\" data by Valid and Tasks or not base on your model performance, however when you evaluation your model, you should to the \"fliter\" process as the host will do that😄",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2225277,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "04/18/2023 03:58:02",
      "content": "<p>There are indeed only three cols in \"testdata files\", but the rest cols are still useful. in <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/evaluation\" target=\"_blank\">evaluation</a>, you can view </p>\n<blockquote>\n  <p>Note that data series in the DeFOG dataset are annotated with Valid and Task labels (in addition to the event labels). Only the portions of the series where both are true should you consider to be annotated. Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).</p>\n</blockquote>\n<p>To another question, I'm not sure either. But I think after you read this notebook: <a href=\"https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets\" target=\"_blank\">Simple EDA on Time for targets</a>, you can notice something. the point is function <code>merge</code>, \"tasks\" is from \"tasks.csv\", the notebook writer merge it with file from test.👀</p>\n<pre><code>\n\nsub[] = \nsubmission = []\n f  test:\n    df = pd.read_csv(f)\n    df.set_index(, drop=, inplace=)\n\n    df[] = f.split()[-].split()[]\n\n    df[]=(df.index/df.index.()).values\n    df = pd.merge(df, tasks[[,]], how=, on=).fillna(-)\n\n    df = pd.merge(df, metadata_complex[[,]+[,,,]], how=, on=).fillna(-)\n    df_feats = fc.calculate(df, return_df=, include_final_window=, approve_sparsity=, window_idx=)\n    df = df.merge(df_feats, how=, left_index=, right_index=)\n    df.fillna(method=, inplace=)\n\n\n    res_vals=[]\n     i_fold  (N_FOLDS):\n        res_val=np.(regs[i_fold].predict(df[cols]).clip(,),)\n        res_vals.append(np.expand_dims(res_val,axis=))\n    res_vals=np.mean(np.concatenate(res_vals,axis=),axis=)\n    res = pd.DataFrame(res_vals, columns=pcols)\n\n    df = pd.concat([df,res], axis=)\n    df[] = df[].astype() +  + df.index.astype()\n    submission.append(df[scols])\nsubmission = pd.concat(submission)\nsubmission = pd.merge(sub[[]], submission, how=, on=).fillna()\nsubmission[scols].to_csv(, index=)\n</code></pre>\n<p>hope it's useful😄</p>",
      "votes": null,
      "replies": [
        {
          "id": 2226430,
          "author_name": "dataanalojisan",
          "author_url": "",
          "post_date": "04/18/2023 22:53:22",
          "content": "<p>Thank you very much for your comment! I completely missed the sentence you pointed out… It was very helpful, thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2229222,
          "author_name": "quimquadrada",
          "author_url": "",
          "post_date": "04/21/2023 07:00:20",
          "content": "<p>Hi Roger,</p>\n<p>I might not understand well this sentence \" Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).\". Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? Still I imagine that the input test data (the hidden data) for DeFOG will not have tha Valid and Tasks columns but for submission they know them and they will take into account only them?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2231031,
              "author_name": "roger92",
              "author_url": "",
              "post_date": "04/23/2023 03:07:11",
              "content": "<blockquote>\n  <p>Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? </p>\n</blockquote>\n<p>I don't think so👀, whether you \"fliter\" data by Valid and Tasks or not base on your model performance, however when you evaluation your model, you should to the \"fliter\" process as the host will do that😄</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2224420": "There is a file with the same name '02ab235146.csv' as in the test/defog directory also present in the train/notype directory. Although the recorded accelerometer values are the same in both files, the file in the train/notype directory contains additional information regarding 'Event', 'Valid', and 'Task'. I'm currently unaware of how to utilize this information for prediction, but is it possible to use these details for prediction?",
    "2225277": "There are indeed only three cols in \"testdata files\", but the rest cols are still useful. in [evaluation](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/evaluation), you can view \n> Note that data series in the DeFOG dataset are annotated with Valid and Task labels (in addition to the event labels). Only the portions of the series where both are true should you consider to be annotated. Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).\n\nTo another question, I'm not sure either. But I think after you read this notebook: [Simple EDA on Time for targets](https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets), you can notice something. the point is function `merge`, \"tasks\" is from \"tasks.csv\", the notebook writer merge it with file from test.👀\n\n```python\n# this code is copied from the very end of the notebook mentioned before\n\nsub['t'] = 0\nsubmission = []\nfor f in test:\n    df = pd.read_csv(f)\n    df.set_index('Time', drop=True, inplace=True)\n\n    df['Id'] = f.split('/')[-1].split('.')[0]\n\n    df['Time_frac']=(df.index/df.index.max()).values#currently the index of data is actually \"Time\"\n    df = pd.merge(df, tasks[['Id','t_kmeans']], how='left', on='Id').fillna(-1)\n\n    df = pd.merge(df, metadata_complex[['Id','Subject']+['Visit','Test','Medication','s_kmeans']], how='left', on='Id').fillna(-1)\n    df_feats = fc.calculate(df, return_df=True, include_final_window=True, approve_sparsity=True, window_idx=\"begin\")\n    df = df.merge(df_feats, how=\"left\", left_index=True, right_index=True)\n    df.fillna(method=\"ffill\", inplace=True)\n\n    \n    res_vals=[]\n    for i_fold in range(N_FOLDS):\n        res_val=np.round(regs[i_fold].predict(df[cols]).clip(0.0,1.0),3)\n        res_vals.append(np.expand_dims(res_val,axis=2))\n    res_vals=np.mean(np.concatenate(res_vals,axis=2),axis=2)\n    res = pd.DataFrame(res_vals, columns=pcols)\n    \n    df = pd.concat([df,res], axis=1)\n    df['Id'] = df['Id'].astype(str) + '_' + df.index.astype(str)\n    submission.append(df[scols])\nsubmission = pd.concat(submission)\nsubmission = pd.merge(sub[['Id']], submission, how='left', on='Id').fillna(0.0)\nsubmission[scols].to_csv('submission.csv', index=False)\n```\n\nhope it's useful😄",
    "2226430": "Thank you very much for your comment! I completely missed the sentence you pointed out... It was very helpful, thank you.",
    "2229222": "Hi Roger,\n\nI might not understand well this sentence \" Though not included in the test set series, the metric is aware of these labels and will ignore predictions on the unannotated portions of these series (where either label is false).\". Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? Still I imagine that the input test data (the hidden data) for DeFOG will not have tha Valid and Tasks columns but for submission they know them and they will take into account only them?",
    "2231031": "> Does it mean that for DeFOG data we should train conisedring only valid the rows that ar both Valid and Tasks true? \n\nI don't think so👀, whether you \"fliter\" data by Valid and Tasks or not base on your model performance, however when you evaluation your model, you should to the \"fliter\" process as the host will do that😄"
  },
  "source": "meta"
}