{
  "id": 394792,
  "title": "Information available during \"inference\"",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/394792",
  "author_name": "",
  "post_date": "2023-03-14T21:39:14.411310900Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Dear hosts,</p>\n<p>I very much appreciate that you provide a lot of information on the <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data\" target=\"_blank\">data</a> page. However, I want to make sure that I understand what kind of raw variables are available when predicting the samples from the test set for the public/private LB. </p>\n<ol>\n<li>An \"Id\" occurs only in exactly one of those sets: train, public test set, private test set? (Asking this because 003f117e14 occurs in the train AND test set in the interactive notebook version)</li>\n<li>For every \"Id\", irrespective in which of the 3 set it occurs (train, public test set, private test set), I can merge  subjects.csv to it and tdcsfog_metadata.csv or defog_metadata.csv subject to the data set the \"Id\" is coming from. Correct?</li>\n<li>Can you please elaborate on \"The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.\"? The words public/hidden confuses me, because we have a public version of these csv-files when opening a kaggle notebook and loading the data set. Then there might by a different public version for scoring the public LB and a third hidden/private version for scoring the private LB.</li>\n</ol>\n<p>Thank you very much!</p>",
  "messages": [
    {
      "id": "2182001",
      "postDate": "03/14/2023 21:39:14",
      "content": "<p>Dear hosts,</p>\n<p>I very much appreciate that you provide a lot of information on the <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data\" target=\"_blank\">data</a> page. However, I want to make sure that I understand what kind of raw variables are available when predicting the samples from the test set for the public/private LB. </p>\n<ol>\n<li>An \"Id\" occurs only in exactly one of those sets: train, public test set, private test set? (Asking this because 003f117e14 occurs in the train AND test set in the interactive notebook version)</li>\n<li>For every \"Id\", irrespective in which of the 3 set it occurs (train, public test set, private test set), I can merge  subjects.csv to it and tdcsfog_metadata.csv or defog_metadata.csv subject to the data set the \"Id\" is coming from. Correct?</li>\n<li>Can you please elaborate on \"The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.\"? The words public/hidden confuses me, because we have a public version of these csv-files when opening a kaggle notebook and loading the data set. Then there might by a different public version for scoring the public LB and a third hidden/private version for scoring the private LB.</li>\n</ol>\n<p>Thank you very much!</p>",
      "rawMarkdown": "Dear hosts,\n\nI very much appreciate that you provide a lot of information on the [data](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data) page. However, I want to make sure that I understand what kind of raw variables are available when predicting the samples from the test set for the public/private LB. \n\n1. An \"Id\" occurs only in exactly one of those sets: train, public test set, private test set? (Asking this because 003f117e14 occurs in the train AND test set in the interactive notebook version)\n2. For every \"Id\", irrespective in which of the 3 set it occurs (train, public test set, private test set), I can merge  subjects.csv to it and tdcsfog_metadata.csv or defog_metadata.csv subject to the data set the \"Id\" is coming from. Correct?\n3. Can you please elaborate on \"The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.\"? The words public/hidden confuses me, because we have a public version of these csv-files when opening a kaggle notebook and loading the data set. Then there might by a different public version for scoring the public LB and a third hidden/private version for scoring the private LB.\n\nThank you very much!",
      "votes": null
    },
    {
      "id": "2183008",
      "postDate": "03/15/2023 12:29:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dv0048\" target=\"_blank\">@dv0048</a> ,</p>\n<p>There are two versions of the dataset:</p>\n<ul>\n<li>The <em>public</em> version, which you see on the Data page and is available to you in interactive notebook sessions. The public version only contains training data. The files in <code>test</code> are only examples in the right format to help you write your code; you should ignore the ids and the actual contents.</li>\n<li>The <em>private</em> or <em>hidden</em> version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the <code>test</code> folder or out of the <code>sample_submission.csv</code> file.</li>\n</ul>\n<p>The hidden (or private) dataset has the entire test data. The words \"public\" and \"private\" here have nothing to do with the public and private leaderboards. You don't know which test id goes with the public or private leaderboard.</p>\n<ol>\n<li>In the hidden dataset, every id occurs only once, either in <code>train</code> or in <code>test</code> (or in <code>unlabeled</code>).</li>\n<li>Correct. This metadata is available for the test set series.</li>\n<li>This is just saying that these files don't contain any information about test set series, only the training set. The training set is the same in the public and hidden versions, so these files are also the same in both versions.</li>\n</ol>\n<p>Hope this answers your questions!</p>",
      "rawMarkdown": "Hi @dv0048 ,\n\nThere are two versions of the dataset:\n- The *public* version, which you see on the Data page and is available to you in interactive notebook sessions. The public version only contains training data. The files in `test` are only examples in the right format to help you write your code; you should ignore the ids and the actual contents.\n- The *private* or *hidden* version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the `test` folder or out of the `sample_submission.csv` file.\n\nThe hidden (or private) dataset has the entire test data. The words \"public\" and \"private\" here have nothing to do with the public and private leaderboards. You don't know which test id goes with the public or private leaderboard.\n\n1. In the hidden dataset, every id occurs only once, either in `train` or in `test` (or in `unlabeled`).\n2. Correct. This metadata is available for the test set series.\n3. This is just saying that these files don't contain any information about test set series, only the training set. The training set is the same in the public and hidden versions, so these files are also the same in both versions.\n\nHope this answers your questions!",
      "votes": null
    },
    {
      "id": "2183361",
      "postDate": "03/15/2023 15:52:11",
      "content": "<p>Yes. Thank you very much for the quick reply!</p>",
      "rawMarkdown": "Yes. Thank you very much for the quick reply!",
      "votes": null
    },
    {
      "id": "2223037",
      "postDate": "04/15/2023 18:41:47",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>\n<p>When you say:</p>\n<blockquote>\n  <p>\"The private or hidden version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the test folder or out of the sample_submission.csv file.\"</p>\n</blockquote>\n<p>I only see 2 unique trial \"Id\" values, 02ab235146 or 003f117e14 from the <code>test</code> folder and the <code>sample_submission.csv</code> file, but we are told in the <strong>Data Splits</strong> section of the \"Data\" tab that:</p>\n<blockquote>\n  <p>\"The test set contains about 250 data series. The series from the tdcsfog and defog are in a proportion similar to that of the training set. These series have subjects that are entirely distinct from those in the training set.\"</p>\n</blockquote>\n<p>So then how can we know what the Id's are for the 250 trials of the private testing set? How then will we accurately name the \"Id\" column in our outputted <code>submission.csv</code> file if we do not know which Id is related to the <em>private test</em> data that we predicting for? Can we obtain the metadata for these trial Id's or subject Id's? What are we actually predicting? (When I say \"trial\" or \"trial id\" I mean the event of time in which the subject is walking, turning, going through door ways, etc).</p>\n<p>I have successfully submitted two notebooks (and received poor scores) which simply have zeros for the \"StartHesitation\", \"Turn\", and \"Walking\" columns (essentially predicting no FOG events throughout the trials). I noticed that the submission score is only based on the <code>submission.csv</code> output that my notebook saves as a .csv with the shape of (286370, 4). I have no statistical model in my notebook submissions. How then is my <code>submission.csv</code> file used to determine the precision for all 250 hidden trials?</p>\n<p>I also noticed that the \"Data Split\" section states the following:</p>\n<blockquote>\n  <p>\"When your submission is scored, this example test data will be replaced with the full test set.\"</p>\n</blockquote>\n<p>What does this mean? What is being replaced? How are the 250 trials of the private test data being predicted on if we do not even have the accelerometer data for those 250 trials? </p>\n<p>I feel as though I am missing something very clear here, even after reading the documents thoroughly. I sincerely appreciate your time and effort in clarifying my questions.</p>",
      "rawMarkdown": "Hello @ryanholbrook \n\nWhen you say:\n\n>\"The private or hidden version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the test folder or out of the sample_submission.csv file.\"\n\nI only see 2 unique trial \"Id\" values, 02ab235146 or 003f117e14 from the `test` folder and the `sample_submission.csv` file, but we are told in the **Data Splits** section of the \"Data\" tab that:\n\n>\"The test set contains about 250 data series. The series from the tdcsfog and defog are in a proportion similar to that of the training set. These series have subjects that are entirely distinct from those in the training set.\"\n\nSo then how can we know what the Id's are for the 250 trials of the private testing set? How then will we accurately name the \"Id\" column in our outputted `submission.csv` file if we do not know which Id is related to the *private test* data that we predicting for? Can we obtain the metadata for these trial Id's or subject Id's? What are we actually predicting? (When I say \"trial\" or \"trial id\" I mean the event of time in which the subject is walking, turning, going through door ways, etc).\n\nI have successfully submitted two notebooks (and received poor scores) which simply have zeros for the \"StartHesitation\", \"Turn\", and \"Walking\" columns (essentially predicting no FOG events throughout the trials). I noticed that the submission score is only based on the `submission.csv` output that my notebook saves as a .csv with the shape of (286370, 4). I have no statistical model in my notebook submissions. How then is my `submission.csv` file used to determine the precision for all 250 hidden trials?\n\nI also noticed that the \"Data Split\" section states the following:\n\n>\"When your submission is scored, this example test data will be replaced with the full test set.\"\n\nWhat does this mean? What is being replaced? How are the 250 trials of the private test data being predicted on if we do not even have the accelerometer data for those 250 trials? \n\nI feel as though I am missing something very clear here, even after reading the documents thoroughly. I sincerely appreciate your time and effort in clarifying my questions.",
      "votes": null
    },
    {
      "id": "2224497",
      "postDate": "04/17/2023 11:38:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/williammcintoshpdx\" target=\"_blank\">@williammcintoshpdx</a> ,</p>\n<p>The files you see in the <code>test</code> set aren't the actual test set. They are just examples in the right format to help you write code. When you submit your notebook, we rerun it against a hidden test set, which you cannot see. This hidden test set has a <code>test</code> folder with all ~250 files like:</p>\n<pre><code>test/defog/\n   abc012.csv\n   def345.csv\n   ...\ntest/tdcsfog/\n   ghi678.csv\n   ...\n</code></pre>\n<p>Since you don't know the actual names of the files, you have to read them programmatically. (You could use <code>os.glob</code> or <code>pathlib.Path.glob</code>, say.)</p>\n<p>Hope this answers your question!</p>\n<p>Ryan</p>",
      "rawMarkdown": "Hi @williammcintoshpdx ,\n\nThe files you see in the `test` set aren't the actual test set. They are just examples in the right format to help you write code. When you submit your notebook, we rerun it against a hidden test set, which you cannot see. This hidden test set has a `test` folder with all ~250 files like:\n\n```\ntest/defog/\n   abc012.csv\n   def345.csv\n   ...\ntest/tdcsfog/\n   ghi678.csv\n   ...\n```\n\nSince you don't know the actual names of the files, you have to read them programmatically. (You could use `os.glob` or `pathlib.Path.glob`, say.)\n\nHope this answers your question!\n\nRyan",
      "votes": null
    },
    {
      "id": "2224738",
      "postDate": "04/17/2023 16:05:08",
      "content": "<p>OH! I think I better understand by your response. When we submit our notebook and set its corresponding output submission.csv file, the contents of the test \"folder\" are replaced during inference time. Is that correct?</p>",
      "rawMarkdown": "OH! I think I better understand by your response. When we submit our notebook and set its corresponding output submission.csv file, the contents of the test \"folder\" are replaced during inference time. Is that correct?",
      "votes": null
    },
    {
      "id": "2224828",
      "postDate": "04/17/2023 17:23:45",
      "content": "<p>Yep! Some of the other files change too if they contain info about the test set. The hidden <code>sample_submission.csv</code> for instance has the full set of ids needed for your <code>submission.csv</code>. See the <em>Data Splits</em> section in the data description for the details.</p>",
      "rawMarkdown": "Yep! Some of the other files change too if they contain info about the test set. The hidden `sample_submission.csv` for instance has the full set of ids needed for your `submission.csv`. See the *Data Splits* section in the data description for the details.",
      "votes": null
    },
    {
      "id": "2225381",
      "postDate": "04/18/2023 06:11:55",
      "content": "<p>Thks for asking such questions, my confusion is same as yours. this discussion really helps me a lot😄</p>",
      "rawMarkdown": "Thks for asking such questions, my confusion is same as yours. this discussion really helps me a lot😄",
      "votes": null
    },
    {
      "id": "2228827",
      "postDate": "04/20/2023 21:05:50",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>!</p>\n<p>Does this also mean that for the defog private test datasets that we will also have access to the <code>Tasks</code> table for those experiments during submission (inference) time?</p>",
      "rawMarkdown": "Thank you @ryanholbrook!\n\nDoes this also mean that for the defog private test datasets that we will also have access to the `Tasks` table for those experiments during submission (inference) time?",
      "votes": null
    },
    {
      "id": "2228873",
      "postDate": "04/20/2023 22:09:50",
      "content": "<p>No, as I mentioned above <code>The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.</code> from the data description implies that info from these files is only available for the training set.</p>",
      "rawMarkdown": "No, as I mentioned above `The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.` from the data description implies that info from these files is only available for the training set.",
      "votes": null
    },
    {
      "id": "2229756",
      "postDate": "04/21/2023 16:47:44",
      "content": "<p>okay thank you for the further clarification!</p>",
      "rawMarkdown": "okay thank you for the further clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2183008,
      "author_name": "ryanholbrook",
      "author_url": "",
      "post_date": "03/15/2023 12:29:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dv0048\" target=\"_blank\">@dv0048</a> ,</p>\n<p>There are two versions of the dataset:</p>\n<ul>\n<li>The <em>public</em> version, which you see on the Data page and is available to you in interactive notebook sessions. The public version only contains training data. The files in <code>test</code> are only examples in the right format to help you write your code; you should ignore the ids and the actual contents.</li>\n<li>The <em>private</em> or <em>hidden</em> version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the <code>test</code> folder or out of the <code>sample_submission.csv</code> file.</li>\n</ul>\n<p>The hidden (or private) dataset has the entire test data. The words \"public\" and \"private\" here have nothing to do with the public and private leaderboards. You don't know which test id goes with the public or private leaderboard.</p>\n<ol>\n<li>In the hidden dataset, every id occurs only once, either in <code>train</code> or in <code>test</code> (or in <code>unlabeled</code>).</li>\n<li>Correct. This metadata is available for the test set series.</li>\n<li>This is just saying that these files don't contain any information about test set series, only the training set. The training set is the same in the public and hidden versions, so these files are also the same in both versions.</li>\n</ol>\n<p>Hope this answers your questions!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2183361,
          "author_name": "dv0048",
          "author_url": "",
          "post_date": "03/15/2023 15:52:11",
          "content": "<p>Yes. Thank you very much for the quick reply!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2223037,
          "author_name": "williammcintoshpdx",
          "author_url": "",
          "post_date": "04/15/2023 18:41:47",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> </p>\n<p>When you say:</p>\n<blockquote>\n  <p>\"The private or hidden version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the test folder or out of the sample_submission.csv file.\"</p>\n</blockquote>\n<p>I only see 2 unique trial \"Id\" values, 02ab235146 or 003f117e14 from the <code>test</code> folder and the <code>sample_submission.csv</code> file, but we are told in the <strong>Data Splits</strong> section of the \"Data\" tab that:</p>\n<blockquote>\n  <p>\"The test set contains about 250 data series. The series from the tdcsfog and defog are in a proportion similar to that of the training set. These series have subjects that are entirely distinct from those in the training set.\"</p>\n</blockquote>\n<p>So then how can we know what the Id's are for the 250 trials of the private testing set? How then will we accurately name the \"Id\" column in our outputted <code>submission.csv</code> file if we do not know which Id is related to the <em>private test</em> data that we predicting for? Can we obtain the metadata for these trial Id's or subject Id's? What are we actually predicting? (When I say \"trial\" or \"trial id\" I mean the event of time in which the subject is walking, turning, going through door ways, etc).</p>\n<p>I have successfully submitted two notebooks (and received poor scores) which simply have zeros for the \"StartHesitation\", \"Turn\", and \"Walking\" columns (essentially predicting no FOG events throughout the trials). I noticed that the submission score is only based on the <code>submission.csv</code> output that my notebook saves as a .csv with the shape of (286370, 4). I have no statistical model in my notebook submissions. How then is my <code>submission.csv</code> file used to determine the precision for all 250 hidden trials?</p>\n<p>I also noticed that the \"Data Split\" section states the following:</p>\n<blockquote>\n  <p>\"When your submission is scored, this example test data will be replaced with the full test set.\"</p>\n</blockquote>\n<p>What does this mean? What is being replaced? How are the 250 trials of the private test data being predicted on if we do not even have the accelerometer data for those 250 trials? </p>\n<p>I feel as though I am missing something very clear here, even after reading the documents thoroughly. I sincerely appreciate your time and effort in clarifying my questions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2224497,
              "author_name": "ryanholbrook",
              "author_url": "",
              "post_date": "04/17/2023 11:38:26",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/williammcintoshpdx\" target=\"_blank\">@williammcintoshpdx</a> ,</p>\n<p>The files you see in the <code>test</code> set aren't the actual test set. They are just examples in the right format to help you write code. When you submit your notebook, we rerun it against a hidden test set, which you cannot see. This hidden test set has a <code>test</code> folder with all ~250 files like:</p>\n<pre><code>test/defog/\n   abc012.csv\n   def345.csv\n   ...\ntest/tdcsfog/\n   ghi678.csv\n   ...\n</code></pre>\n<p>Since you don't know the actual names of the files, you have to read them programmatically. (You could use <code>os.glob</code> or <code>pathlib.Path.glob</code>, say.)</p>\n<p>Hope this answers your question!</p>\n<p>Ryan</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2224738,
                  "author_name": "williammcintoshpdx",
                  "author_url": "",
                  "post_date": "04/17/2023 16:05:08",
                  "content": "<p>OH! I think I better understand by your response. When we submit our notebook and set its corresponding output submission.csv file, the contents of the test \"folder\" are replaced during inference time. Is that correct?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2224828,
                      "author_name": "ryanholbrook",
                      "author_url": "",
                      "post_date": "04/17/2023 17:23:45",
                      "content": "<p>Yep! Some of the other files change too if they contain info about the test set. The hidden <code>sample_submission.csv</code> for instance has the full set of ids needed for your <code>submission.csv</code>. See the <em>Data Splits</em> section in the data description for the details.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2228827,
                          "author_name": "williammcintoshpdx",
                          "author_url": "",
                          "post_date": "04/20/2023 21:05:50",
                          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>!</p>\n<p>Does this also mean that for the defog private test datasets that we will also have access to the <code>Tasks</code> table for those experiments during submission (inference) time?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2228873,
                              "author_name": "ryanholbrook",
                              "author_url": "",
                              "post_date": "04/20/2023 22:09:50",
                              "content": "<p>No, as I mentioned above <code>The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.</code> from the data description implies that info from these files is only available for the training set.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2229756,
                                  "author_name": "williammcintoshpdx",
                                  "author_url": "",
                                  "post_date": "04/21/2023 16:47:44",
                                  "content": "<p>okay thank you for the further clarification!</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2225381,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "04/18/2023 06:11:55",
      "content": "<p>Thks for asking such questions, my confusion is same as yours. this discussion really helps me a lot😄</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2182001": "Dear hosts,\n\nI very much appreciate that you provide a lot of information on the [data](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/data) page. However, I want to make sure that I understand what kind of raw variables are available when predicting the samples from the test set for the public/private LB. \n\n1. An \"Id\" occurs only in exactly one of those sets: train, public test set, private test set? (Asking this because 003f117e14 occurs in the train AND test set in the interactive notebook version)\n2. For every \"Id\", irrespective in which of the 3 set it occurs (train, public test set, private test set), I can merge  subjects.csv to it and tdcsfog_metadata.csv or defog_metadata.csv subject to the data set the \"Id\" is coming from. Correct?\n3. Can you please elaborate on \"The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.\"? The words public/hidden confuses me, because we have a public version of these csv-files when opening a kaggle notebook and loading the data set. Then there might by a different public version for scoring the public LB and a third hidden/private version for scoring the private LB.\n\nThank you very much!",
    "2183008": "Hi @dv0048 ,\n\nThere are two versions of the dataset:\n- The *public* version, which you see on the Data page and is available to you in interactive notebook sessions. The public version only contains training data. The files in `test` are only examples in the right format to help you write your code; you should ignore the ids and the actual contents.\n- The *private* or *hidden* version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the `test` folder or out of the `sample_submission.csv` file.\n\nThe hidden (or private) dataset has the entire test data. The words \"public\" and \"private\" here have nothing to do with the public and private leaderboards. You don't know which test id goes with the public or private leaderboard.\n\n1. In the hidden dataset, every id occurs only once, either in `train` or in `test` (or in `unlabeled`).\n2. Correct. This metadata is available for the test set series.\n3. This is just saying that these files don't contain any information about test set series, only the training set. The training set is the same in the public and hidden versions, so these files are also the same in both versions.\n\nHope this answers your questions!",
    "2183361": "Yes. Thank you very much for the quick reply!",
    "2223037": "Hello @ryanholbrook \n\nWhen you say:\n\n>\"The private or hidden version, which you cannot see but is available to your notebook when you submit. This version contains both the training data and the actual test data. To get the ids of the test series, you can either read them out of the test folder or out of the sample_submission.csv file.\"\n\nI only see 2 unique trial \"Id\" values, 02ab235146 or 003f117e14 from the `test` folder and the `sample_submission.csv` file, but we are told in the **Data Splits** section of the \"Data\" tab that:\n\n>\"The test set contains about 250 data series. The series from the tdcsfog and defog are in a proportion similar to that of the training set. These series have subjects that are entirely distinct from those in the training set.\"\n\nSo then how can we know what the Id's are for the 250 trials of the private testing set? How then will we accurately name the \"Id\" column in our outputted `submission.csv` file if we do not know which Id is related to the *private test* data that we predicting for? Can we obtain the metadata for these trial Id's or subject Id's? What are we actually predicting? (When I say \"trial\" or \"trial id\" I mean the event of time in which the subject is walking, turning, going through door ways, etc).\n\nI have successfully submitted two notebooks (and received poor scores) which simply have zeros for the \"StartHesitation\", \"Turn\", and \"Walking\" columns (essentially predicting no FOG events throughout the trials). I noticed that the submission score is only based on the `submission.csv` output that my notebook saves as a .csv with the shape of (286370, 4). I have no statistical model in my notebook submissions. How then is my `submission.csv` file used to determine the precision for all 250 hidden trials?\n\nI also noticed that the \"Data Split\" section states the following:\n\n>\"When your submission is scored, this example test data will be replaced with the full test set.\"\n\nWhat does this mean? What is being replaced? How are the 250 trials of the private test data being predicted on if we do not even have the accelerometer data for those 250 trials? \n\nI feel as though I am missing something very clear here, even after reading the documents thoroughly. I sincerely appreciate your time and effort in clarifying my questions.",
    "2224497": "Hi @williammcintoshpdx ,\n\nThe files you see in the `test` set aren't the actual test set. They are just examples in the right format to help you write code. When you submit your notebook, we rerun it against a hidden test set, which you cannot see. This hidden test set has a `test` folder with all ~250 files like:\n\n```\ntest/defog/\n   abc012.csv\n   def345.csv\n   ...\ntest/tdcsfog/\n   ghi678.csv\n   ...\n```\n\nSince you don't know the actual names of the files, you have to read them programmatically. (You could use `os.glob` or `pathlib.Path.glob`, say.)\n\nHope this answers your question!\n\nRyan",
    "2224738": "OH! I think I better understand by your response. When we submit our notebook and set its corresponding output submission.csv file, the contents of the test \"folder\" are replaced during inference time. Is that correct?",
    "2224828": "Yep! Some of the other files change too if they contain info about the test set. The hidden `sample_submission.csv` for instance has the full set of ids needed for your `submission.csv`. See the *Data Splits* section in the data description for the details.",
    "2225381": "Thks for asking such questions, my confusion is same as yours. this discussion really helps me a lot😄",
    "2228827": "Thank you @ryanholbrook!\n\nDoes this also mean that for the defog private test datasets that we will also have access to the `Tasks` table for those experiments during submission (inference) time?",
    "2228873": "No, as I mentioned above `The events.csv, defog_tasks.csv, and daily_metadata.csv files are the same in both the public and hidden datasets.` from the data description implies that info from these files is only available for the training set.",
    "2229756": "okay thank you for the further clarification!"
  },
  "source": "meta"
}