{
  "id": 522988,
  "title": "Submission file",
  "url": "/competitions/bci-initiative-alvi-hci-challenge/discussion/522988",
  "author_name": "",
  "post_date": "2024-07-29T14:59:06.646950600Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Dear organizers,</p>\n<p>we are currently preparing our submission table and noticed that in the description (see <code>Overview</code> page), there is no information on the time dimension. Are we supposed to only hand in the angles of the final time step?</p>\n<p>A full example submission sheet would be very helpful.</p>\n<p>Best regards,<br>\nMoritz</p>",
  "messages": [
    {
      "id": "2939855",
      "postDate": "07/29/2024 14:59:06",
      "content": "<p>Dear organizers,</p>\n<p>we are currently preparing our submission table and noticed that in the description (see <code>Overview</code> page), there is no information on the time dimension. Are we supposed to only hand in the angles of the final time step?</p>\n<p>A full example submission sheet would be very helpful.</p>\n<p>Best regards,<br>\nMoritz</p>",
      "rawMarkdown": "Dear organizers,\n\nwe are currently preparing our submission table and noticed that in the description (see `Overview` page), there is no information on the time dimension. Are we supposed to only hand in the angles of the final time step?\n\nA full example submission sheet would be very helpful.\n\nBest regards,\nMoritz",
      "votes": null
    },
    {
      "id": "2939872",
      "postDate": "07/29/2024 15:10:06",
      "content": "<p>Hi! </p>\n<p>There's more details about the Data and how to prepare the submission file in the <a href=\"https://www.kaggle.com/competitions/bci-initiative-alvi-hci-challenge/data\" target=\"_blank\">Data</a> tab on this page. </p>\n<blockquote>\n  <p>The output of your model should be a 20-dimensional time-series with the predicted angle for each hand joint at each time stamp. <br>\n  The data, both EMG and hand pose, is sampled at 200Hz. However, each model's predictions should be down sampled to 25Hz. When preparing your submissions, make sure that your model's predictions are down sampled accordingly. We share some code to illustrate this in detail. </p>\n</blockquote>\n<p>The <a href=\"https://github.com/BCI-I/BCI_ALVI_challenge/blob/main/tutorials/04_submit_predictions.ipynb\" target=\"_blank\">example code</a> also shows how to prepare a submission. </p>\n<p>I hope this helps, let us know if you need any additional information!</p>",
      "rawMarkdown": "Hi! \n\nThere's more details about the Data and how to prepare the submission file in the [Data](https://www.kaggle.com/competitions/bci-initiative-alvi-hci-challenge/data) tab on this page. \n>The output of your model should be a 20-dimensional time-series with the predicted angle for each hand joint at each time stamp. \n>The data, both EMG and hand pose, is sampled at 200Hz. However, each model's predictions should be down sampled to 25Hz. When preparing your submissions, make sure that your model's predictions are down sampled accordingly. We share some code to illustrate this in detail. \n\n\nThe [example code](https://github.com/BCI-I/BCI_ALVI_challenge/blob/main/tutorials/04_submit_predictions.ipynb) also shows how to prepare a submission. \n\nI hope this helps, let us know if you need any additional information!",
      "votes": null
    },
    {
      "id": "2939943",
      "postDate": "07/29/2024 16:10:45",
      "content": "<p>Thank you very much for looking into the request. </p>\n<p>Indeed we were also reading the Data section and were therefore even more puzzled: </p>\n<p>The <code>submission file</code> example presented in the <code>Overview</code> section does not correspond to the example code you present:</p>\n<ul>\n<li>sample_id starts to be indexed with 0 (starts with 1 in the source code)</li>\n<li>it contains header names, which are not defined in the </li>\n<li>the example suggests 21 columns (sample_id and angles 0-19), assuming that each column name is unique (which is a common convention). given that each row seems to correspond to one of the 72 submit samples, there is no room for the time dimension.</li>\n</ul>\n<p>Given your response, I assume that the first row (column names) are discarded and there are not 21 but 21 * n_time_steps/8. This, however, would lead to a varying number of columns, since the <code>submit</code> samples have varying time dimensions.</p>\n<p>Since we don't have any insights into your evaluation logic, we would very much appreciate if you could clarify these deviations. </p>\n<p>Again, thank you very much for your time to look into this.</p>",
      "rawMarkdown": "Thank you very much for looking into the request. \n\nIndeed we were also reading the Data section and were therefore even more puzzled: \n\nThe `submission file` example presented in the `Overview` section does not correspond to the example code you present:\n- sample_id starts to be indexed with 0 (starts with 1 in the source code)\n- it contains header names, which are not defined in the \n- the example suggests 21 columns (sample_id and angles 0-19), assuming that each column name is unique (which is a common convention). given that each row seems to correspond to one of the 72 submit samples, there is no room for the time dimension.\n\nGiven your response, I assume that the first row (column names) are discarded and there are not 21 but 21 * n_time_steps/8. This, however, would lead to a varying number of columns, since the `submit` samples have varying time dimensions.\n\nSince we don't have any insights into your evaluation logic, we would very much appreciate if you could clarify these deviations. \n\nAgain, thank you very much for your time to look into this.",
      "votes": null
    },
    {
      "id": "2940123",
      "postDate": "07/29/2024 19:17:10",
      "content": "<p>Your predictions should be saved as a <code>.csv</code> file. <code>.csv</code> files typically have a header with the name of each column of data.<br>\nThe first column is called <code>sample_id</code> - <code>sample_id</code> should be a number between 1 and T with the id of each sample in the test dataset. If you're using a <code>DataFrame</code> to represent your model's predictions, you can add <code>sample_id</code> with:</p>\n<pre><code>df.insert(, , (,  + (df)))\n</code></pre>\n<p>There's no need to provide \"time\" information, all you need is the <code>sample_id</code> to match each sample in your predictions to the ground truth data. </p>\n<p>The other 20 columns have the name of the corresponding 20 joints/targets. The number is constant and the column names are the same for the entire dataset. The name of these columns doesn't matter: the order should match the order of the angles you're predicting.</p>\n<p>The number of rows in the <code>.csv</code> file depends on the number of samples your predictions generate. Keep in mind that predictions are downsampled relative to the inputs your model will receive. The example code shows one way in which such downsampling can be carried out. </p>\n<p>I hope this helps, please let me know otherwise. </p>",
      "rawMarkdown": "Your predictions should be saved as a `.csv` file. `.csv` files typically have a header with the name of each column of data.\nThe first column is called `sample_id` - `sample_id` should be a number between 1 and T with the id of each sample in the test dataset. If you're using a `DataFrame` to represent your model's predictions, you can add `sample_id` with:\n```python\ndf.insert(0, \"sample_id\", range(1, 1 + len(df)))\n```\nThere's no need to provide \"time\" information, all you need is the `sample_id` to match each sample in your predictions to the ground truth data. \n\nThe other 20 columns have the name of the corresponding 20 joints/targets. The number is constant and the column names are the same for the entire dataset. The name of these columns doesn't matter: the order should match the order of the angles you're predicting.\n\nThe number of rows in the `.csv` file depends on the number of samples your predictions generate. Keep in mind that predictions are downsampled relative to the inputs your model will receive. The example code shows one way in which such downsampling can be carried out. \n\n\nI hope this helps, please let me know otherwise.",
      "votes": null
    },
    {
      "id": "2942956",
      "postDate": "08/01/2024 08:10:42",
      "content": "<p>Thank you very much for explaining once more. My misunderstanding arose from the <code>sample_id</code> column, which seems to indicate that it holds a single row for each sample. In reality, the <code>sample_id</code> reflects timepoints from all 72 samples concatenated together.</p>\n<p>Again thank you very much for your time.</p>",
      "rawMarkdown": "Thank you very much for explaining once more. My misunderstanding arose from the `sample_id` column, which seems to indicate that it holds a single row for each sample. In reality, the `sample_id` reflects timepoints from all 72 samples concatenated together.\n\nAgain thank you very much for your time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2939872,
      "author_name": "bciinitiative",
      "author_url": "",
      "post_date": "07/29/2024 15:10:06",
      "content": "<p>Hi! </p>\n<p>There's more details about the Data and how to prepare the submission file in the <a href=\"https://www.kaggle.com/competitions/bci-initiative-alvi-hci-challenge/data\" target=\"_blank\">Data</a> tab on this page. </p>\n<blockquote>\n  <p>The output of your model should be a 20-dimensional time-series with the predicted angle for each hand joint at each time stamp. <br>\n  The data, both EMG and hand pose, is sampled at 200Hz. However, each model's predictions should be down sampled to 25Hz. When preparing your submissions, make sure that your model's predictions are down sampled accordingly. We share some code to illustrate this in detail. </p>\n</blockquote>\n<p>The <a href=\"https://github.com/BCI-I/BCI_ALVI_challenge/blob/main/tutorials/04_submit_predictions.ipynb\" target=\"_blank\">example code</a> also shows how to prepare a submission. </p>\n<p>I hope this helps, let us know if you need any additional information!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2939943,
          "author_name": "moritzschfer",
          "author_url": "",
          "post_date": "07/29/2024 16:10:45",
          "content": "<p>Thank you very much for looking into the request. </p>\n<p>Indeed we were also reading the Data section and were therefore even more puzzled: </p>\n<p>The <code>submission file</code> example presented in the <code>Overview</code> section does not correspond to the example code you present:</p>\n<ul>\n<li>sample_id starts to be indexed with 0 (starts with 1 in the source code)</li>\n<li>it contains header names, which are not defined in the </li>\n<li>the example suggests 21 columns (sample_id and angles 0-19), assuming that each column name is unique (which is a common convention). given that each row seems to correspond to one of the 72 submit samples, there is no room for the time dimension.</li>\n</ul>\n<p>Given your response, I assume that the first row (column names) are discarded and there are not 21 but 21 * n_time_steps/8. This, however, would lead to a varying number of columns, since the <code>submit</code> samples have varying time dimensions.</p>\n<p>Since we don't have any insights into your evaluation logic, we would very much appreciate if you could clarify these deviations. </p>\n<p>Again, thank you very much for your time to look into this.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2940123,
              "author_name": "bciinitiative",
              "author_url": "",
              "post_date": "07/29/2024 19:17:10",
              "content": "<p>Your predictions should be saved as a <code>.csv</code> file. <code>.csv</code> files typically have a header with the name of each column of data.<br>\nThe first column is called <code>sample_id</code> - <code>sample_id</code> should be a number between 1 and T with the id of each sample in the test dataset. If you're using a <code>DataFrame</code> to represent your model's predictions, you can add <code>sample_id</code> with:</p>\n<pre><code>df.insert(, , (,  + (df)))\n</code></pre>\n<p>There's no need to provide \"time\" information, all you need is the <code>sample_id</code> to match each sample in your predictions to the ground truth data. </p>\n<p>The other 20 columns have the name of the corresponding 20 joints/targets. The number is constant and the column names are the same for the entire dataset. The name of these columns doesn't matter: the order should match the order of the angles you're predicting.</p>\n<p>The number of rows in the <code>.csv</code> file depends on the number of samples your predictions generate. Keep in mind that predictions are downsampled relative to the inputs your model will receive. The example code shows one way in which such downsampling can be carried out. </p>\n<p>I hope this helps, please let me know otherwise. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2942956,
                  "author_name": "moritzschfer",
                  "author_url": "",
                  "post_date": "08/01/2024 08:10:42",
                  "content": "<p>Thank you very much for explaining once more. My misunderstanding arose from the <code>sample_id</code> column, which seems to indicate that it holds a single row for each sample. In reality, the <code>sample_id</code> reflects timepoints from all 72 samples concatenated together.</p>\n<p>Again thank you very much for your time.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2939855": "Dear organizers,\n\nwe are currently preparing our submission table and noticed that in the description (see `Overview` page), there is no information on the time dimension. Are we supposed to only hand in the angles of the final time step?\n\nA full example submission sheet would be very helpful.\n\nBest regards,\nMoritz",
    "2939872": "Hi! \n\nThere's more details about the Data and how to prepare the submission file in the [Data](https://www.kaggle.com/competitions/bci-initiative-alvi-hci-challenge/data) tab on this page. \n>The output of your model should be a 20-dimensional time-series with the predicted angle for each hand joint at each time stamp. \n>The data, both EMG and hand pose, is sampled at 200Hz. However, each model's predictions should be down sampled to 25Hz. When preparing your submissions, make sure that your model's predictions are down sampled accordingly. We share some code to illustrate this in detail. \n\n\nThe [example code](https://github.com/BCI-I/BCI_ALVI_challenge/blob/main/tutorials/04_submit_predictions.ipynb) also shows how to prepare a submission. \n\nI hope this helps, let us know if you need any additional information!",
    "2939943": "Thank you very much for looking into the request. \n\nIndeed we were also reading the Data section and were therefore even more puzzled: \n\nThe `submission file` example presented in the `Overview` section does not correspond to the example code you present:\n- sample_id starts to be indexed with 0 (starts with 1 in the source code)\n- it contains header names, which are not defined in the \n- the example suggests 21 columns (sample_id and angles 0-19), assuming that each column name is unique (which is a common convention). given that each row seems to correspond to one of the 72 submit samples, there is no room for the time dimension.\n\nGiven your response, I assume that the first row (column names) are discarded and there are not 21 but 21 * n_time_steps/8. This, however, would lead to a varying number of columns, since the `submit` samples have varying time dimensions.\n\nSince we don't have any insights into your evaluation logic, we would very much appreciate if you could clarify these deviations. \n\nAgain, thank you very much for your time to look into this.",
    "2940123": "Your predictions should be saved as a `.csv` file. `.csv` files typically have a header with the name of each column of data.\nThe first column is called `sample_id` - `sample_id` should be a number between 1 and T with the id of each sample in the test dataset. If you're using a `DataFrame` to represent your model's predictions, you can add `sample_id` with:\n```python\ndf.insert(0, \"sample_id\", range(1, 1 + len(df)))\n```\nThere's no need to provide \"time\" information, all you need is the `sample_id` to match each sample in your predictions to the ground truth data. \n\nThe other 20 columns have the name of the corresponding 20 joints/targets. The number is constant and the column names are the same for the entire dataset. The name of these columns doesn't matter: the order should match the order of the angles you're predicting.\n\nThe number of rows in the `.csv` file depends on the number of samples your predictions generate. Keep in mind that predictions are downsampled relative to the inputs your model will receive. The example code shows one way in which such downsampling can be carried out. \n\n\nI hope this helps, please let me know otherwise.",
    "2942956": "Thank you very much for explaining once more. My misunderstanding arose from the `sample_id` column, which seems to indicate that it holds a single row for each sample. In reality, the `sample_id` reflects timepoints from all 72 samples concatenated together.\n\nAgain thank you very much for your time."
  },
  "source": "meta"
}