{
  "id": 250070,
  "title": "Training CSV with paths to .npy files",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/250070",
  "author_name": "",
  "post_date": "2021-07-01T04:14:44.092875500Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In the input data, there is a complex system of subfolders, while in the <code>training_labels.csv</code> each data point is referenced by an ID corresponding to the file name.</p>\n<p>As we don't have filepaths, it is not straightforward to get a balanced subsample of training data.</p>\n<p>Just in case it can save someone some time, I shared <a href=\"https://www.kaggle.com/samusram/g2net-mapping-id-to-file-path\" target=\"_blank\">computation of a mapping from each ID to the full file path</a>. Per kind suggestion by Laura <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> , I turned the outputs into dataset. So that everyone would be able to quickly get the training CSV file containing full file paths to the raw input files. </p>\n<p>The dataset can found <a href=\"https://www.kaggle.com/samusram/g2net-gravitational-wave-detection-file-paths\" target=\"_blank\"><strong>here</strong></a>.<br>\nI provide a simple usage example in the dataset's notebook. Basically, the CSV can serve as a drop-in replacement for the official training data CSV:<br>\n<code>training_labels = pd.read_csv('../input/g2net-gravitational-wave-detection-file-paths/training_labels_with_paths.csv')</code>.</p>",
  "messages": [
    {
      "id": "1371515",
      "postDate": "07/01/2021 04:14:44",
      "content": "<p>In the input data, there is a complex system of subfolders, while in the <code>training_labels.csv</code> each data point is referenced by an ID corresponding to the file name.</p>\n<p>As we don't have filepaths, it is not straightforward to get a balanced subsample of training data.</p>\n<p>Just in case it can save someone some time, I shared <a href=\"https://www.kaggle.com/samusram/g2net-mapping-id-to-file-path\" target=\"_blank\">computation of a mapping from each ID to the full file path</a>. Per kind suggestion by Laura <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> , I turned the outputs into dataset. So that everyone would be able to quickly get the training CSV file containing full file paths to the raw input files. </p>\n<p>The dataset can found <a href=\"https://www.kaggle.com/samusram/g2net-gravitational-wave-detection-file-paths\" target=\"_blank\"><strong>here</strong></a>.<br>\nI provide a simple usage example in the dataset's notebook. Basically, the CSV can serve as a drop-in replacement for the official training data CSV:<br>\n<code>training_labels = pd.read_csv('../input/g2net-gravitational-wave-detection-file-paths/training_labels_with_paths.csv')</code>.</p>",
      "rawMarkdown": "In the input data, there is a complex system of subfolders, while in the `training_labels.csv` each data point is referenced by an ID corresponding to the file name.\n\nAs we don't have filepaths, it is not straightforward to get a balanced subsample of training data.\n\nJust in case it can save someone some time, I shared [computation of a mapping from each ID to the full file path](https://www.kaggle.com/samusram/g2net-mapping-id-to-file-path). Per kind suggestion by Laura @allunia , I turned the outputs into dataset. So that everyone would be able to quickly get the training CSV file containing full file paths to the raw input files. \n\nThe dataset can found [**here**](https://www.kaggle.com/samusram/g2net-gravitational-wave-detection-file-paths).\nI provide a simple usage example in the dataset's notebook. Basically, the CSV can serve as a drop-in replacement for the official training data CSV:\n`training_labels = pd.read_csv('../input/g2net-gravitational-wave-detection-file-paths/training_labels_with_paths.csv')`.",
      "votes": null
    },
    {
      "id": "1561306",
      "postDate": "10/27/2021 13:13:41",
      "content": "<p>Hey,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1561306,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 13:13:41",
      "content": "<p>Hey,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1371515": "In the input data, there is a complex system of subfolders, while in the `training_labels.csv` each data point is referenced by an ID corresponding to the file name.\n\nAs we don't have filepaths, it is not straightforward to get a balanced subsample of training data.\n\nJust in case it can save someone some time, I shared [computation of a mapping from each ID to the full file path](https://www.kaggle.com/samusram/g2net-mapping-id-to-file-path). Per kind suggestion by Laura @allunia , I turned the outputs into dataset. So that everyone would be able to quickly get the training CSV file containing full file paths to the raw input files. \n\nThe dataset can found [**here**](https://www.kaggle.com/samusram/g2net-gravitational-wave-detection-file-paths).\nI provide a simple usage example in the dataset's notebook. Basically, the CSV can serve as a drop-in replacement for the official training data CSV:\n`training_labels = pd.read_csv('../input/g2net-gravitational-wave-detection-file-paths/training_labels_with_paths.csv')`.",
    "1561306": "Hey,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}