{
  "id": 383734,
  "title": "Dataset splitting train_meta",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/383734",
  "author_name": "",
  "post_date": "2023-02-04T22:42:13.722224600Z",
  "votes": 17,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The train_meta.parquet is quite large and takes a long time to load, and consumes a lot of memory while you are holding it.  Since it is a parquet file (and the row group is quite large), you cannot easily read it batch by batch.  So I created a dataset that split this file into 660 files, one for each batch.</p>\n<p><a href=\"https://www.kaggle.com/datasets/solverworld/train-meta-parquet\" target=\"_blank\">train-meta-parquet</a></p>\n<p>If you want to see how it is used, see this notebook where I just read in 20 batches:<br>\n<a href=\"https://www.kaggle.com/code/solverworld/distribution-of-neutrino-arrival-directions\" target=\"_blank\">Distribution of Neutrino Arrival Directions</a></p>\n<p>The code to split it was </p>\n<pre><code>DATA_DIR='/path/to/train_meta file'\nOUTPUT_DIR=/path/to/place/output'\nMETA_FILE = os.path.join(DATA_DIR, 'train_meta.parquet')\nstart=time.time()            \nmeta_train = pd.read_parquet(os.path.join(DATA_DIR, 'train_meta.parquet'))\nfor idx, df in meta_train.groupby('batch_id'):\n    file = os.path.join(OUTPUT_DIR, 'meta', f'train_meta_{idx}.parquet')\n    df = df[['event_id', 'first_pulse_index', 'last_pulse_index', 'azimuth', 'zenith']]\n    df.to_parquet(file)\n    if idx &lt;= 2:\n        print(df)\n    print(f'{time.time()-start:8.1f} wrote {file}')\n</code></pre>",
  "messages": [
    {
      "id": "2129825",
      "postDate": "02/04/2023 22:42:13",
      "content": "<p>The train_meta.parquet is quite large and takes a long time to load, and consumes a lot of memory while you are holding it.  Since it is a parquet file (and the row group is quite large), you cannot easily read it batch by batch.  So I created a dataset that split this file into 660 files, one for each batch.</p>\n<p><a href=\"https://www.kaggle.com/datasets/solverworld/train-meta-parquet\" target=\"_blank\">train-meta-parquet</a></p>\n<p>If you want to see how it is used, see this notebook where I just read in 20 batches:<br>\n<a href=\"https://www.kaggle.com/code/solverworld/distribution-of-neutrino-arrival-directions\" target=\"_blank\">Distribution of Neutrino Arrival Directions</a></p>\n<p>The code to split it was </p>\n<pre><code>DATA_DIR='/path/to/train_meta file'\nOUTPUT_DIR=/path/to/place/output'\nMETA_FILE = os.path.join(DATA_DIR, 'train_meta.parquet')\nstart=time.time()            \nmeta_train = pd.read_parquet(os.path.join(DATA_DIR, 'train_meta.parquet'))\nfor idx, df in meta_train.groupby('batch_id'):\n    file = os.path.join(OUTPUT_DIR, 'meta', f'train_meta_{idx}.parquet')\n    df = df[['event_id', 'first_pulse_index', 'last_pulse_index', 'azimuth', 'zenith']]\n    df.to_parquet(file)\n    if idx &lt;= 2:\n        print(df)\n    print(f'{time.time()-start:8.1f} wrote {file}')\n</code></pre>",
      "rawMarkdown": "The train_meta.parquet is quite large and takes a long time to load, and consumes a lot of memory while you are holding it.  Since it is a parquet file (and the row group is quite large), you cannot easily read it batch by batch.  So I created a dataset that split this file into 660 files, one for each batch.\n\n[train-meta-parquet](https://www.kaggle.com/datasets/solverworld/train-meta-parquet)\n\nIf you want to see how it is used, see this notebook where I just read in 20 batches:\n[Distribution of Neutrino Arrival Directions](https://www.kaggle.com/code/solverworld/distribution-of-neutrino-arrival-directions)\n\nThe code to split it was \n```\nDATA_DIR='/path/to/train_meta file'\nOUTPUT_DIR=/path/to/place/output'\nMETA_FILE = os.path.join(DATA_DIR, 'train_meta.parquet')\nstart=time.time()            \nmeta_train = pd.read_parquet(os.path.join(DATA_DIR, 'train_meta.parquet'))\nfor idx, df in meta_train.groupby('batch_id'):\n    file = os.path.join(OUTPUT_DIR, 'meta', f'train_meta_{idx}.parquet')\n    df = df[['event_id', 'first_pulse_index', 'last_pulse_index', 'azimuth', 'zenith']]\n    df.to_parquet(file)\n    if idx <= 2:\n        print(df)\n    print(f'{time.time()-start:8.1f} wrote {file}')\n\n```",
      "votes": null
    },
    {
      "id": "2131209",
      "postDate": "02/06/2023 01:35:46",
      "content": "<p>Very cool.</p>\n<p>I also created a dataset containing the metadata for each batch of training data.</p>\n<p>You can find it here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/dschettler8845/batched-ndi-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/dschettler8845/batched-ndi-metadata</a></p>",
      "rawMarkdown": "Very cool.\n\nI also created a dataset containing the metadata for each batch of training data.\n\nYou can find it here:\n\nhttps://www.kaggle.com/datasets/dschettler8845/batched-ndi-metadata",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2131209,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "02/06/2023 01:35:46",
      "content": "<p>Very cool.</p>\n<p>I also created a dataset containing the metadata for each batch of training data.</p>\n<p>You can find it here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/dschettler8845/batched-ndi-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/dschettler8845/batched-ndi-metadata</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2129825": "The train_meta.parquet is quite large and takes a long time to load, and consumes a lot of memory while you are holding it.  Since it is a parquet file (and the row group is quite large), you cannot easily read it batch by batch.  So I created a dataset that split this file into 660 files, one for each batch.\n\n[train-meta-parquet](https://www.kaggle.com/datasets/solverworld/train-meta-parquet)\n\nIf you want to see how it is used, see this notebook where I just read in 20 batches:\n[Distribution of Neutrino Arrival Directions](https://www.kaggle.com/code/solverworld/distribution-of-neutrino-arrival-directions)\n\nThe code to split it was \n```\nDATA_DIR='/path/to/train_meta file'\nOUTPUT_DIR=/path/to/place/output'\nMETA_FILE = os.path.join(DATA_DIR, 'train_meta.parquet')\nstart=time.time()            \nmeta_train = pd.read_parquet(os.path.join(DATA_DIR, 'train_meta.parquet'))\nfor idx, df in meta_train.groupby('batch_id'):\n    file = os.path.join(OUTPUT_DIR, 'meta', f'train_meta_{idx}.parquet')\n    df = df[['event_id', 'first_pulse_index', 'last_pulse_index', 'azimuth', 'zenith']]\n    df.to_parquet(file)\n    if idx <= 2:\n        print(df)\n    print(f'{time.time()-start:8.1f} wrote {file}')\n\n```",
    "2131209": "Very cool.\n\nI also created a dataset containing the metadata for each batch of training data.\n\nYou can find it here:\n\nhttps://www.kaggle.com/datasets/dschettler8845/batched-ndi-metadata"
  },
  "source": "meta"
}