{
  "id": 467038,
  "title": "Do you have any tips on handling the large number of files",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467038",
  "author_name": "",
  "post_date": "2024-01-10T23:03:37.171047300Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Here is my current code and it uses up the 30GB quite quickly. Also this is only for the eegs there are so many other files:</p>\n<pre><code>trainEegDf = pd()\n\ndirectory = \ndirectoryFiles = os(directory)\ndirectoryFiles()\n\nreadFileArr = *(directoryFiles)\n\n idx  (((directoryFiles))):\n    path = os(directory, directoryFiles)\n\n    addDf = pd(path, engine=) \n    addDf = (path)\n\n    readFileArr = addDf()\n\ntrainEegDf = pd(np(readFileArr))\n</code></pre>",
  "messages": [
    {
      "id": "2596162",
      "postDate": "01/10/2024 23:03:37",
      "content": "<p>Here is my current code and it uses up the 30GB quite quickly. Also this is only for the eegs there are so many other files:</p>\n<pre><code>trainEegDf = pd()\n\ndirectory = \ndirectoryFiles = os(directory)\ndirectoryFiles()\n\nreadFileArr = *(directoryFiles)\n\n idx  (((directoryFiles))):\n    path = os(directory, directoryFiles)\n\n    addDf = pd(path, engine=) \n    addDf = (path)\n\n    readFileArr = addDf()\n\ntrainEegDf = pd(np(readFileArr))\n</code></pre>",
      "rawMarkdown": "Here is my current code and it uses up the 30GB quite quickly. Also this is only for the eegs there are so many other files:\n\n    trainEegDf = pd.DataFrame()\n\n    directory = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs\"\n    directoryFiles = os.listdir(directory)\n    directoryFiles.sort()\n\n    readFileArr = [0]*len(directoryFiles[:30])\n\n    for idx in tqdm(range(len(directoryFiles[:30]))):\n        path = os.path.join(directory, directoryFiles[idx])\n    \n        addDf = pd.read_parquet(path, engine=\"fastparquet\") \n        addDf['id'] = int(path[67:-8])\n    \n        readFileArr[idx] = addDf.to_numpy()\n    \n    trainEegDf = pd.DataFrame(np.concatenate(readFileArr))",
      "votes": null
    },
    {
      "id": "2596199",
      "postDate": "01/10/2024 23:49:17",
      "content": "<p>If you work with tensorflow- transform everything to TFRecords and then load in batches when you train.<br>\nIf you work with pytorch- do whatever  the equivalent is in pytorch…</p>",
      "rawMarkdown": "If you work with tensorflow- transform everything to TFRecords and then load in batches when you train.\nIf you work with pytorch- do whatever  the equivalent is in pytorch...",
      "votes": null
    },
    {
      "id": "2596884",
      "postDate": "01/11/2024 11:44:58",
      "content": "<p>What if I wanted to use a model like xgboost? I don't think it supports batch training</p>",
      "rawMarkdown": "What if I wanted to use a model like xgboost? I don't think it supports batch training",
      "votes": null
    },
    {
      "id": "2599547",
      "postDate": "01/12/2024 23:35:52",
      "content": "<p>if you can pay, consider using cloud-based Machine Learning services, such as Azure Machine Learning Service. Once your model is ready, you can then export/import into Kaggle. </p>",
      "rawMarkdown": "if you can pay, consider using cloud-based Machine Learning services, such as Azure Machine Learning Service. Once your model is ready, you can then export/import into Kaggle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2596199,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "01/10/2024 23:49:17",
      "content": "<p>If you work with tensorflow- transform everything to TFRecords and then load in batches when you train.<br>\nIf you work with pytorch- do whatever  the equivalent is in pytorch…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2596884,
          "author_name": "kawaiicoderuwu",
          "author_url": "",
          "post_date": "01/11/2024 11:44:58",
          "content": "<p>What if I wanted to use a model like xgboost? I don't think it supports batch training</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2599547,
      "author_name": "narareddy",
      "author_url": "",
      "post_date": "01/12/2024 23:35:52",
      "content": "<p>if you can pay, consider using cloud-based Machine Learning services, such as Azure Machine Learning Service. Once your model is ready, you can then export/import into Kaggle. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2596162": "Here is my current code and it uses up the 30GB quite quickly. Also this is only for the eegs there are so many other files:\n\n    trainEegDf = pd.DataFrame()\n\n    directory = \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs\"\n    directoryFiles = os.listdir(directory)\n    directoryFiles.sort()\n\n    readFileArr = [0]*len(directoryFiles[:30])\n\n    for idx in tqdm(range(len(directoryFiles[:30]))):\n        path = os.path.join(directory, directoryFiles[idx])\n    \n        addDf = pd.read_parquet(path, engine=\"fastparquet\") \n        addDf['id'] = int(path[67:-8])\n    \n        readFileArr[idx] = addDf.to_numpy()\n    \n    trainEegDf = pd.DataFrame(np.concatenate(readFileArr))",
    "2596199": "If you work with tensorflow- transform everything to TFRecords and then load in batches when you train.\nIf you work with pytorch- do whatever  the equivalent is in pytorch...",
    "2596884": "What if I wanted to use a model like xgboost? I don't think it supports batch training",
    "2599547": "if you can pay, consider using cloud-based Machine Learning services, such as Azure Machine Learning Service. Once your model is ready, you can then export/import into Kaggle."
  },
  "source": "meta"
}