{
  "id": 477920,
  "title": "RAM requirements",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477920",
  "author_name": "",
  "post_date": "2024-02-18T12:54:52.394925500Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Can anyone suggest a good way to minimize RAM after data split? Its always exceeding 30GB. Its causing problem while training the model. Any suggestions will be appreciated.</p>",
  "messages": [
    {
      "id": "2657322",
      "postDate": "02/18/2024 12:54:52",
      "content": "<p>Can anyone suggest a good way to minimize RAM after data split? Its always exceeding 30GB. Its causing problem while training the model. Any suggestions will be appreciated.</p>",
      "rawMarkdown": "Can anyone suggest a good way to minimize RAM after data split? Its always exceeding 30GB. Its causing problem while training the model. Any suggestions will be appreciated.",
      "votes": null
    },
    {
      "id": "2657444",
      "postDate": "02/18/2024 14:23:59",
      "content": "<p>Lazy load. There is no need to retain all data at memmory. Read and split ids and load the specific pointed files when you need them.</p>",
      "rawMarkdown": "Lazy load. There is no need to retain all data at memmory. Read and split ids and load the specific pointed files when you need them.",
      "votes": null
    },
    {
      "id": "2657461",
      "postDate": "02/18/2024 14:35:37",
      "content": "<p>Sort of the obvious option of buying/renting bigger hardware (I haven't gone that route), here are my two cents:</p>\n<ul>\n<li>Ensure data is stored using the \"right\" precision. Often python converts data to float64. I haven't yet found case in which more than float32 was needed. In this competition I wouldn't feel comfortable using float16, but I haven't tried it.</li>\n<li>Run experiments with a subset of the data loaded in memory. Once you figure out what settings you want to use, read data from disk instead of keeping it in memory. It will be a lot slower, but it will work. I'd like to note that this shortcut is not without risks. In the past I have discarded ideas because then didn't seem to make a difference (in a small dataset) to later find that they worked well when applied to a larger dataset.</li>\n</ul>\n<p>Good luck.</p>",
      "rawMarkdown": "Sort of the obvious option of buying/renting bigger hardware (I haven't gone that route), here are my two cents:\n- Ensure data is stored using the \"right\" precision. Often python converts data to float64. I haven't yet found case in which more than float32 was needed. In this competition I wouldn't feel comfortable using float16, but I haven't tried it.\n- Run experiments with a subset of the data loaded in memory. Once you figure out what settings you want to use, read data from disk instead of keeping it in memory. It will be a lot slower, but it will work. I'd like to note that this shortcut is not without risks. In the past I have discarded ideas because then didn't seem to make a difference (in a small dataset) to later find that they worked well when applied to a larger dataset.\n\nGood luck.",
      "votes": null
    },
    {
      "id": "2657555",
      "postDate": "02/18/2024 15:27:08",
      "content": "<p>If you want to use custom spectrograms, - convert them on the fly using torchaudio [e.g. torchaudio.transforms.MelSpectrogram], it can be done on GPU and is really fast, way faster than librosa implementation. </p>\n<p>As far as I remember, you can get same spectrogram from torchaudio compared to librosa if you set parameters right.</p>",
      "rawMarkdown": "If you want to use custom spectrograms, - convert them on the fly using torchaudio [e.g. torchaudio.transforms.MelSpectrogram], it can be done on GPU and is really fast, way faster than librosa implementation. \n\nAs far as I remember, you can get same spectrogram from torchaudio compared to librosa if you set parameters right.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2657444,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/18/2024 14:23:59",
      "content": "<p>Lazy load. There is no need to retain all data at memmory. Read and split ids and load the specific pointed files when you need them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2657461,
      "author_name": "vialactea",
      "author_url": "",
      "post_date": "02/18/2024 14:35:37",
      "content": "<p>Sort of the obvious option of buying/renting bigger hardware (I haven't gone that route), here are my two cents:</p>\n<ul>\n<li>Ensure data is stored using the \"right\" precision. Often python converts data to float64. I haven't yet found case in which more than float32 was needed. In this competition I wouldn't feel comfortable using float16, but I haven't tried it.</li>\n<li>Run experiments with a subset of the data loaded in memory. Once you figure out what settings you want to use, read data from disk instead of keeping it in memory. It will be a lot slower, but it will work. I'd like to note that this shortcut is not without risks. In the past I have discarded ideas because then didn't seem to make a difference (in a small dataset) to later find that they worked well when applied to a larger dataset.</li>\n</ul>\n<p>Good luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2657555,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "02/18/2024 15:27:08",
      "content": "<p>If you want to use custom spectrograms, - convert them on the fly using torchaudio [e.g. torchaudio.transforms.MelSpectrogram], it can be done on GPU and is really fast, way faster than librosa implementation. </p>\n<p>As far as I remember, you can get same spectrogram from torchaudio compared to librosa if you set parameters right.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2657322": "Can anyone suggest a good way to minimize RAM after data split? Its always exceeding 30GB. Its causing problem while training the model. Any suggestions will be appreciated.",
    "2657444": "Lazy load. There is no need to retain all data at memmory. Read and split ids and load the specific pointed files when you need them.",
    "2657461": "Sort of the obvious option of buying/renting bigger hardware (I haven't gone that route), here are my two cents:\n- Ensure data is stored using the \"right\" precision. Often python converts data to float64. I haven't yet found case in which more than float32 was needed. In this competition I wouldn't feel comfortable using float16, but I haven't tried it.\n- Run experiments with a subset of the data loaded in memory. Once you figure out what settings you want to use, read data from disk instead of keeping it in memory. It will be a lot slower, but it will work. I'd like to note that this shortcut is not without risks. In the past I have discarded ideas because then didn't seem to make a difference (in a small dataset) to later find that they worked well when applied to a larger dataset.\n\nGood luck.",
    "2657555": "If you want to use custom spectrograms, - convert them on the fly using torchaudio [e.g. torchaudio.transforms.MelSpectrogram], it can be done on GPU and is really fast, way faster than librosa implementation. \n\nAs far as I remember, you can get same spectrogram from torchaudio compared to librosa if you set parameters right."
  },
  "source": "meta"
}