{
  "id": 160693,
  "title": "MP3 Files stuck in memory?",
  "url": "/competitions/birdsong-recognition/discussion/160693",
  "author_name": "Mark Wijkhuizen",
  "post_date": "2020-06-22T10:08:04.150000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I did run into some memory usage problems when loading the files. When a file is loaded it seems to be kept in memory, even after it is closed. The script below reads all mp3 files, but does not return anything and should thus not cause RAM usage to grow.</p>\n\n<p>When running this sript locally the RAM usage does indeed not grow, however on a Kaggle notebook this is the case, which makes me wonder how the file management is done and if we can prevent this behaviour. I also asked a <a href=\"https://stackoverflow.com/questions/62511033/file-stays-in-memory-after-being-closed\">question on StackOverflow</a> regarding this problem and the behaviour in the Kaggle notebooks is indeed wrong.</p>\n\n<p>The current behaviour is highly inconvenient as we are unable to process 25GB of audio files with just 16GB of RAM when all audio files are kept in memory after being read.</p>\n\n<p>```\ndef read_mp3_as_bin(fname):\n    with open(fname, \"rb\") as f:\n        data = f.read() # when using 'pass' memory usage doesn't grow\n    print(f.closed)\n    return</p>\n\n<p>for fname in file_names: # file_names are 25K paths to the mp3 files\n    read_mp3_as_bin(fname)\n```</p>",
  "messages": [
    {
      "id": 896609,
      "postDate": "2020-06-22T10:08:04.150Z",
      "content": "<p>I did run into some memory usage problems when loading the files. When a file is loaded it seems to be kept in memory, even after it is closed. The script below reads all mp3 files, but does not return anything and should thus not cause RAM usage to grow.</p>\n\n<p>When running this sript locally the RAM usage does indeed not grow, however on a Kaggle notebook this is the case, which makes me wonder how the file management is done and if we can prevent this behaviour. I also asked a <a href=\"https://stackoverflow.com/questions/62511033/file-stays-in-memory-after-being-closed\">question on StackOverflow</a> regarding this problem and the behaviour in the Kaggle notebooks is indeed wrong.</p>\n\n<p>The current behaviour is highly inconvenient as we are unable to process 25GB of audio files with just 16GB of RAM when all audio files are kept in memory after being read.</p>\n\n<p>```\ndef read_mp3_as_bin(fname):\n    with open(fname, \"rb\") as f:\n        data = f.read() # when using 'pass' memory usage doesn't grow\n    print(f.closed)\n    return</p>\n\n<p>for fname in file_names: # file_names are 25K paths to the mp3 files\n    read_mp3_as_bin(fname)\n```</p>",
      "rawMarkdown": "I did run into some memory usage problems when loading the files. When a file is loaded it seems to be kept in memory, even after it is closed. The script below reads all mp3 files, but does not return anything and should thus not cause RAM usage to grow.\n\nWhen running this sript locally the RAM usage does indeed not grow, however on a Kaggle notebook this is the case, which makes me wonder how the file management is done and if we can prevent this behaviour. I also asked a [question on StackOverflow](https://stackoverflow.com/questions/62511033/file-stays-in-memory-after-being-closed) regarding this problem and the behaviour in the Kaggle notebooks is indeed wrong.\n\nThe current behaviour is highly inconvenient as we are unable to process 25GB of audio files with just 16GB of RAM when all audio files are kept in memory after being read.\n\n```\ndef read_mp3_as_bin(fname):\n    with open(fname, \"rb\") as f:\n        data = f.read() # when using 'pass' memory usage doesn't grow\n    print(f.closed)\n    return\n\nfor fname in file_names: # file_names are 25K paths to the mp3 files\n    read_mp3_as_bin(fname)\n```"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "896609": "I did run into some memory usage problems when loading the files. When a file is loaded it seems to be kept in memory, even after it is closed. The script below reads all mp3 files, but does not return anything and should thus not cause RAM usage to grow.\n\nWhen running this sript locally the RAM usage does indeed not grow, however on a Kaggle notebook this is the case, which makes me wonder how the file management is done and if we can prevent this behaviour. I also asked a [question on StackOverflow](https://stackoverflow.com/questions/62511033/file-stays-in-memory-after-being-closed) regarding this problem and the behaviour in the Kaggle notebooks is indeed wrong.\n\nThe current behaviour is highly inconvenient as we are unable to process 25GB of audio files with just 16GB of RAM when all audio files are kept in memory after being read.\n\n```\ndef read_mp3_as_bin(fname):\n    with open(fname, \"rb\") as f:\n        data = f.read() # when using 'pass' memory usage doesn't grow\n    print(f.closed)\n    return\n\nfor fname in file_names: # file_names are 25K paths to the mp3 files\n    read_mp3_as_bin(fname)\n```"
  }
}