{
  "id": 513022,
  "title": "Memory Issue (Read parquet's files)",
  "url": "/competitions/uspto-explainable-ai/discussion/513022",
  "author_name": "",
  "post_date": "2024-06-18T07:42:33.788361400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Faced a problem with RAM usage.</p>\n<h1>Get information about each patent</h1>\n<p>All information about any patent is contained in parquet in the patent_data/ folder. When working with test data, we will not be able to upload all the information about each unique patent into RAM (we will not have enough of RAM). Therefore, 'as I understand it' we will have to work with patents step by step (we have processed 1 patent -&gt; move to the next one). I don't quite understand how RAM works.</p>\n<h1>Notebook issue</h1>\n<p>Here's a link to a notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/qurusx/memory-issue</a> where I elementary read information about each parquet in patent_data/. Why is RAM constantly increasing, eventually we end up with an excess of available RAM if we read the parquet -&gt; delete the reference (variable) to it -&gt; do rubbish collection (gc.collect). The most weighty file is 2023_6.parquet (1.09 GB In folder -&gt; 11.3 GB RAM), so any parquet can be loaded into RAM.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F2fc86baa7bd1b04d5edb2b1a86d73243%2Fphoto_2024-06-18_10-31-10.jpg?generation=1718696219062757&amp;alt=media\"></p>\n<h1>Questions</h1>\n<p>Is there any explanation for this? (RAM work in Python)<br>\nUsing only pandas (without polars and other libraries), is it possible to get information about the needed patent faster and less memory consuming than next method. Get publication_number in test.csv -&gt; get publication_date in patent_metadata.parquet -&gt; get info in patent_data/publication_date.parquet (Load full parquet)?<br>\n<strong>Thanks for any help, explanations, links</strong></p>",
  "messages": [
    {
      "id": "2877119",
      "postDate": "06/18/2024 07:42:33",
      "content": "<p>Faced a problem with RAM usage.</p>\n<h1>Get information about each patent</h1>\n<p>All information about any patent is contained in parquet in the patent_data/ folder. When working with test data, we will not be able to upload all the information about each unique patent into RAM (we will not have enough of RAM). Therefore, 'as I understand it' we will have to work with patents step by step (we have processed 1 patent -&gt; move to the next one). I don't quite understand how RAM works.</p>\n<h1>Notebook issue</h1>\n<p>Here's a link to a notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/qurusx/memory-issue</a> where I elementary read information about each parquet in patent_data/. Why is RAM constantly increasing, eventually we end up with an excess of available RAM if we read the parquet -&gt; delete the reference (variable) to it -&gt; do rubbish collection (gc.collect). The most weighty file is 2023_6.parquet (1.09 GB In folder -&gt; 11.3 GB RAM), so any parquet can be loaded into RAM.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F2fc86baa7bd1b04d5edb2b1a86d73243%2Fphoto_2024-06-18_10-31-10.jpg?generation=1718696219062757&amp;alt=media\"></p>\n<h1>Questions</h1>\n<p>Is there any explanation for this? (RAM work in Python)<br>\nUsing only pandas (without polars and other libraries), is it possible to get information about the needed patent faster and less memory consuming than next method. Get publication_number in test.csv -&gt; get publication_date in patent_metadata.parquet -&gt; get info in patent_data/publication_date.parquet (Load full parquet)?<br>\n<strong>Thanks for any help, explanations, links</strong></p>",
      "rawMarkdown": "Faced a problem with RAM usage.\n\n# Get information about each patent\nAll information about any patent is contained in parquet in the patent_data/ folder. When working with test data, we will not be able to upload all the information about each unique patent into RAM (we will not have enough of RAM). Therefore, 'as I understand it' we will have to work with patents step by step (we have processed 1 patent -> move to the next one). I don't quite understand how RAM works.\n\n# Notebook issue\nHere's a link to a notebook [https://www.kaggle.com/code/qurusx/memory-issue](url) where I elementary read information about each parquet in patent_data/. Why is RAM constantly increasing, eventually we end up with an excess of available RAM if we read the parquet -> delete the reference (variable) to it -> do rubbish collection (gc.collect). The most weighty file is 2023_6.parquet (1.09 GB In folder -> 11.3 GB RAM), so any parquet can be loaded into RAM.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F2fc86baa7bd1b04d5edb2b1a86d73243%2Fphoto_2024-06-18_10-31-10.jpg?generation=1718696219062757&alt=media)\n\n# Questions\nIs there any explanation for this? (RAM work in Python)\nUsing only pandas (without polars and other libraries), is it possible to get information about the needed patent faster and less memory consuming than next method. Get publication_number in test.csv -> get publication_date in patent_metadata.parquet -> get info in patent_data/publication_date.parquet (Load full parquet)?\n**Thanks for any help, explanations, links**",
      "votes": null
    },
    {
      "id": "2878565",
      "postDate": "06/19/2024 04:26:48",
      "content": "<p>Garbage collector in Python doesn't always fully release unused memory. There are ways to deal with it, e.g. <br>\nlibc = ctypes.CDLL(\"libc.so.6\") and then libc.malloc_trim(0)</p>\n<p>You may read the files by chunks so that they definitely fit into the memory. Even better yet, you may read only required columns, that will improve both memory usage and performance.</p>",
      "rawMarkdown": "Garbage collector in Python doesn't always fully release unused memory. There are ways to deal with it, e.g. \nlibc = ctypes.CDLL(\"libc.so.6\") and then libc.malloc_trim(0)\n\nYou may read the files by chunks so that they definitely fit into the memory. Even better yet, you may read only required columns, that will improve both memory usage and performance.",
      "votes": null
    },
    {
      "id": "2878601",
      "postDate": "06/19/2024 04:49:45",
      "content": "<p>Thanks, I'll try to apply the memory tip.</p>\n<p>If you mean reading the parquet line by line or chunk by chunk, it will take a long time to find the right line. I usually load the whole parquet, set 'publication_number' as an index and quickly find the required row by index. Doesn't seem to be a faster way?</p>\n<p>As for columns, I want to work not only with title and abstract. As far as I know, loading only 'publication_number' column and making it an index, we can't load the rest of information (claims, description etc.) for the line we NEED. That's why we have to load the whole parquet.</p>",
      "rawMarkdown": "Thanks, I'll try to apply the memory tip.\n\nIf you mean reading the parquet line by line or chunk by chunk, it will take a long time to find the right line. I usually load the whole parquet, set 'publication_number' as an index and quickly find the required row by index. Doesn't seem to be a faster way?\n\nAs for columns, I want to work not only with title and abstract. As far as I know, loading only 'publication_number' column and making it an index, we can't load the rest of information (claims, description etc.) for the line we NEED. That's why we have to load the whole parquet.",
      "votes": null
    },
    {
      "id": "2878909",
      "postDate": "06/19/2024 08:35:01",
      "content": "<p>Try using polars it will reduce the memory ,As I faced this issue at the beginning of the competition , And of course this competition needs a good knowledge of how to deal with memory </p>",
      "rawMarkdown": "Try using polars it will reduce the memory ,As I faced this issue at the beginning of the competition , And of course this competition needs a good knowledge of how to deal with memory",
      "votes": null
    },
    {
      "id": "2878917",
      "postDate": "06/19/2024 08:42:40",
      "content": "<p>For example, there is a way to split publication_date.parquet into smaller subfiles and save them, and refer to the dictionary of patent-&gt;filename.</p>\n<p>publication_number -&gt; dict_patent2file -&gt; smaller_target_file</p>",
      "rawMarkdown": "For example, there is a way to split publication_date.parquet into smaller subfiles and save them, and refer to the dictionary of patent->filename.\n\npublication_number -> dict_patent2file -> smaller_target_file",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2878565,
      "author_name": "apparition",
      "author_url": "",
      "post_date": "06/19/2024 04:26:48",
      "content": "<p>Garbage collector in Python doesn't always fully release unused memory. There are ways to deal with it, e.g. <br>\nlibc = ctypes.CDLL(\"libc.so.6\") and then libc.malloc_trim(0)</p>\n<p>You may read the files by chunks so that they definitely fit into the memory. Even better yet, you may read only required columns, that will improve both memory usage and performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2878601,
          "author_name": "qurusx",
          "author_url": "",
          "post_date": "06/19/2024 04:49:45",
          "content": "<p>Thanks, I'll try to apply the memory tip.</p>\n<p>If you mean reading the parquet line by line or chunk by chunk, it will take a long time to find the right line. I usually load the whole parquet, set 'publication_number' as an index and quickly find the required row by index. Doesn't seem to be a faster way?</p>\n<p>As for columns, I want to work not only with title and abstract. As far as I know, loading only 'publication_number' column and making it an index, we can't load the rest of information (claims, description etc.) for the line we NEED. That's why we have to load the whole parquet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2878909,
      "author_name": "mahmoudsaber",
      "author_url": "",
      "post_date": "06/19/2024 08:35:01",
      "content": "<p>Try using polars it will reduce the memory ,As I faced this issue at the beginning of the competition , And of course this competition needs a good knowledge of how to deal with memory </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2878917,
      "author_name": "ryotayoshinobu",
      "author_url": "",
      "post_date": "06/19/2024 08:42:40",
      "content": "<p>For example, there is a way to split publication_date.parquet into smaller subfiles and save them, and refer to the dictionary of patent-&gt;filename.</p>\n<p>publication_number -&gt; dict_patent2file -&gt; smaller_target_file</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2877119": "Faced a problem with RAM usage.\n\n# Get information about each patent\nAll information about any patent is contained in parquet in the patent_data/ folder. When working with test data, we will not be able to upload all the information about each unique patent into RAM (we will not have enough of RAM). Therefore, 'as I understand it' we will have to work with patents step by step (we have processed 1 patent -> move to the next one). I don't quite understand how RAM works.\n\n# Notebook issue\nHere's a link to a notebook [https://www.kaggle.com/code/qurusx/memory-issue](url) where I elementary read information about each parquet in patent_data/. Why is RAM constantly increasing, eventually we end up with an excess of available RAM if we read the parquet -> delete the reference (variable) to it -> do rubbish collection (gc.collect). The most weighty file is 2023_6.parquet (1.09 GB In folder -> 11.3 GB RAM), so any parquet can be loaded into RAM.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F2fc86baa7bd1b04d5edb2b1a86d73243%2Fphoto_2024-06-18_10-31-10.jpg?generation=1718696219062757&alt=media)\n\n# Questions\nIs there any explanation for this? (RAM work in Python)\nUsing only pandas (without polars and other libraries), is it possible to get information about the needed patent faster and less memory consuming than next method. Get publication_number in test.csv -> get publication_date in patent_metadata.parquet -> get info in patent_data/publication_date.parquet (Load full parquet)?\n**Thanks for any help, explanations, links**",
    "2878565": "Garbage collector in Python doesn't always fully release unused memory. There are ways to deal with it, e.g. \nlibc = ctypes.CDLL(\"libc.so.6\") and then libc.malloc_trim(0)\n\nYou may read the files by chunks so that they definitely fit into the memory. Even better yet, you may read only required columns, that will improve both memory usage and performance.",
    "2878601": "Thanks, I'll try to apply the memory tip.\n\nIf you mean reading the parquet line by line or chunk by chunk, it will take a long time to find the right line. I usually load the whole parquet, set 'publication_number' as an index and quickly find the required row by index. Doesn't seem to be a faster way?\n\nAs for columns, I want to work not only with title and abstract. As far as I know, loading only 'publication_number' column and making it an index, we can't load the rest of information (claims, description etc.) for the line we NEED. That's why we have to load the whole parquet.",
    "2878909": "Try using polars it will reduce the memory ,As I faced this issue at the beginning of the competition , And of course this competition needs a good knowledge of how to deal with memory",
    "2878917": "For example, there is a way to split publication_date.parquet into smaller subfiles and save them, and refer to the dictionary of patent->filename.\n\npublication_number -> dict_patent2file -> smaller_target_file"
  },
  "source": "meta"
}