{
  "id": 328245,
  "title": "Is there a shorter version?",
  "url": "/competitions/amex-default-prediction/discussion/328245",
  "author_name": "",
  "post_date": "2022-05-31T15:34:56.819238700Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have 16GB RAM but the training set is still way too large to open.<br>\nI suppose even a few hundred entries would give some very good solutions, and this is all that matters from the host's point of view.<br>\nBut you say that you also seek to recruit skilful individuals, and by making the size of the dataset prohibitive you're missing many such individuals.</p>",
  "messages": [
    {
      "id": "1806937",
      "postDate": "05/31/2022 15:34:56",
      "content": "<p>I have 16GB RAM but the training set is still way too large to open.<br>\nI suppose even a few hundred entries would give some very good solutions, and this is all that matters from the host's point of view.<br>\nBut you say that you also seek to recruit skilful individuals, and by making the size of the dataset prohibitive you're missing many such individuals.</p>",
      "rawMarkdown": "I have 16GB RAM but the training set is still way too large to open.\nI suppose even a few hundred entries would give some very good solutions, and this is all that matters from the host's point of view.\nBut you say that you also seek to recruit skilful individuals, and by making the size of the dataset prohibitive you're missing many such individuals.",
      "votes": null
    },
    {
      "id": "1807005",
      "postDate": "05/31/2022 16:32:46",
      "content": "<p>I understand this frustration, I was doing the Home Credit competition with 16GB RAM at first as well. Here is what I did back then in the Home Credit competition, hopefully, this can help you.<br>\nMost of the time the memory usage peak is at joining, assigning columns, or aggregation. Most of the time, you only need to deal with those memory peaks.</p>\n<ol>\n<li>Increase the swap space in my Windows machine so that all the peak memory usage landed on my hard disk instead of raising an error.ca</li>\n<li>Keep deleting unused data frames and calling gc.collect.</li>\n<li>You can use fp16 data, but just before your feature engineering, covert the column to fp32, and after the feature, engineering converts it back to fp16. (I shared this in the other post as well).</li>\n<li>Save different features by files, and do your feature selection on a feature file-level instead of all features. so in the end, before combining all feature files, you can control/select how many features you want to use for training that can fit into your RAM + Swap.<br>\nHope it is helpful. In the end, what I did was buy a second-handed motherboard that can host 64GB of ram, you will be surprised how many IT companies or university research labs are selling old machine parts.~~</li>\n</ol>",
      "rawMarkdown": "I understand this frustration, I was doing the Home Credit competition with 16GB RAM at first as well. Here is what I did back then in the Home Credit competition, hopefully, this can help you.\n\nMost of the time the memory usage peak is at joining, assigning columns, or aggregation. Most of the time, you only need to deal with those memory peaks.\n\n1. Increase the swap space in my Windows machine so that all the peak memory usage landed on my hard disk instead of raising an error.ca\n\n2. Keep deleting unused data frames and calling gc.collect.\n\n3. You can use fp16 data, but just before your feature engineering, covert the column to fp32, and after the feature, engineering converts it back to fp16. (I shared this in the other post as well).\n\n4. Save different features by files, and do your feature selection on a feature file-level instead of all features. so in the end, before combining all feature files, you can control/select how many features you want to use for training that can fit into your RAM + Swap.\n\n\nHope it is helpful. In the end, what I did was buy a second-handed motherboard that can host 64GB of ram, you will be surprised how many IT companies or university research labs are selling old machine parts.~~",
      "votes": null
    },
    {
      "id": "1807166",
      "postDate": "05/31/2022 19:19:10",
      "content": "<p>Yes there is, <a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Yes there is, [here](https://www.kaggle.com/datasets/munumbutt/amexfeather)",
      "votes": null
    },
    {
      "id": "1807274",
      "postDate": "05/31/2022 21:59:07",
      "content": "<p>Put a very large swap on an SSD - works great on Ubuntu - I am using 200GB - slow but works.  Not sure how well large virtual memory on Windows will work since I left Windows AI a couple of years ago (Windows grabs to much GPU memory and will not let go).</p>",
      "rawMarkdown": "Put a very large swap on an SSD - works great on Ubuntu - I am using 200GB - slow but works.  Not sure how well large virtual memory on Windows will work since I left Windows AI a couple of years ago (Windows grabs to much GPU memory and will not let go).",
      "votes": null
    },
    {
      "id": "1807276",
      "postDate": "05/31/2022 22:00:23",
      "content": "<pre><code># create larger swap file on system\n# Resize Swap to 200GB\n# when I try to do these steps by putting sudo in front of each it often failed\n# but once I set sudo as show here than it worked\nsudo -s\n# Turn swap off\n# This moves stuff in swap to the main memory and might take several minutes\nswapoff -a\n# Create an empty swapfile\n# Note that \"1G\" is basically just the unit and count is an integer.\n# Together, they define the size. In this case 200GB.\n# this step takes a long time - 5-15 minutes\ndd if=/dev/zero of=/swapfile bs=1G count=200\nmkswap /swapfile  # Set up a Linux swap area\nswapon /swapfile  # Turn the swap on\n# Check if it worked\ngrep Swap /proc/meminfo\n</code></pre>",
      "rawMarkdown": "```\n# create larger swap file on system\n# Resize Swap to 200GB\n\n# when I try to do these steps by putting sudo in front of each it often failed\n# but once I set sudo as show here than it worked\n\nsudo -s\n\n# Turn swap off\n# This moves stuff in swap to the main memory and might take several minutes\nswapoff -a\n\n# Create an empty swapfile\n# Note that \"1G\" is basically just the unit and count is an integer.\n# Together, they define the size. In this case 200GB.\n\n# this step takes a long time - 5-15 minutes\ndd if=/dev/zero of=/swapfile bs=1G count=200\n\nmkswap /swapfile  # Set up a Linux swap area\nswapon /swapfile  # Turn the swap on\n\n# Check if it worked\n\ngrep Swap /proc/meminfo\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1807005,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "05/31/2022 16:32:46",
      "content": "<p>I understand this frustration, I was doing the Home Credit competition with 16GB RAM at first as well. Here is what I did back then in the Home Credit competition, hopefully, this can help you.<br>\nMost of the time the memory usage peak is at joining, assigning columns, or aggregation. Most of the time, you only need to deal with those memory peaks.</p>\n<ol>\n<li>Increase the swap space in my Windows machine so that all the peak memory usage landed on my hard disk instead of raising an error.ca</li>\n<li>Keep deleting unused data frames and calling gc.collect.</li>\n<li>You can use fp16 data, but just before your feature engineering, covert the column to fp32, and after the feature, engineering converts it back to fp16. (I shared this in the other post as well).</li>\n<li>Save different features by files, and do your feature selection on a feature file-level instead of all features. so in the end, before combining all feature files, you can control/select how many features you want to use for training that can fit into your RAM + Swap.<br>\nHope it is helpful. In the end, what I did was buy a second-handed motherboard that can host 64GB of ram, you will be surprised how many IT companies or university research labs are selling old machine parts.~~</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1807166,
      "author_name": "munumbutt",
      "author_url": "",
      "post_date": "05/31/2022 19:19:10",
      "content": "<p>Yes there is, <a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1807274,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/31/2022 21:59:07",
      "content": "<p>Put a very large swap on an SSD - works great on Ubuntu - I am using 200GB - slow but works.  Not sure how well large virtual memory on Windows will work since I left Windows AI a couple of years ago (Windows grabs to much GPU memory and will not let go).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1807276,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "05/31/2022 22:00:23",
          "content": "<pre><code># create larger swap file on system\n# Resize Swap to 200GB\n# when I try to do these steps by putting sudo in front of each it often failed\n# but once I set sudo as show here than it worked\nsudo -s\n# Turn swap off\n# This moves stuff in swap to the main memory and might take several minutes\nswapoff -a\n# Create an empty swapfile\n# Note that \"1G\" is basically just the unit and count is an integer.\n# Together, they define the size. In this case 200GB.\n# this step takes a long time - 5-15 minutes\ndd if=/dev/zero of=/swapfile bs=1G count=200\nmkswap /swapfile  # Set up a Linux swap area\nswapon /swapfile  # Turn the swap on\n# Check if it worked\ngrep Swap /proc/meminfo\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1806937": "I have 16GB RAM but the training set is still way too large to open.\nI suppose even a few hundred entries would give some very good solutions, and this is all that matters from the host's point of view.\nBut you say that you also seek to recruit skilful individuals, and by making the size of the dataset prohibitive you're missing many such individuals.",
    "1807005": "I understand this frustration, I was doing the Home Credit competition with 16GB RAM at first as well. Here is what I did back then in the Home Credit competition, hopefully, this can help you.\n\nMost of the time the memory usage peak is at joining, assigning columns, or aggregation. Most of the time, you only need to deal with those memory peaks.\n\n1. Increase the swap space in my Windows machine so that all the peak memory usage landed on my hard disk instead of raising an error.ca\n\n2. Keep deleting unused data frames and calling gc.collect.\n\n3. You can use fp16 data, but just before your feature engineering, covert the column to fp32, and after the feature, engineering converts it back to fp16. (I shared this in the other post as well).\n\n4. Save different features by files, and do your feature selection on a feature file-level instead of all features. so in the end, before combining all feature files, you can control/select how many features you want to use for training that can fit into your RAM + Swap.\n\n\nHope it is helpful. In the end, what I did was buy a second-handed motherboard that can host 64GB of ram, you will be surprised how many IT companies or university research labs are selling old machine parts.~~",
    "1807166": "Yes there is, [here](https://www.kaggle.com/datasets/munumbutt/amexfeather)",
    "1807274": "Put a very large swap on an SSD - works great on Ubuntu - I am using 200GB - slow but works.  Not sure how well large virtual memory on Windows will work since I left Windows AI a couple of years ago (Windows grabs to much GPU memory and will not let go).",
    "1807276": "```\n# create larger swap file on system\n# Resize Swap to 200GB\n\n# when I try to do these steps by putting sudo in front of each it often failed\n# but once I set sudo as show here than it worked\n\nsudo -s\n\n# Turn swap off\n# This moves stuff in swap to the main memory and might take several minutes\nswapoff -a\n\n# Create an empty swapfile\n# Note that \"1G\" is basically just the unit and count is an integer.\n# Together, they define the size. In this case 200GB.\n\n# this step takes a long time - 5-15 minutes\ndd if=/dev/zero of=/swapfile bs=1G count=200\n\nmkswap /swapfile  # Set up a Linux swap area\nswapon /swapfile  # Turn the swap on\n\n# Check if it worked\n\ngrep Swap /proc/meminfo\n```"
  },
  "source": "meta"
}