{
  "id": 384359,
  "title": "Low memory Kaggle servers love this weird trick. Click ...",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/384359",
  "author_name": "DennisSakva",
  "post_date": "2023-02-07T16:28:30.178000",
  "votes": 101,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Clickbait at its finest :)<br>\nI see many Kagglers struggling with memory errors while doing pre-processing. Use this code to load competition data into memory with like 10x smaller footprint. Don't do column format conversion after loading data, do it while loading data. It works automagically. <br>\nThen do all your pre-processing with this much more memory-efficient dataframe. <br>\nCheers.</p>\n<pre><code>dtypes={:, \n:np.int32,\n    :,\n    :,\n    :np.uint8,\n    :,\n    :np.float32,\n    :np.float32,\n    :np.float32,\n    :np.float32,\n    :np.float32,\n     :,\n     :,\n     :,\n     :,\n     :,\n     :,\n     :,\n     :}\ntrain_df=pd.read_csv(work_path+, dtype=dtypes)\n</code></pre>",
  "messages": [
    {
      "id": 2133848,
      "postDate": "2023-02-07T16:28:30.180Z",
      "content": "<p>Clickbait at its finest :)<br>\nI see many Kagglers struggling with memory errors while doing pre-processing. Use this code to load competition data into memory with like 10x smaller footprint. Don't do column format conversion after loading data, do it while loading data. It works automagically. <br>\nThen do all your pre-processing with this much more memory-efficient dataframe. <br>\nCheers.</p>\n<pre><code>dtypes={:, \n:np.int32,\n    :,\n    :,\n    :np.uint8,\n    :,\n    :np.float32,\n    :np.float32,\n    :np.float32,\n    :np.float32,\n    :np.float32,\n     :,\n     :,\n     :,\n     :,\n     :,\n     :,\n     :,\n     :}\ntrain_df=pd.read_csv(work_path+, dtype=dtypes)\n</code></pre>",
      "rawMarkdown": "Clickbait at its finest :)\nI see many Kagglers struggling with memory errors while doing pre-processing. Use this code to load competition data into memory with like 10x smaller footprint. Don't do column format conversion after loading data, do it while loading data. It works automagically. \nThen do all your pre-processing with this much more memory-efficient dataframe. \nCheers.\n\n```python\ndtypes={'session_id':'category', \n'elapsed_time':np.int32,\n    'event_name':'category',\n    'name':'category',\n    'level':np.uint8,\n    'page':'category',\n    'room_coor_x':np.float32,\n    'room_coor_y':np.float32,\n    'screen_coor_x':np.float32,\n    'screen_coor_y':np.float32,\n    'hover_duration':np.float32,\n     'text':'category',\n     'fqid':'category',\n     'room_fqid':'category',\n     'text_fqid':'category',\n     'fullscreen':'category',\n     'hq':'category',\n     'music':'category',\n     'level_group':'category'}\ntrain_df=pd.read_csv(work_path+'train.csv', dtype=dtypes)\n```\n",
      "votes": 101
    },
    {
      "id": 2134524,
      "postDate": "2023-02-08T03:54:45.077Z",
      "content": "<p>Thank you for sharing your idea! Additionaly, ignoring <code>fullscreen</code>, <code>hq</code> and <code>music</code> is good way to save RAM usage because these column don't have any valid value (all values are missing).</p>",
      "rawMarkdown": "Thank you for sharing your idea! Additionaly, ignoring `fullscreen`, `hq` and `music` is good way to save RAM usage because these column don't have any valid value (all values are missing).",
      "votes": 8
    },
    {
      "id": 2265163,
      "postDate": "2023-05-19T03:17:00.970Z",
      "content": "<p>Why not np.float16 for the numerical columns? The range of the columns are all within the range of np.float16</p>",
      "rawMarkdown": "Why not np.float16 for the numerical columns? The range of the columns are all within the range of np.float16",
      "replies": [
        {
          "id": 2265171,
          "postDate": "2023-05-19T03:30:49.260Z",
          "content": "<p>Also, is there anyway to deal with the 'page' column? It's meant to be an integer column but it contains missing values so pandas treats it as a float column which cannot be downcasted to int columns. <br>\nMuch appreciated!</p>",
          "rawMarkdown": "Also, is there anyway to deal with the 'page' column? It's meant to be an integer column but it contains missing values so pandas treats it as a float column which cannot be downcasted to int columns. \nMuch appreciated!",
          "replies": [
            {
              "id": 2265737,
              "postDate": "2023-05-19T13:35:04.550Z",
              "content": "<p>Use category for page. It's a very natural choice.<br>\nAf for f16, you better try and compare results. Fp32 is usually quite robust compared to f64, yet fp16 calculations are more fragile. Sometimes they work, sometimes they don't. It's not about range, its about numeric precision.</p>",
              "rawMarkdown": "Use category for page. It's a very natural choice.\nAf for f16, you better try and compare results. Fp32 is usually quite robust compared to f64, yet fp16 calculations are more fragile. Sometimes they work, sometimes they don't. It's not about range, its about numeric precision.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2224730,
      "postDate": "2023-04-17T15:52:42.070Z",
      "content": "<p>As I'd just learnt from my own bone and bloody experience, this perfectly works. Especially when pandas changes data type itself, this help reserve the data type as desired with lesssss, much more lessss than post processing. But, why do we bother reading CSV, read the parquet instead. Parquet provides smaller size of file in most cases compares to csv. More than enough, parquet also save the data types of each column. Ah, for parquet, Kaggle doesn't provide data with parquet, so I think parquet is for practical work only</p>",
      "rawMarkdown": "As I'd just learnt from my own bone and bloody experience, this perfectly works. Especially when pandas changes data type itself, this help reserve the data type as desired with lesssss, much more lessss than post processing. But, why do we bother reading CSV, read the parquet instead. Parquet provides smaller size of file in most cases compares to csv. More than enough, parquet also save the data types of each column. Ah, for parquet, Kaggle doesn't provide data with parquet, so I think parquet is for practical work only"
    },
    {
      "id": 2220069,
      "postDate": "2023-04-13T05:25:52.973Z",
      "content": "<p>Thank you for sharing your idea!  I am new to python and would like to ask you a few questions. I always suffer from it. how did you learn  memory kaggle? <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a></p>",
      "rawMarkdown": "Thank you for sharing your idea!  I am new to python and would like to ask you a few questions. I always suffer from it. how did you learn  memory kaggle? @sakvaua"
    },
    {
      "id": 2135002,
      "postDate": "2023-02-08T11:48:14.780Z",
      "content": "<p>Yet another day where i get to learn something that is very useful , thanks a lot <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> </p>",
      "rawMarkdown": "Yet another day where i get to learn something that is very useful , thanks a lot @sakvaua "
    },
    {
      "id": 2134455,
      "postDate": "2023-02-08T01:55:21.303Z",
      "content": "<p>Haha - you got me Dennis. But really, I think this is a great approach. Nothing less than float32 and use of categories where it applies. One question on room coordinates, which are precise in the data. Do you think np.float32 is OK vs. 64? I haven't closely looked at the data to know either way.</p>",
      "rawMarkdown": "Haha - you got me Dennis. But really, I think this is a great approach. Nothing less than float32 and use of categories where it applies. One question on room coordinates, which are precise in the data. Do you think np.float32 is OK vs. 64? I haven't closely looked at the data to know either way.",
      "replies": [
        {
          "id": 2134664,
          "postDate": "2023-02-08T07:10:20.893Z",
          "content": "<p>If I learned one thing in machine learning it's that you won't know until you try :) But my experience tells me that the difference between float64 and float32 doesn't matter in most cases. It is negligible and in the order of 1e-5 for our scale of numbers. That's the difference of like 1 pixel on a 10,000x10,000 image.<br>\nEven more, I'm pretty sure that these values can (and probably should) be discretized even further, but that's something yet to be seen.<br>\nRegards.</p>",
          "rawMarkdown": "If I learned one thing in machine learning it's that you won't know until you try :) But my experience tells me that the difference between float64 and float32 doesn't matter in most cases. It is negligible and in the order of 1e-5 for our scale of numbers. That's the difference of like 1 pixel on a 10,000x10,000 image.\nEven more, I'm pretty sure that these values can (and probably should) be discretized even further, but that's something yet to be seen.\nRegards.",
          "votes": 6,
          "replies": [
            {
              "id": 2134689,
              "postDate": "2023-02-08T07:38:12.223Z",
              "content": "<p>As for the precision of room coordinates in the csv file, I think It has more to do with the representation of floats in binary and decimal formats. Like 1/10 converted to float64bit and back to decimal representation equals to 0.1000000000000000055511151231257827021181583404541015625, but that doesn't mean that we need this level of precision for our purposes.</p>",
              "rawMarkdown": "As for the precision of room coordinates in the csv file, I think It has more to do with the representation of floats in binary and decimal formats. Like 1/10 converted to float64bit and back to decimal representation equals to 0.1000000000000000055511151231257827021181583404541015625, but that doesn't mean that we need this level of precision for our purposes.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2190541,
      "postDate": "2023-03-21T10:28:20.617Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2133863,
      "postDate": "2023-02-07T16:38:51.897Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2235977,
      "postDate": "2023-04-26T13:06:17.897Z",
      "content": "<p>Thank You very much!</p>",
      "rawMarkdown": "Thank You very much!",
      "votes": 1
    },
    {
      "id": 2289164,
      "postDate": "2023-06-06T01:11:21.693Z",
      "content": "<p>Thank you for sharing your idea! </p>",
      "rawMarkdown": "Thank you for sharing your idea! "
    },
    {
      "id": 2231320,
      "postDate": "2023-04-23T08:23:34.203Z",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!\n"
    },
    {
      "id": 2218151,
      "postDate": "2023-04-11T13:16:28.253Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing."
    },
    {
      "id": 2206270,
      "postDate": "2023-04-02T12:51:47.193Z",
      "content": "<p>Thank you for your valuable comment !</p>",
      "rawMarkdown": "Thank you for your valuable comment !"
    },
    {
      "id": 2136277,
      "postDate": "2023-02-09T08:02:35.720Z",
      "content": "<p>Thank you for sharing.<br>\nSo useful!</p>",
      "rawMarkdown": "Thank you for sharing.\nSo useful!\n"
    }
  ],
  "comments": [
    {
      "id": 2134524,
      "author_name": "Quvotha",
      "author_url": "",
      "post_date": "2023-02-08T03:54:45.077000",
      "content": "<p>Thank you for sharing your idea! Additionaly, ignoring <code>fullscreen</code>, <code>hq</code> and <code>music</code> is good way to save RAM usage because these column don't have any valid value (all values are missing).</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 2265163,
      "author_name": "Guolun Li",
      "author_url": "",
      "post_date": "2023-05-19T03:17:00.970000",
      "content": "<p>Why not np.float16 for the numerical columns? The range of the columns are all within the range of np.float16</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2265171,
          "author_name": "Guolun Li",
          "author_url": "",
          "post_date": "2023-05-19T03:30:49.260000",
          "content": "<p>Also, is there anyway to deal with the 'page' column? It's meant to be an integer column but it contains missing values so pandas treats it as a float column which cannot be downcasted to int columns. <br>\nMuch appreciated!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2265737,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2023-05-19T13:35:04.550000",
              "content": "<p>Use category for page. It's a very natural choice.<br>\nAf for f16, you better try and compare results. Fp32 is usually quite robust compared to f64, yet fp16 calculations are more fragile. Sometimes they work, sometimes they don't. It's not about range, its about numeric precision.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2224730,
      "author_name": "Thái Lê",
      "author_url": "",
      "post_date": "2023-04-17T15:52:42.070000",
      "content": "<p>As I'd just learnt from my own bone and bloody experience, this perfectly works. Especially when pandas changes data type itself, this help reserve the data type as desired with lesssss, much more lessss than post processing. But, why do we bother reading CSV, read the parquet instead. Parquet provides smaller size of file in most cases compares to csv. More than enough, parquet also save the data types of each column. Ah, for parquet, Kaggle doesn't provide data with parquet, so I think parquet is for practical work only</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2220069,
      "author_name": "riicon",
      "author_url": "",
      "post_date": "2023-04-13T05:25:52.973000",
      "content": "<p>Thank you for sharing your idea!  I am new to python and would like to ask you a few questions. I always suffer from it. how did you learn  memory kaggle? <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2135002,
      "author_name": "Ashvanth s",
      "author_url": "",
      "post_date": "2023-02-08T11:48:14.780000",
      "content": "<p>Yet another day where i get to learn something that is very useful , thanks a lot <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2134455,
      "author_name": "JohnM",
      "author_url": "",
      "post_date": "2023-02-08T01:55:21.303000",
      "content": "<p>Haha - you got me Dennis. But really, I think this is a great approach. Nothing less than float32 and use of categories where it applies. One question on room coordinates, which are precise in the data. Do you think np.float32 is OK vs. 64? I haven't closely looked at the data to know either way.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2134664,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2023-02-08T07:10:20.893000",
          "content": "<p>If I learned one thing in machine learning it's that you won't know until you try :) But my experience tells me that the difference between float64 and float32 doesn't matter in most cases. It is negligible and in the order of 1e-5 for our scale of numbers. That's the difference of like 1 pixel on a 10,000x10,000 image.<br>\nEven more, I'm pretty sure that these values can (and probably should) be discretized even further, but that's something yet to be seen.<br>\nRegards.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2134689,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2023-02-08T07:38:12.223000",
              "content": "<p>As for the precision of room coordinates in the csv file, I think It has more to do with the representation of floats in binary and decimal formats. Like 1/10 converted to float64bit and back to decimal representation equals to 0.1000000000000000055511151231257827021181583404541015625, but that doesn't mean that we need this level of precision for our purposes.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2190541,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-21T10:28:20.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2133863,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-07T16:38:51.897000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2235977,
      "author_name": "Andrey Kulikov",
      "author_url": "",
      "post_date": "2023-04-26T13:06:17.897000",
      "content": "<p>Thank You very much!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2289164,
      "author_name": "liuyanghe",
      "author_url": "",
      "post_date": "2023-06-06T01:11:21.693000",
      "content": "<p>Thank you for sharing your idea! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2231320,
      "author_name": "Natalie Kalina",
      "author_url": "",
      "post_date": "2023-04-23T08:23:34.203000",
      "content": "<p>Thank you very much!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2218151,
      "author_name": "Byungeun Hwang",
      "author_url": "",
      "post_date": "2023-04-11T13:16:28.253000",
      "content": "<p>Thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2206270,
      "author_name": "Shoma Tateno",
      "author_url": "",
      "post_date": "2023-04-02T12:51:47.193000",
      "content": "<p>Thank you for your valuable comment !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2136277,
      "author_name": "shun takinami",
      "author_url": "",
      "post_date": "2023-02-09T08:02:35.720000",
      "content": "<p>Thank you for sharing.<br>\nSo useful!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2133848": "Clickbait at its finest :)\nI see many Kagglers struggling with memory errors while doing pre-processing. Use this code to load competition data into memory with like 10x smaller footprint. Don't do column format conversion after loading data, do it while loading data. It works automagically. \nThen do all your pre-processing with this much more memory-efficient dataframe. \nCheers.\n\n```python\ndtypes={'session_id':'category', \n'elapsed_time':np.int32,\n    'event_name':'category',\n    'name':'category',\n    'level':np.uint8,\n    'page':'category',\n    'room_coor_x':np.float32,\n    'room_coor_y':np.float32,\n    'screen_coor_x':np.float32,\n    'screen_coor_y':np.float32,\n    'hover_duration':np.float32,\n     'text':'category',\n     'fqid':'category',\n     'room_fqid':'category',\n     'text_fqid':'category',\n     'fullscreen':'category',\n     'hq':'category',\n     'music':'category',\n     'level_group':'category'}\ntrain_df=pd.read_csv(work_path+'train.csv', dtype=dtypes)\n```\n",
    "2134524": "Thank you for sharing your idea! Additionaly, ignoring `fullscreen`, `hq` and `music` is good way to save RAM usage because these column don't have any valid value (all values are missing).",
    "2265163": "Why not np.float16 for the numerical columns? The range of the columns are all within the range of np.float16",
    "2224730": "As I'd just learnt from my own bone and bloody experience, this perfectly works. Especially when pandas changes data type itself, this help reserve the data type as desired with lesssss, much more lessss than post processing. But, why do we bother reading CSV, read the parquet instead. Parquet provides smaller size of file in most cases compares to csv. More than enough, parquet also save the data types of each column. Ah, for parquet, Kaggle doesn't provide data with parquet, so I think parquet is for practical work only",
    "2220069": "Thank you for sharing your idea!  I am new to python and would like to ask you a few questions. I always suffer from it. how did you learn  memory kaggle? @sakvaua",
    "2135002": "Yet another day where i get to learn something that is very useful , thanks a lot @sakvaua ",
    "2134455": "Haha - you got me Dennis. But really, I think this is a great approach. Nothing less than float32 and use of categories where it applies. One question on room coordinates, which are precise in the data. Do you think np.float32 is OK vs. 64? I haven't closely looked at the data to know either way.",
    "2190541": "",
    "2133863": "",
    "2235977": "Thank You very much!",
    "2289164": "Thank you for sharing your idea! ",
    "2231320": "Thank you very much!\n",
    "2218151": "Thank you for sharing.",
    "2206270": "Thank you for your valuable comment !",
    "2136277": "Thank you for sharing.\nSo useful!\n"
  }
}