{
  "id": 384475,
  "title": "Memory reduction  - watch out for pitfalls!",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/384475",
  "author_name": "",
  "post_date": "2023-02-08T02:23:32.130742500Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The method you use for memory reduction can impact your results. Float values in particular carry a certain amount of precision which may or may not be needed. Converting all floats to float16, or even float32 in some cases, removes information that could help your model.</p>\n<p>There's a popular function on Kaggle, often used to reduce the memory footprint for dataframes. With all due respect to the author, I would only use it with great caution, mainly because it changes all floats to 16 bits.  In this competition, room coordinates are shown to be very precise. Is the precision needed? I don't know, but I would still stay with float32 at the least until investigating.<br>\nDiscussions on SO here:</p>\n<p><a href=\"https://stackoverflow.com/questions/32465481/what-exactly-is-the-resolution-parameter-of-numpy-float\" target=\"_blank\">https://stackoverflow.com/questions/32465481/what-exactly-is-the-resolution-parameter-of-numpy-float</a><br>\nand  <a href=\"https://stackoverflow.com/questions/43440821/the-real-difference-between-float32-and-float64A\" target=\"_blank\">https://stackoverflow.com/questions/43440821/the-real-difference-between-float32-and-float64A</a>.</p>\n<p>So what to do? I think the easiest way is to explicitly name datatypes for each column. <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> has it waiting for you at <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359</a> . </p>\n<p>You can also use the <code>skmem</code> utility script here on Kaggle, which is useful for wide dataframes where listing each column dtype could be tedious.</p>\n<p>Maybe someone has findings to share for this dataset? It could be that precision isn't needed at all with these magnitudes and ranges.</p>",
  "messages": [
    {
      "id": "2134477",
      "postDate": "02/08/2023 02:23:32",
      "content": "<p>The method you use for memory reduction can impact your results. Float values in particular carry a certain amount of precision which may or may not be needed. Converting all floats to float16, or even float32 in some cases, removes information that could help your model.</p>\n<p>There's a popular function on Kaggle, often used to reduce the memory footprint for dataframes. With all due respect to the author, I would only use it with great caution, mainly because it changes all floats to 16 bits.  In this competition, room coordinates are shown to be very precise. Is the precision needed? I don't know, but I would still stay with float32 at the least until investigating.<br>\nDiscussions on SO here:</p>\n<p><a href=\"https://stackoverflow.com/questions/32465481/what-exactly-is-the-resolution-parameter-of-numpy-float\" target=\"_blank\">https://stackoverflow.com/questions/32465481/what-exactly-is-the-resolution-parameter-of-numpy-float</a><br>\nand  <a href=\"https://stackoverflow.com/questions/43440821/the-real-difference-between-float32-and-float64A\" target=\"_blank\">https://stackoverflow.com/questions/43440821/the-real-difference-between-float32-and-float64A</a>.</p>\n<p>So what to do? I think the easiest way is to explicitly name datatypes for each column. <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> has it waiting for you at <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359</a> . </p>\n<p>You can also use the <code>skmem</code> utility script here on Kaggle, which is useful for wide dataframes where listing each column dtype could be tedious.</p>\n<p>Maybe someone has findings to share for this dataset? It could be that precision isn't needed at all with these magnitudes and ranges.</p>",
      "rawMarkdown": "The method you use for memory reduction can impact your results. Float values in particular carry a certain amount of precision which may or may not be needed. Converting all floats to float16, or even float32 in some cases, removes information that could help your model.\n\nThere's a popular function on Kaggle, often used to reduce the memory footprint for dataframes. With all due respect to the author, I would only use it with great caution, mainly because it changes all floats to 16 bits.  In this competition, room coordinates are shown to be very precise. Is the precision needed? I don't know, but I would still stay with float32 at the least until investigating.\nDiscussions on SO here:\n\nhttps://stackoverflow.com/questions/32465481/what-exactly-is-the-resolution-parameter-of-numpy-float\nand  https://stackoverflow.com/questions/43440821/the-real-difference-between-float32-and-float64A.\n\n\nSo what to do? I think the easiest way is to explicitly name datatypes for each column. @sakvaua has it waiting for you at https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359 . \n\nYou can also use the `skmem` utility script here on Kaggle, which is useful for wide dataframes where listing each column dtype could be tedious.\n\nMaybe someone has findings to share for this dataset? It could be that precision isn't needed at all with these magnitudes and ranges.",
      "votes": null
    },
    {
      "id": "2134688",
      "postDate": "02/08/2023 07:38:02",
      "content": "<p>This is the reason I love Kaggle community. <br>\nEveryone have an equal opportunity to express their findings with solid proofs.</p>\n<p>I was intrigued by the memory reduction technique. However I will now implement with a pinch of salt.</p>\n<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> Thank you for pointing this out.</p>",
      "rawMarkdown": "This is the reason I love Kaggle community. \nEveryone have an equal opportunity to express their findings with solid proofs.\n\nI was intrigued by the memory reduction technique. However I will now implement with a pinch of salt.\n\n@jpmiller Thank you for pointing this out.",
      "votes": null
    },
    {
      "id": "2135353",
      "postDate": "02/08/2023 15:25:01",
      "content": "<p>This is great and clearly explained with reasonable explanation. Thanks <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> </p>",
      "rawMarkdown": "This is great and clearly explained with reasonable explanation. Thanks @jpmiller",
      "votes": null
    },
    {
      "id": "2135466",
      "postDate": "02/08/2023 16:48:23",
      "content": "<p>Thank you, Suraj. Have you found that the dytpe matters at all for model accuracy in this competition? I suspect it may be an academic discussion when it comes to this particular dataset.</p>",
      "rawMarkdown": "Thank you, Suraj. Have you found that the dytpe matters at all for model accuracy in this competition? I suspect it may be an academic discussion when it comes to this particular dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2134688,
      "author_name": "surajdengale",
      "author_url": "",
      "post_date": "02/08/2023 07:38:02",
      "content": "<p>This is the reason I love Kaggle community. <br>\nEveryone have an equal opportunity to express their findings with solid proofs.</p>\n<p>I was intrigued by the memory reduction technique. However I will now implement with a pinch of salt.</p>\n<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> Thank you for pointing this out.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2135466,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "02/08/2023 16:48:23",
          "content": "<p>Thank you, Suraj. Have you found that the dytpe matters at all for model accuracy in this competition? I suspect it may be an academic discussion when it comes to this particular dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2135353,
      "author_name": "zahrizhalali",
      "author_url": "",
      "post_date": "02/08/2023 15:25:01",
      "content": "<p>This is great and clearly explained with reasonable explanation. Thanks <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2134477": "The method you use for memory reduction can impact your results. Float values in particular carry a certain amount of precision which may or may not be needed. Converting all floats to float16, or even float32 in some cases, removes information that could help your model.\n\nThere's a popular function on Kaggle, often used to reduce the memory footprint for dataframes. With all due respect to the author, I would only use it with great caution, mainly because it changes all floats to 16 bits.  In this competition, room coordinates are shown to be very precise. Is the precision needed? I don't know, but I would still stay with float32 at the least until investigating.\nDiscussions on SO here:\n\nhttps://stackoverflow.com/questions/32465481/what-exactly-is-the-resolution-parameter-of-numpy-float\nand  https://stackoverflow.com/questions/43440821/the-real-difference-between-float32-and-float64A.\n\n\nSo what to do? I think the easiest way is to explicitly name datatypes for each column. @sakvaua has it waiting for you at https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359 . \n\nYou can also use the `skmem` utility script here on Kaggle, which is useful for wide dataframes where listing each column dtype could be tedious.\n\nMaybe someone has findings to share for this dataset? It could be that precision isn't needed at all with these magnitudes and ranges.",
    "2134688": "This is the reason I love Kaggle community. \nEveryone have an equal opportunity to express their findings with solid proofs.\n\nI was intrigued by the memory reduction technique. However I will now implement with a pinch of salt.\n\n@jpmiller Thank you for pointing this out.",
    "2135353": "This is great and clearly explained with reasonable explanation. Thanks @jpmiller",
    "2135466": "Thank you, Suraj. Have you found that the dytpe matters at all for model accuracy in this competition? I suspect it may be an academic discussion when it comes to this particular dataset."
  },
  "source": "meta"
}