{
  "id": 206924,
  "title": "Distortion of (float)data when reducing memory usage?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206924",
  "author_name": "jwc",
  "post_date": "2020-12-27T08:51:47.960000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In many of Kaggle competitions, <a href=\"https://www.kaggle.com/rinnqd/reduce-memory-usage\" target=\"_blank\">this code sample</a> is used to reduce the memory size of <code>DataFrame</code>.</p>\n<p>I found out strange thing when converting <code>float</code> data: There is a conditional statement for converting <code>float</code> data:  <code>c_min &gt; np.finfo(np.float32).min and c_max &lt; np.finfo(np.float32).max</code>. If you check <code>np.finfo(np.float32).max</code> out manually, it is equal to <code>3.4028235e+38</code>,  <strong>which represents a really huge value</strong>. I think that almost all float values we are faced with in the kaggle competition is less than <code>3.4028235e+38</code>, which means any <code>float</code> data has no chance to be converted to <code>float64</code> dtype when using this code.</p>\n<p>What do you think?</p>",
  "messages": [
    {
      "id": 1128198,
      "postDate": "2020-12-27T08:51:47.960Z",
      "content": "<p>In many of Kaggle competitions, <a href=\"https://www.kaggle.com/rinnqd/reduce-memory-usage\" target=\"_blank\">this code sample</a> is used to reduce the memory size of <code>DataFrame</code>.</p>\n<p>I found out strange thing when converting <code>float</code> data: There is a conditional statement for converting <code>float</code> data:  <code>c_min &gt; np.finfo(np.float32).min and c_max &lt; np.finfo(np.float32).max</code>. If you check <code>np.finfo(np.float32).max</code> out manually, it is equal to <code>3.4028235e+38</code>,  <strong>which represents a really huge value</strong>. I think that almost all float values we are faced with in the kaggle competition is less than <code>3.4028235e+38</code>, which means any <code>float</code> data has no chance to be converted to <code>float64</code> dtype when using this code.</p>\n<p>What do you think?</p>",
      "rawMarkdown": "In many of Kaggle competitions, [this code sample](https://www.kaggle.com/rinnqd/reduce-memory-usage) is used to reduce the memory size of `DataFrame`.\n\nI found out strange thing when converting `float` data: There is a conditional statement for converting `float` data:  `c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max`. If you check `np.finfo(np.float32).max` out manually, it is equal to `3.4028235e+38`,  **which represents a really huge value**. I think that almost all float values we are faced with in the kaggle competition is less than `3.4028235e+38`, which means any `float` data has no chance to be converted to `float64` dtype when using this code.\n\nWhat do you think?\n\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1128198": "In many of Kaggle competitions, [this code sample](https://www.kaggle.com/rinnqd/reduce-memory-usage) is used to reduce the memory size of `DataFrame`.\n\nI found out strange thing when converting `float` data: There is a conditional statement for converting `float` data:  `c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max`. If you check `np.finfo(np.float32).max` out manually, it is equal to `3.4028235e+38`,  **which represents a really huge value**. I think that almost all float values we are faced with in the kaggle competition is less than `3.4028235e+38`, which means any `float` data has no chance to be converted to `float64` dtype when using this code.\n\nWhat do you think?\n\n"
  }
}