{
  "id": 328138,
  "title": "Risk of overflow with Float 16 conversion",
  "url": "/competitions/amex-default-prediction/discussion/328138",
  "author_name": "",
  "post_date": "2022-05-31T06:15:02.354977900Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have seen several posts regarding data compression by converting numeric columns to float 16, doing this causes overflow errors on summary functions for e.g<br>\n<code>\nnp.mean(df.loc[df.R_9.notnull(),'R_9'].astype(np.float16))\n</code><br>\nReturns NaN due to overflow, for details check <a href=\"https://github.com/pandas-dev/pandas/issues/20642\" target=\"_blank\">this</a> </p>\n<p>This might effect model training as well. Float32 is well supported so its a safer conversion</p>",
  "messages": [
    {
      "id": "1806409",
      "postDate": "05/31/2022 06:15:02",
      "content": "<p>I have seen several posts regarding data compression by converting numeric columns to float 16, doing this causes overflow errors on summary functions for e.g<br>\n<code>\nnp.mean(df.loc[df.R_9.notnull(),'R_9'].astype(np.float16))\n</code><br>\nReturns NaN due to overflow, for details check <a href=\"https://github.com/pandas-dev/pandas/issues/20642\" target=\"_blank\">this</a> </p>\n<p>This might effect model training as well. Float32 is well supported so its a safer conversion</p>",
      "rawMarkdown": "I have seen several posts regarding data compression by converting numeric columns to float 16, doing this causes overflow errors on summary functions for e.g\n`\nnp.mean(df.loc[df.R_9.notnull(),'R_9'].astype(np.float16))\n`\nReturns NaN due to overflow, for details check [this](https://github.com/pandas-dev/pandas/issues/20642) \n\nThis might effect model training as well. Float32 is well supported so its a safer conversion",
      "votes": null
    },
    {
      "id": "1806864",
      "postDate": "05/31/2022 14:06:18",
      "content": "<p>It is definitely an important point, thank you. I never noticed that before.</p>",
      "rawMarkdown": "It is definitely an important point, thank you. I never noticed that before.",
      "votes": null
    },
    {
      "id": "1806880",
      "postDate": "05/31/2022 14:26:26",
      "content": "<p>Yes, conversion to <code>float16</code> causes overflow for some functions. I changed the data type to <code>float32</code> due to this issue.</p>",
      "rawMarkdown": "Yes, conversion to `float16` causes overflow for some functions. I changed the data type to `float32` due to this issue.",
      "votes": null
    },
    {
      "id": "1806889",
      "postDate": "05/31/2022 14:38:48",
      "content": "<p>You can convert your feature to float32 (1 column) before the computation. But this way you cannot do the multi-column aggregation, because casting all numerical columns to float32 will cause OOM on a Kaggle notebook.</p>\n<p>For example:</p>\n<pre><code>temp_df = df[[\"customer_ID\",  feature]]\ntemp_df[feature] = temp_df[feature].astype(np.float32)\ntemp_df.groupby(\"customer_ID\")[feature].mean().astype(np.float16)\n</code></pre>\n<p>But this is too troublesome, I switched to using float32 completely with my local machine instead.</p>\n<p>Haha, this reminds me of the famous interview question of computing the mid point during binary search<br>\n<code>mid = low + ((high - low) / 2);</code><br>\nvs<br>\n<code>mid = (low + high) / 2;</code></p>\n<p>The second middle point calculation has a chance of overflow if both low and high are close to the max of its variable type limit.</p>",
      "rawMarkdown": "You can convert your feature to float32 (1 column) before the computation. But this way you cannot do the multi-column aggregation, because casting all numerical columns to float32 will cause OOM on a Kaggle notebook.\n\nFor example:\n```\ntemp_df = df[[\"customer_ID\",  feature]]\ntemp_df[feature] = temp_df[feature].astype(np.float32)\ntemp_df.groupby(\"customer_ID\")[feature].mean().astype(np.float16)\n```\n\nBut this is too troublesome, I switched to using float32 completely with my local machine instead.\n\n\nHaha, this reminds me of the famous interview question of computing the mid point during binary search\n`mid = low + ((high - low) / 2);`\nvs\n`mid = (low + high) / 2;`\n\nThe second middle point calculation has a chance of overflow if both low and high are close to the max of its variable type limit.",
      "votes": null
    },
    {
      "id": "1812473",
      "postDate": "06/05/2022 21:05:11",
      "content": "<p>What a good point, I was just experiencing that, thank you very much for sharing.</p>",
      "rawMarkdown": "What a good point, I was just experiencing that, thank you very much for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1806864,
      "author_name": "meli19",
      "author_url": "",
      "post_date": "05/31/2022 14:06:18",
      "content": "<p>It is definitely an important point, thank you. I never noticed that before.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1806880,
      "author_name": "itacdonev",
      "author_url": "",
      "post_date": "05/31/2022 14:26:26",
      "content": "<p>Yes, conversion to <code>float16</code> causes overflow for some functions. I changed the data type to <code>float32</code> due to this issue.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1806889,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "05/31/2022 14:38:48",
      "content": "<p>You can convert your feature to float32 (1 column) before the computation. But this way you cannot do the multi-column aggregation, because casting all numerical columns to float32 will cause OOM on a Kaggle notebook.</p>\n<p>For example:</p>\n<pre><code>temp_df = df[[\"customer_ID\",  feature]]\ntemp_df[feature] = temp_df[feature].astype(np.float32)\ntemp_df.groupby(\"customer_ID\")[feature].mean().astype(np.float16)\n</code></pre>\n<p>But this is too troublesome, I switched to using float32 completely with my local machine instead.</p>\n<p>Haha, this reminds me of the famous interview question of computing the mid point during binary search<br>\n<code>mid = low + ((high - low) / 2);</code><br>\nvs<br>\n<code>mid = (low + high) / 2;</code></p>\n<p>The second middle point calculation has a chance of overflow if both low and high are close to the max of its variable type limit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1812473,
      "author_name": "maxdiazbattan",
      "author_url": "",
      "post_date": "06/05/2022 21:05:11",
      "content": "<p>What a good point, I was just experiencing that, thank you very much for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1806409": "I have seen several posts regarding data compression by converting numeric columns to float 16, doing this causes overflow errors on summary functions for e.g\n`\nnp.mean(df.loc[df.R_9.notnull(),'R_9'].astype(np.float16))\n`\nReturns NaN due to overflow, for details check [this](https://github.com/pandas-dev/pandas/issues/20642) \n\nThis might effect model training as well. Float32 is well supported so its a safer conversion",
    "1806864": "It is definitely an important point, thank you. I never noticed that before.",
    "1806880": "Yes, conversion to `float16` causes overflow for some functions. I changed the data type to `float32` due to this issue.",
    "1806889": "You can convert your feature to float32 (1 column) before the computation. But this way you cannot do the multi-column aggregation, because casting all numerical columns to float32 will cause OOM on a Kaggle notebook.\n\nFor example:\n```\ntemp_df = df[[\"customer_ID\",  feature]]\ntemp_df[feature] = temp_df[feature].astype(np.float32)\ntemp_df.groupby(\"customer_ID\")[feature].mean().astype(np.float16)\n```\n\nBut this is too troublesome, I switched to using float32 completely with my local machine instead.\n\n\nHaha, this reminds me of the famous interview question of computing the mid point during binary search\n`mid = low + ((high - low) / 2);`\nvs\n`mid = (low + high) / 2;`\n\nThe second middle point calculation has a chance of overflow if both low and high are close to the max of its variable type limit.",
    "1812473": "What a good point, I was just experiencing that, thank you very much for sharing."
  },
  "source": "meta"
}