{
  "id": 328057,
  "title": "Analysis of Information loss during conversion from float64 to float16",
  "url": "/competitions/amex-default-prediction/discussion/328057",
  "author_name": "",
  "post_date": "2022-05-30T17:03:34.087762800Z",
  "votes": 28,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, since we have a huge train and test dataset, one way of reducing the memory space, is by converting columns with float64 to float16. I also did this analysis <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142\" target=\"_blank\">in this post</a>. Totally, we have <strong>190 features, of which 185 columns are with float64 data type.</strong></p>\n<p>By this conversion from <code>float64</code> to <code>float16</code>, the precision of values in these columns gets reduced drastically (gets truncated). This leads to information loss. Since the rows are high, I got curious to know the amount of information loss due to this precision loss (datatype from float64 to float16)</p>\n<p>So, what did I do? I imported the train data in small chunks with each column individually and with the original datatype - float64. Now, after dropping NaN, I captured the variance of that imported column. Similarly, I did this again by importing the data in small chunks, this time with datatype - float16, and captured the variance.</p>\n<p>The difference in the variance of these two datatypes is the information loss during data type conversion. Below are the top 5 features with max loss in percentage.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>feature</th>\n<th>var_float64</th>\n<th>var_float16</th>\n<th>loss_pct</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>D_106</td>\n<td>1.581541</td>\n<td>1.582124</td>\n<td><strong>0.036892</strong></td>\n</tr>\n<tr>\n<td>2</td>\n<td>B_12</td>\n<td>0.6731</td>\n<td>0.672979</td>\n<td><strong>0.018006</strong></td>\n</tr>\n<tr>\n<td>3</td>\n<td>B_10</td>\n<td>23.038513</td>\n<td>23.034945</td>\n<td><strong>0.015488</strong></td>\n</tr>\n<tr>\n<td>4</td>\n<td>D_69</td>\n<td>256.309</td>\n<td>256.284049</td>\n<td><strong>0.009735</strong></td>\n</tr>\n<tr>\n<td>5</td>\n<td>S_12</td>\n<td>0.06286</td>\n<td>0.062866</td>\n<td><strong>0.009104</strong></td>\n</tr>\n</tbody>\n</table>\n<p>The total sum of this loss_pct column for all 185 columns, is going to give the total loss of information in percentage because of datatype conversion in the training dataset. And, this value is <strong>0.263%</strong></p>\n<p>Finally, <strong>there is a loss of 0.263% variance out of total variance in the train dataset, due to datatype conversion from float64 to float16</strong>, and for sure <strong>this is not going to impact the model performance during the training</strong> process. Also, on the brighter side, we are able to see a huge reduction in memory space, which helps in managing the computation for model training. </p>",
  "messages": [
    {
      "id": "1805985",
      "postDate": "05/30/2022 17:03:34",
      "content": "<p>Hi, since we have a huge train and test dataset, one way of reducing the memory space, is by converting columns with float64 to float16. I also did this analysis <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142\" target=\"_blank\">in this post</a>. Totally, we have <strong>190 features, of which 185 columns are with float64 data type.</strong></p>\n<p>By this conversion from <code>float64</code> to <code>float16</code>, the precision of values in these columns gets reduced drastically (gets truncated). This leads to information loss. Since the rows are high, I got curious to know the amount of information loss due to this precision loss (datatype from float64 to float16)</p>\n<p>So, what did I do? I imported the train data in small chunks with each column individually and with the original datatype - float64. Now, after dropping NaN, I captured the variance of that imported column. Similarly, I did this again by importing the data in small chunks, this time with datatype - float16, and captured the variance.</p>\n<p>The difference in the variance of these two datatypes is the information loss during data type conversion. Below are the top 5 features with max loss in percentage.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>feature</th>\n<th>var_float64</th>\n<th>var_float16</th>\n<th>loss_pct</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>D_106</td>\n<td>1.581541</td>\n<td>1.582124</td>\n<td><strong>0.036892</strong></td>\n</tr>\n<tr>\n<td>2</td>\n<td>B_12</td>\n<td>0.6731</td>\n<td>0.672979</td>\n<td><strong>0.018006</strong></td>\n</tr>\n<tr>\n<td>3</td>\n<td>B_10</td>\n<td>23.038513</td>\n<td>23.034945</td>\n<td><strong>0.015488</strong></td>\n</tr>\n<tr>\n<td>4</td>\n<td>D_69</td>\n<td>256.309</td>\n<td>256.284049</td>\n<td><strong>0.009735</strong></td>\n</tr>\n<tr>\n<td>5</td>\n<td>S_12</td>\n<td>0.06286</td>\n<td>0.062866</td>\n<td><strong>0.009104</strong></td>\n</tr>\n</tbody>\n</table>\n<p>The total sum of this loss_pct column for all 185 columns, is going to give the total loss of information in percentage because of datatype conversion in the training dataset. And, this value is <strong>0.263%</strong></p>\n<p>Finally, <strong>there is a loss of 0.263% variance out of total variance in the train dataset, due to datatype conversion from float64 to float16</strong>, and for sure <strong>this is not going to impact the model performance during the training</strong> process. Also, on the brighter side, we are able to see a huge reduction in memory space, which helps in managing the computation for model training. </p>",
      "rawMarkdown": "Hi, since we have a huge train and test dataset, one way of reducing the memory space, is by converting columns with float64 to float16. I also did this analysis [in this post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142). Totally, we have **190 features, of which 185 columns are with float64 data type.**\n\nBy this conversion from `float64` to `float16`, the precision of values in these columns gets reduced drastically (gets truncated). This leads to information loss. Since the rows are high, I got curious to know the amount of information loss due to this precision loss (datatype from float64 to float16)\n\nSo, what did I do? I imported the train data in small chunks with each column individually and with the original datatype - float64. Now, after dropping NaN, I captured the variance of that imported column. Similarly, I did this again by importing the data in small chunks, this time with datatype - float16, and captured the variance.\n\nThe difference in the variance of these two datatypes is the information loss during data type conversion. Below are the top 5 features with max loss in percentage.\n\n|\t|feature|\tvar_float64|\tvar_float16|\tloss_pct|\n| ---|---|---|---|\n|1|D_106|\t1.581541|\t1.582124|\t**0.036892**|\n|2|\tB_12|\t0.6731|\t0.672979|\t**0.018006**|\n|3|\tB_10|\t23.038513|\t23.034945|\t**0.015488**|\n|4|\tD_69|\t256.309|\t256.284049|\t**0.009735**|\n|5|\tS_12\t|0.06286|\t0.062866|\t**0.009104**|\n\nThe total sum of this loss_pct column for all 185 columns, is going to give the total loss of information in percentage because of datatype conversion in the training dataset. And, this value is **0.263%**\n\nFinally, **there is a loss of 0.263% variance out of total variance in the train dataset, due to datatype conversion from float64 to float16**, and for sure **this is not going to impact the model performance during the training** process. Also, on the brighter side, we are able to see a huge reduction in memory space, which helps in managing the computation for model training.",
      "votes": null
    },
    {
      "id": "1805988",
      "postDate": "05/30/2022 17:05:17",
      "content": "<p>I think this should not have a great impact in the training process. The advantage of using a smaller sized train data outperforms the associated data difference. </p>",
      "rawMarkdown": "I think this should not have a great impact in the training process. The advantage of using a smaller sized train data outperforms the associated data difference.",
      "votes": null
    },
    {
      "id": "1805991",
      "postDate": "05/30/2022 17:08:21",
      "content": "<p>Yes, of course. But, I got curious to do this analysis.</p>",
      "rawMarkdown": "Yes, of course. But, I got curious to do this analysis.",
      "votes": null
    },
    {
      "id": "1827701",
      "postDate": "06/21/2022 08:36:59",
      "content": "<p><strong>Good training <a href=\"https://www.kaggle.com/balabaskar\" target=\"_blank\">@balabaskar</a> 👍 keep your progress up 😊</strong></p>",
      "rawMarkdown": "**Good training @balabaskar 👍 keep your progress up 😊**",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1805988,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/30/2022 17:05:17",
      "content": "<p>I think this should not have a great impact in the training process. The advantage of using a smaller sized train data outperforms the associated data difference. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1805991,
          "author_name": "balabaskar",
          "author_url": "",
          "post_date": "05/30/2022 17:08:21",
          "content": "<p>Yes, of course. But, I got curious to do this analysis.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1827701,
      "author_name": "nancyalaswad90",
      "author_url": "",
      "post_date": "06/21/2022 08:36:59",
      "content": "<p><strong>Good training <a href=\"https://www.kaggle.com/balabaskar\" target=\"_blank\">@balabaskar</a> 👍 keep your progress up 😊</strong></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1805985": "Hi, since we have a huge train and test dataset, one way of reducing the memory space, is by converting columns with float64 to float16. I also did this analysis [in this post](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327142). Totally, we have **190 features, of which 185 columns are with float64 data type.**\n\nBy this conversion from `float64` to `float16`, the precision of values in these columns gets reduced drastically (gets truncated). This leads to information loss. Since the rows are high, I got curious to know the amount of information loss due to this precision loss (datatype from float64 to float16)\n\nSo, what did I do? I imported the train data in small chunks with each column individually and with the original datatype - float64. Now, after dropping NaN, I captured the variance of that imported column. Similarly, I did this again by importing the data in small chunks, this time with datatype - float16, and captured the variance.\n\nThe difference in the variance of these two datatypes is the information loss during data type conversion. Below are the top 5 features with max loss in percentage.\n\n|\t|feature|\tvar_float64|\tvar_float16|\tloss_pct|\n| ---|---|---|---|\n|1|D_106|\t1.581541|\t1.582124|\t**0.036892**|\n|2|\tB_12|\t0.6731|\t0.672979|\t**0.018006**|\n|3|\tB_10|\t23.038513|\t23.034945|\t**0.015488**|\n|4|\tD_69|\t256.309|\t256.284049|\t**0.009735**|\n|5|\tS_12\t|0.06286|\t0.062866|\t**0.009104**|\n\nThe total sum of this loss_pct column for all 185 columns, is going to give the total loss of information in percentage because of datatype conversion in the training dataset. And, this value is **0.263%**\n\nFinally, **there is a loss of 0.263% variance out of total variance in the train dataset, due to datatype conversion from float64 to float16**, and for sure **this is not going to impact the model performance during the training** process. Also, on the brighter side, we are able to see a huge reduction in memory space, which helps in managing the computation for model training.",
    "1805988": "I think this should not have a great impact in the training process. The advantage of using a smaller sized train data outperforms the associated data difference.",
    "1805991": "Yes, of course. But, I got curious to do this analysis.",
    "1827701": "**Good training @balabaskar 👍 keep your progress up 😊**"
  },
  "source": "meta"
}