{
  "id": 473919,
  "title": "Minimizng Memory Usage through Optimal Datatype Selection",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/473919",
  "author_name": "",
  "post_date": "2024-02-06T14:09:20.489830900Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I think decreasing the memory usage to the lowest possible level is important in this competition. So, the following notebooks may be relevant here. They choose the optimal datatype for each variable:</p>\n<p>Pandas Version:<br>\n<a href=\"https://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb\" target=\"_blank\">https://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb</a></p>\n<p>Polars Version (Thanks to <a href=\"https://www.kaggle.com/demche\" target=\"_blank\">@demche</a>):<br>\n<a href=\"https://www.kaggle.com/code/demche/polars-memory-usage-optimization\" target=\"_blank\">https://www.kaggle.com/code/demche/polars-memory-usage-optimization</a></p>\n<p>Also if you want to use \"Year\" somewhere as a feature in your dataset, this line might be relevant:<br>\n<code>dataset['Year'] = (dataset['Year'] - 2000).astype('int8')     #Memory Efficient</code></p>",
  "messages": [
    {
      "id": "2638841",
      "postDate": "02/06/2024 14:09:20",
      "content": "<p>I think decreasing the memory usage to the lowest possible level is important in this competition. So, the following notebooks may be relevant here. They choose the optimal datatype for each variable:</p>\n<p>Pandas Version:<br>\n<a href=\"https://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb\" target=\"_blank\">https://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb</a></p>\n<p>Polars Version (Thanks to <a href=\"https://www.kaggle.com/demche\" target=\"_blank\">@demche</a>):<br>\n<a href=\"https://www.kaggle.com/code/demche/polars-memory-usage-optimization\" target=\"_blank\">https://www.kaggle.com/code/demche/polars-memory-usage-optimization</a></p>\n<p>Also if you want to use \"Year\" somewhere as a feature in your dataset, this line might be relevant:<br>\n<code>dataset['Year'] = (dataset['Year'] - 2000).astype('int8')     #Memory Efficient</code></p>",
      "rawMarkdown": "I think decreasing the memory usage to the lowest possible level is important in this competition. So, the following notebooks may be relevant here. They choose the optimal datatype for each variable:\n\nPandas Version:\nhttps://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb\n\nPolars Version (Thanks to @demche):\nhttps://www.kaggle.com/code/demche/polars-memory-usage-optimization\n\nAlso if you want to use \"Year\" somewhere as a feature in your dataset, this line might be relevant:\n`dataset['Year'] = (dataset['Year'] - 2000).astype('int8')     #Memory Efficient`",
      "votes": null
    },
    {
      "id": "2638944",
      "postDate": "02/06/2024 15:17:59",
      "content": "<p>That is clever, or you could give a chance to polars :)</p>",
      "rawMarkdown": "That is clever, or you could give a chance to polars :)",
      "votes": null
    },
    {
      "id": "2638984",
      "postDate": "02/06/2024 15:50:10",
      "content": "<p>Or maybe polars + this <br>\n=)</p>",
      "rawMarkdown": "Or maybe polars + this \n=)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2638944,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/06/2024 15:17:59",
      "content": "<p>That is clever, or you could give a chance to polars :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2638984,
          "author_name": "mohammad2012191",
          "author_url": "",
          "post_date": "02/06/2024 15:50:10",
          "content": "<p>Or maybe polars + this <br>\n=)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2638841": "I think decreasing the memory usage to the lowest possible level is important in this competition. So, the following notebooks may be relevant here. They choose the optimal datatype for each variable:\n\nPandas Version:\nhttps://www.kaggle.com/code/mohammad2012191/reduce-memory-usage-2gb-780mb\n\nPolars Version (Thanks to @demche):\nhttps://www.kaggle.com/code/demche/polars-memory-usage-optimization\n\nAlso if you want to use \"Year\" somewhere as a feature in your dataset, this line might be relevant:\n`dataset['Year'] = (dataset['Year'] - 2000).astype('int8')     #Memory Efficient`",
    "2638944": "That is clever, or you could give a chance to polars :)",
    "2638984": "Or maybe polars + this \n=)"
  },
  "source": "meta"
}