{
  "id": 327106,
  "title": "Training data starter",
  "url": "/competitions/amex-default-prediction/discussion/327106",
  "author_name": "Konrad Banachewicz",
  "post_date": "2022-05-25T15:54:25.101000",
  "votes": 35,
  "comment_count": 7,
  "views": 0,
  "content": "<p>It went roughly like this:</p>\n<ul>\n<li>tweet from <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> about this comp</li>\n<li>ecstasy</li>\n<li>start a new notebook, load entire data</li>\n<li>!@#%^)_* kind of reaction</li>\n</ul>\n<p>In order to spare others the same path (except for the first step perhaps, since you are already here ;-) here's a dataset compressed towards something manageable:</p>\n<ul>\n<li>all numerical columns cast to float16</li>\n<li>pickle </li>\n</ul>\n<p><a href=\"https://www.kaggle.com/datasets/konradb/pickled-shrunken-training-data\" target=\"_blank\">https://www.kaggle.com/datasets/konradb/pickled-shrunken-training-data</a></p>\n<p>Good luck.</p>",
  "messages": [
    {
      "id": 1801308,
      "postDate": "2022-05-25T15:54:25.100Z",
      "content": "<p>It went roughly like this:</p>\n<ul>\n<li>tweet from <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> about this comp</li>\n<li>ecstasy</li>\n<li>start a new notebook, load entire data</li>\n<li>!@#%^)_* kind of reaction</li>\n</ul>\n<p>In order to spare others the same path (except for the first step perhaps, since you are already here ;-) here's a dataset compressed towards something manageable:</p>\n<ul>\n<li>all numerical columns cast to float16</li>\n<li>pickle </li>\n</ul>\n<p><a href=\"https://www.kaggle.com/datasets/konradb/pickled-shrunken-training-data\" target=\"_blank\">https://www.kaggle.com/datasets/konradb/pickled-shrunken-training-data</a></p>\n<p>Good luck.</p>",
      "rawMarkdown": "It went roughly like this:\n- tweet from @inversion about this comp\n- ecstasy\n- start a new notebook, load entire data\n- !@#%^)_* kind of reaction\n\nIn order to spare others the same path (except for the first step perhaps, since you are already here ;-) here's a dataset compressed towards something manageable:\n\n- all numerical columns cast to float16\n- pickle \n\nhttps://www.kaggle.com/datasets/konradb/pickled-shrunken-training-data\n\nGood luck.\n",
      "votes": 35
    },
    {
      "id": 1801470,
      "postDate": "2022-05-25T19:33:29.467Z",
      "content": "<p>Thanks for the pickle <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> </p>\n<p>You can reduce the size even more (~400 Mb) if you convert the categorical columns.</p>\n<pre><code>categories = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\ndf[categories] = df[categories].astype(\"category\")\n</code></pre>",
      "rawMarkdown": "Thanks for the pickle @konradb \n\nYou can reduce the size even more (~400 Mb) if you convert the categorical columns.\n\n```\ncategories = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\ndf[categories] = df[categories].astype(\"category\")\n```",
      "votes": 4,
      "replies": [
        {
          "id": 1801984,
          "postDate": "2022-05-26T10:02:26Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1801988,
          "postDate": "2022-05-26T10:15:52.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> has created a dataset <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">with these column transformations</a>.</p>",
          "rawMarkdown": "@ruchi798 has created a dataset [with these column transformations](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143).\n",
          "votes": 1
        },
        {
          "id": 1802002,
          "postDate": "2022-05-26T10:39:13.747Z",
          "content": "<p><a href=\"https://www.kaggle.com/yashlab\" target=\"_blank\">@yashlab</a> No, I haven't. Probably <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> will update his dataset.</p>",
          "rawMarkdown": "@yashlab No, I haven't. Probably @konradb will update his dataset."
        }
      ]
    },
    {
      "id": 1801722,
      "postDate": "2022-05-26T05:00:10.667Z",
      "content": "<p>Hi,</p>\n<p>I have created pickled dataset (train &amp; test) through the notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/aninda/creating-smaller-train-test-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/aninda/creating-smaller-train-test-data/notebook</a></p>",
      "rawMarkdown": "Hi,\n\nI have created pickled dataset (train & test) through the notebook.\n\nhttps://www.kaggle.com/code/aninda/creating-smaller-train-test-data/notebook\n",
      "votes": 1
    },
    {
      "id": 1900117,
      "postDate": "2022-08-15T17:59:44.433Z",
      "content": "<p><a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> thanks a lot for the file! A laggard speaking here: does this compression to float16 affect in any way the quality of the dataset? I assume the original data was 64 bits float.</p>",
      "rawMarkdown": "@konradb thanks a lot for the file! A laggard speaking here: does this compression to float16 affect in any way the quality of the dataset? I assume the original data was 64 bits float.",
      "replies": [
        {
          "id": 1903365,
          "postDate": "2022-08-17T11:06:07.950Z",
          "content": "<p>Either that or 'object', yes.</p>",
          "rawMarkdown": "Either that or 'object', yes.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1801470,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2022-05-25T19:33:29.467000",
      "content": "<p>Thanks for the pickle <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> </p>\n<p>You can reduce the size even more (~400 Mb) if you convert the categorical columns.</p>\n<pre><code>categories = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\ndf[categories] = df[categories].astype(\"category\")\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 1801984,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-26T10:02:26",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1801988,
          "author_name": "DivyaRawat07",
          "author_url": "",
          "post_date": "2022-05-26T10:15:52.737000",
          "content": "<p><a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> has created a dataset <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">with these column transformations</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1802002,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2022-05-26T10:39:13.747000",
          "content": "<p><a href=\"https://www.kaggle.com/yashlab\" target=\"_blank\">@yashlab</a> No, I haven't. Probably <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> will update his dataset.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1801722,
      "author_name": "Aninda Goswamy",
      "author_url": "",
      "post_date": "2022-05-26T05:00:10.667000",
      "content": "<p>Hi,</p>\n<p>I have created pickled dataset (train &amp; test) through the notebook.</p>\n<p><a href=\"https://www.kaggle.com/code/aninda/creating-smaller-train-test-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/aninda/creating-smaller-train-test-data/notebook</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1900117,
      "author_name": "Antonio Intini",
      "author_url": "",
      "post_date": "2022-08-15T17:59:44.433000",
      "content": "<p><a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> thanks a lot for the file! A laggard speaking here: does this compression to float16 affect in any way the quality of the dataset? I assume the original data was 64 bits float.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1903365,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2022-08-17T11:06:07.950000",
          "content": "<p>Either that or 'object', yes.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1801308": "It went roughly like this:\n- tweet from @inversion about this comp\n- ecstasy\n- start a new notebook, load entire data\n- !@#%^)_* kind of reaction\n\nIn order to spare others the same path (except for the first step perhaps, since you are already here ;-) here's a dataset compressed towards something manageable:\n\n- all numerical columns cast to float16\n- pickle \n\nhttps://www.kaggle.com/datasets/konradb/pickled-shrunken-training-data\n\nGood luck.\n",
    "1801470": "Thanks for the pickle @konradb \n\nYou can reduce the size even more (~400 Mb) if you convert the categorical columns.\n\n```\ncategories = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\ndf[categories] = df[categories].astype(\"category\")\n```",
    "1801722": "Hi,\n\nI have created pickled dataset (train & test) through the notebook.\n\nhttps://www.kaggle.com/code/aninda/creating-smaller-train-test-data/notebook\n",
    "1900117": "@konradb thanks a lot for the file! A laggard speaking here: does this compression to float16 affect in any way the quality of the dataset? I assume the original data was 64 bits float."
  }
}