{
  "id": 327228,
  "title": "6.53GB Dataset + Kaggle Notebook training and Inference Pipeline",
  "url": "/competitions/amex-default-prediction/discussion/327228",
  "author_name": "",
  "post_date": "2022-05-26T08:59:22.099717800Z",
  "votes": 21,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Here is my contribution to the data compression game.</p>\n<h3>Methods</h3>\n<ul>\n<li><p>Label Encode <code>customer_ID</code><br>\nI think some of the Kagglers already noticed that the <code>customer_ID</code> is very huge. </p></li>\n<li><p>THE Kaggle <code>reduce_mem_usage</code> function<br>\nI modified this function to optionally compress the data into float16.</p></li>\n<li><p>Pickle files with protocol = 4</p></li>\n</ul>\n<h3>Datasets</h3>\n<p>Overall, I created 3 datasets. <br>\nThe actual code for generating the dataset is under the description of the dataset page.</p>\n<ul>\n<li><p>Only Label Encode <code>customer_ID</code> without compression [~22GB]<br>\n<a href=\"https://drive.google.com/drive/folders/1Hsobf--EYaG26G4yNQ0YcBCR6E_3zW-7?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1Hsobf--EYaG26G4yNQ0YcBCR6E_3zW-7?usp=sharing</a></p></li>\n<li><p>Label Encode <code>customer_ID</code> + at most float32 compression [12.78GB]<br>\n<a href=\"https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp32\" target=\"_blank\">AE Credit ID Encoded Dataset [FP32]</a></p></li>\n<li><p>Label Encode <code>customer_ID</code> + at most float16 compression [6.53GB]<br>\n<a href=\"https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp16\" target=\"_blank\">AE Credit ID Encoded Dataset [FP16]</a></p></li>\n</ul>\n<h3>Inverse Encode the <code>customer_ID</code></h3>\n<p>When we make a submission, we can to inverse_transform the id encoding:</p>\n<pre><code>loaded_encoder = LabelEncoder()\nloaded_encoder.classes_ = np.load(f\"{base_path}/id_encodings.npy\", allow_pickle=True)\nsample_submission[\"customer_ID\"] = loaded_encoder.inverse_transform(sample_submission[\"customer_ID\"])\n</code></pre>\n<h3>Notebooks</h3>\n<p>The FP16 version can be fully loaded into a Kaggle Notebook. Here is a Logistic Regression Train + Predict notebook [It nearly makes all 0s haha]:<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/logistic-regression-id-encoded-fp16-dataset\" target=\"_blank\">Logistic Regression ID Encoded FP16 Dataset</a></p>\n<p>I applied the FP16 dataset to the current best public notebook (WoE Baseline):<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/amex-woe-baseline-with-id-encoded-fp16-dataset\" target=\"_blank\">AMEX - WoE Baseline with ID encoded FP16 Dataset</a></p>\n<ul>\n<li>I just replace the dataset, no feature selection is done, so the score is significantly lower.</li>\n<li>Notice that, in the original <a href=\"https://www.kaggle.com/code/lucasmorin/amex-woe-baseline\" target=\"_blank\">AMEX - WoE Baseline Notebook</a>, the private dataset contains only 354 columns after <code>prepare_df</code>.</li>\n</ul>\n<p>I also applied the dataset to <a href=\"https://www.kaggle.com/code/aninda/first-submission-using-catboost\" target=\"_blank\">First submission using CatBoost</a> notebook and both the CV and LB  becomes slightly better:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/kingychiu/catboost-with-id-encoded-fp16-dataset?scriptVersionId=96645239\" target=\"_blank\">CatBoost with ID encoded FP16 Dataset </a></li>\n</ul>",
  "messages": [
    {
      "id": "1801895",
      "postDate": "05/26/2022 08:59:22",
      "content": "<p>Here is my contribution to the data compression game.</p>\n<h3>Methods</h3>\n<ul>\n<li><p>Label Encode <code>customer_ID</code><br>\nI think some of the Kagglers already noticed that the <code>customer_ID</code> is very huge. </p></li>\n<li><p>THE Kaggle <code>reduce_mem_usage</code> function<br>\nI modified this function to optionally compress the data into float16.</p></li>\n<li><p>Pickle files with protocol = 4</p></li>\n</ul>\n<h3>Datasets</h3>\n<p>Overall, I created 3 datasets. <br>\nThe actual code for generating the dataset is under the description of the dataset page.</p>\n<ul>\n<li><p>Only Label Encode <code>customer_ID</code> without compression [~22GB]<br>\n<a href=\"https://drive.google.com/drive/folders/1Hsobf--EYaG26G4yNQ0YcBCR6E_3zW-7?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1Hsobf--EYaG26G4yNQ0YcBCR6E_3zW-7?usp=sharing</a></p></li>\n<li><p>Label Encode <code>customer_ID</code> + at most float32 compression [12.78GB]<br>\n<a href=\"https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp32\" target=\"_blank\">AE Credit ID Encoded Dataset [FP32]</a></p></li>\n<li><p>Label Encode <code>customer_ID</code> + at most float16 compression [6.53GB]<br>\n<a href=\"https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp16\" target=\"_blank\">AE Credit ID Encoded Dataset [FP16]</a></p></li>\n</ul>\n<h3>Inverse Encode the <code>customer_ID</code></h3>\n<p>When we make a submission, we can to inverse_transform the id encoding:</p>\n<pre><code>loaded_encoder = LabelEncoder()\nloaded_encoder.classes_ = np.load(f\"{base_path}/id_encodings.npy\", allow_pickle=True)\nsample_submission[\"customer_ID\"] = loaded_encoder.inverse_transform(sample_submission[\"customer_ID\"])\n</code></pre>\n<h3>Notebooks</h3>\n<p>The FP16 version can be fully loaded into a Kaggle Notebook. Here is a Logistic Regression Train + Predict notebook [It nearly makes all 0s haha]:<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/logistic-regression-id-encoded-fp16-dataset\" target=\"_blank\">Logistic Regression ID Encoded FP16 Dataset</a></p>\n<p>I applied the FP16 dataset to the current best public notebook (WoE Baseline):<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/amex-woe-baseline-with-id-encoded-fp16-dataset\" target=\"_blank\">AMEX - WoE Baseline with ID encoded FP16 Dataset</a></p>\n<ul>\n<li>I just replace the dataset, no feature selection is done, so the score is significantly lower.</li>\n<li>Notice that, in the original <a href=\"https://www.kaggle.com/code/lucasmorin/amex-woe-baseline\" target=\"_blank\">AMEX - WoE Baseline Notebook</a>, the private dataset contains only 354 columns after <code>prepare_df</code>.</li>\n</ul>\n<p>I also applied the dataset to <a href=\"https://www.kaggle.com/code/aninda/first-submission-using-catboost\" target=\"_blank\">First submission using CatBoost</a> notebook and both the CV and LB  becomes slightly better:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/kingychiu/catboost-with-id-encoded-fp16-dataset?scriptVersionId=96645239\" target=\"_blank\">CatBoost with ID encoded FP16 Dataset </a></li>\n</ul>",
      "rawMarkdown": "Here is my contribution to the data compression game.\n\n### Methods\n- Label Encode `customer_ID`\nI think some of the Kagglers already noticed that the `customer_ID` is very huge. \n\n- THE Kaggle `reduce_mem_usage` function\nI modified this function to optionally compress the data into float16.\n\n- Pickle files with protocol = 4\n\n\n### Datasets\nOverall, I created 3 datasets. \nThe actual code for generating the dataset is under the description of the dataset page.\n\n- Only Label Encode `customer_ID` without compression [~22GB]\nhttps://drive.google.com/drive/folders/1Hsobf--EYaG26G4yNQ0YcBCR6E_3zW-7?usp=sharing\n\n- Label Encode `customer_ID` + at most float32 compression [12.78GB]\n[AE Credit ID Encoded Dataset [FP32]](https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp32)\n\n- Label Encode `customer_ID` + at most float16 compression [6.53GB]\n[AE Credit ID Encoded Dataset [FP16]](https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp16)\n\n\n### Inverse Encode the `customer_ID`\nWhen we make a submission, we can to inverse_transform the id encoding:\n```python\nloaded_encoder = LabelEncoder()\nloaded_encoder.classes_ = np.load(f\"{base_path}/id_encodings.npy\", allow_pickle=True)\nsample_submission[\"customer_ID\"] = loaded_encoder.inverse_transform(sample_submission[\"customer_ID\"])\n```\n\n### Notebooks\nThe FP16 version can be fully loaded into a Kaggle Notebook. Here is a Logistic Regression Train + Predict notebook [It nearly makes all 0s haha]:\n[Logistic Regression ID Encoded FP16 Dataset](https://www.kaggle.com/code/kingychiu/logistic-regression-id-encoded-fp16-dataset)\n\n\nI applied the FP16 dataset to the current best public notebook (WoE Baseline):\n[AMEX - WoE Baseline with ID encoded FP16 Dataset](https://www.kaggle.com/code/kingychiu/amex-woe-baseline-with-id-encoded-fp16-dataset)\n\n\n- I just replace the dataset, no feature selection is done, so the score is significantly lower.\n- Notice that, in the original [AMEX - WoE Baseline Notebook](https://www.kaggle.com/code/lucasmorin/amex-woe-baseline), the private dataset contains only 354 columns after `prepare_df`.\n\n\nI also applied the dataset to [First submission using CatBoost](https://www.kaggle.com/code/aninda/first-submission-using-catboost) notebook and both the CV and LB  becomes slightly better:\n- [CatBoost with ID encoded FP16 Dataset ](https://www.kaggle.com/code/kingychiu/catboost-with-id-encoded-fp16-dataset?scriptVersionId=96645239)",
      "votes": null
    },
    {
      "id": "1802089",
      "postDate": "05/26/2022 12:53:20",
      "content": "<p>Thanks. These are the files I need.</p>",
      "rawMarkdown": "Thanks. These are the files I need.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1802089,
      "author_name": "yus002",
      "author_url": "",
      "post_date": "05/26/2022 12:53:20",
      "content": "<p>Thanks. These are the files I need.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801895": "Here is my contribution to the data compression game.\n\n### Methods\n- Label Encode `customer_ID`\nI think some of the Kagglers already noticed that the `customer_ID` is very huge. \n\n- THE Kaggle `reduce_mem_usage` function\nI modified this function to optionally compress the data into float16.\n\n- Pickle files with protocol = 4\n\n\n### Datasets\nOverall, I created 3 datasets. \nThe actual code for generating the dataset is under the description of the dataset page.\n\n- Only Label Encode `customer_ID` without compression [~22GB]\nhttps://drive.google.com/drive/folders/1Hsobf--EYaG26G4yNQ0YcBCR6E_3zW-7?usp=sharing\n\n- Label Encode `customer_ID` + at most float32 compression [12.78GB]\n[AE Credit ID Encoded Dataset [FP32]](https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp32)\n\n- Label Encode `customer_ID` + at most float16 compression [6.53GB]\n[AE Credit ID Encoded Dataset [FP16]](https://www.kaggle.com/datasets/kingychiu/ae-credit-id-encoded-dataset-fp16)\n\n\n### Inverse Encode the `customer_ID`\nWhen we make a submission, we can to inverse_transform the id encoding:\n```python\nloaded_encoder = LabelEncoder()\nloaded_encoder.classes_ = np.load(f\"{base_path}/id_encodings.npy\", allow_pickle=True)\nsample_submission[\"customer_ID\"] = loaded_encoder.inverse_transform(sample_submission[\"customer_ID\"])\n```\n\n### Notebooks\nThe FP16 version can be fully loaded into a Kaggle Notebook. Here is a Logistic Regression Train + Predict notebook [It nearly makes all 0s haha]:\n[Logistic Regression ID Encoded FP16 Dataset](https://www.kaggle.com/code/kingychiu/logistic-regression-id-encoded-fp16-dataset)\n\n\nI applied the FP16 dataset to the current best public notebook (WoE Baseline):\n[AMEX - WoE Baseline with ID encoded FP16 Dataset](https://www.kaggle.com/code/kingychiu/amex-woe-baseline-with-id-encoded-fp16-dataset)\n\n\n- I just replace the dataset, no feature selection is done, so the score is significantly lower.\n- Notice that, in the original [AMEX - WoE Baseline Notebook](https://www.kaggle.com/code/lucasmorin/amex-woe-baseline), the private dataset contains only 354 columns after `prepare_df`.\n\n\nI also applied the dataset to [First submission using CatBoost](https://www.kaggle.com/code/aninda/first-submission-using-catboost) notebook and both the CV and LB  becomes slightly better:\n- [CatBoost with ID encoded FP16 Dataset ](https://www.kaggle.com/code/kingychiu/catboost-with-id-encoded-fp16-dataset?scriptVersionId=96645239)",
    "1802089": "Thanks. These are the files I need."
  },
  "source": "meta"
}