{
  "id": 331778,
  "title": "Compressed Data Set for Amex - Default Prediction",
  "url": "/competitions/amex-default-prediction/discussion/331778",
  "author_name": "",
  "post_date": "2022-06-18T19:16:48.366585500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi folks,  please find the link for the compressed dataset below:</p>\n<p><strong><a href=\"https://www.kaggle.com/datasets/khangjrakpamarjun/datasets-train-test-for-amex-parquet\" target=\"_blank\">https://www.kaggle.com/datasets/khangjrakpamarjun/datasets-train-test-for-amex-parquet</a></strong></p>\n<p>Following is how I have compressed the data:</p>\n<ol>\n<li><p>The Data was compressed by downcasting the int and float variable to its smallest possible dtype. No change was done for the object type variables.</p></li>\n<li><p>Compressed data was saved using feather format. </p></li>\n</ol>\n<p>No missing values were imputed in both the train and test datasets. There are exactly the same no. of rows and columns in both the train and test datasets.</p>",
  "messages": [
    {
      "id": "1824940",
      "postDate": "06/18/2022 19:16:48",
      "content": "<p>Hi folks,  please find the link for the compressed dataset below:</p>\n<p><strong><a href=\"https://www.kaggle.com/datasets/khangjrakpamarjun/datasets-train-test-for-amex-parquet\" target=\"_blank\">https://www.kaggle.com/datasets/khangjrakpamarjun/datasets-train-test-for-amex-parquet</a></strong></p>\n<p>Following is how I have compressed the data:</p>\n<ol>\n<li><p>The Data was compressed by downcasting the int and float variable to its smallest possible dtype. No change was done for the object type variables.</p></li>\n<li><p>Compressed data was saved using feather format. </p></li>\n</ol>\n<p>No missing values were imputed in both the train and test datasets. There are exactly the same no. of rows and columns in both the train and test datasets.</p>",
      "rawMarkdown": "Hi folks,  please find the link for the compressed dataset below:\n\n**https://www.kaggle.com/datasets/khangjrakpamarjun/datasets-train-test-for-amex-parquet**\n\nFollowing is how I have compressed the data:\n \n1. The Data was compressed by downcasting the int and float variable to its smallest possible dtype. No change was done for the object type variables.\n\n2. Compressed data was saved using feather format. \n\nNo missing values were imputed in both the train and test datasets. There are exactly the same no. of rows and columns in both the train and test datasets.",
      "votes": null
    },
    {
      "id": "1824999",
      "postDate": "06/18/2022 20:22:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/khangjrakpamarjun\" target=\"_blank\">@khangjrakpamarjun</a> Perhaps you could describe how you compressed the dataset so that people can take an informed decision whether they want to use it.</p>",
      "rawMarkdown": "Hi @khangjrakpamarjun Perhaps you could describe how you compressed the dataset so that people can take an informed decision whether they want to use it.",
      "votes": null
    },
    {
      "id": "1825271",
      "postDate": "06/19/2022 07:08:52",
      "content": "<p>Kindly describe the compression logic please</p>",
      "rawMarkdown": "Kindly describe the compression logic please",
      "votes": null
    },
    {
      "id": "1825287",
      "postDate": "06/19/2022 07:46:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> sure, I have added it to the discription. Do let me know if you have further suggestions.</p>",
      "rawMarkdown": "Hi @ambrosm sure, I have added it to the discription. Do let me know if you have further suggestions.",
      "votes": null
    },
    {
      "id": "1825290",
      "postDate": "06/19/2022 07:47:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> , I have added it to the discription. Let me know if you have further queries.</p>",
      "rawMarkdown": "Hi @ravi20076 , I have added it to the discription. Let me know if you have further queries.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1824999,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "06/18/2022 20:22:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/khangjrakpamarjun\" target=\"_blank\">@khangjrakpamarjun</a> Perhaps you could describe how you compressed the dataset so that people can take an informed decision whether they want to use it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1825287,
          "author_name": "khangjrakpamarjun",
          "author_url": "",
          "post_date": "06/19/2022 07:46:19",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> sure, I have added it to the discription. Do let me know if you have further suggestions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1825271,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/19/2022 07:08:52",
      "content": "<p>Kindly describe the compression logic please</p>",
      "votes": null,
      "replies": [
        {
          "id": 1825290,
          "author_name": "khangjrakpamarjun",
          "author_url": "",
          "post_date": "06/19/2022 07:47:17",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> , I have added it to the discription. Let me know if you have further queries.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1824940": "Hi folks,  please find the link for the compressed dataset below:\n\n**https://www.kaggle.com/datasets/khangjrakpamarjun/datasets-train-test-for-amex-parquet**\n\nFollowing is how I have compressed the data:\n \n1. The Data was compressed by downcasting the int and float variable to its smallest possible dtype. No change was done for the object type variables.\n\n2. Compressed data was saved using feather format. \n\nNo missing values were imputed in both the train and test datasets. There are exactly the same no. of rows and columns in both the train and test datasets.",
    "1824999": "Hi @khangjrakpamarjun Perhaps you could describe how you compressed the dataset so that people can take an informed decision whether they want to use it.",
    "1825271": "Kindly describe the compression logic please",
    "1825287": "Hi @ambrosm sure, I have added it to the discription. Do let me know if you have further suggestions.",
    "1825290": "Hi @ravi20076 , I have added it to the discription. Let me know if you have further queries."
  },
  "source": "meta"
}