{
  "id": 333864,
  "title": "Data (train, test)for best performing model? ",
  "url": "/competitions/amex-default-prediction/discussion/333864",
  "author_name": "",
  "post_date": "2022-06-28T15:30:12.739165Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi Folks,</p>\n<p>I am wondering what does the best pre-processed data looks like which could be used to build model that could achieve best results. </p>\n<p>Appreciate you all sharing how you went about preparing train and test data to build your best performing model. </p>",
  "messages": [
    {
      "id": "1836382",
      "postDate": "06/28/2022 15:30:12",
      "content": "<p>Hi Folks,</p>\n<p>I am wondering what does the best pre-processed data looks like which could be used to build model that could achieve best results. </p>\n<p>Appreciate you all sharing how you went about preparing train and test data to build your best performing model. </p>",
      "rawMarkdown": "Hi Folks,\n\nI am wondering what does the best pre-processed data looks like which could be used to build model that could achieve best results. \n\nAppreciate you all sharing how you went about preparing train and test data to build your best performing model.",
      "votes": null
    },
    {
      "id": "1836654",
      "postDate": "06/28/2022 22:57:13",
      "content": "<p>I've shared a notebook for generic feature engineering: <a href=\"https://www.kaggle.com/code/lucasmorin/amex-feature-engineering\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-feature-engineering</a><br>\nThe idea is to work in chunks to avoid oom errors. </p>",
      "rawMarkdown": "I've shared a notebook for generic feature engineering: https://www.kaggle.com/code/lucasmorin/amex-feature-engineering\nThe idea is to work in chunks to avoid oom errors.",
      "votes": null
    },
    {
      "id": "1836922",
      "postDate": "06/29/2022 08:06:48",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/bmkhandu\" target=\"_blank\">@bmkhandu</a>, </p>\n<p>I recommend using <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">parquet data with noise reduction.</a> You can check out his discussion board <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. </p>\n<p>I used these parquet files without noise to aggregate the data and do feature engineering, which improved my score a bit more. You can find my dataset <a href=\"https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise\" target=\"_blank\">here</a> and my <a href=\"https://www.kaggle.com/code/heyspaceturtle/data-aggregation\" target=\"_blank\">notebook </a></p>\n<p>Cheers!</p>",
      "rawMarkdown": "Dear @bmkhandu, \n\nI recommend using @raddar [parquet data with noise reduction.](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) You can check out his discussion board [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514). \n\nI used these parquet files without noise to aggregate the data and do feature engineering, which improved my score a bit more. You can find my dataset [here](https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise) and my [notebook ](https://www.kaggle.com/code/heyspaceturtle/data-aggregation)\n\nCheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1836654,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "06/28/2022 22:57:13",
      "content": "<p>I've shared a notebook for generic feature engineering: <a href=\"https://www.kaggle.com/code/lucasmorin/amex-feature-engineering\" target=\"_blank\">https://www.kaggle.com/code/lucasmorin/amex-feature-engineering</a><br>\nThe idea is to work in chunks to avoid oom errors. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836922,
      "author_name": "heyspaceturtle",
      "author_url": "",
      "post_date": "06/29/2022 08:06:48",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/bmkhandu\" target=\"_blank\">@bmkhandu</a>, </p>\n<p>I recommend using <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">parquet data with noise reduction.</a> You can check out his discussion board <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. </p>\n<p>I used these parquet files without noise to aggregate the data and do feature engineering, which improved my score a bit more. You can find my dataset <a href=\"https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise\" target=\"_blank\">here</a> and my <a href=\"https://www.kaggle.com/code/heyspaceturtle/data-aggregation\" target=\"_blank\">notebook </a></p>\n<p>Cheers!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1836382": "Hi Folks,\n\nI am wondering what does the best pre-processed data looks like which could be used to build model that could achieve best results. \n\nAppreciate you all sharing how you went about preparing train and test data to build your best performing model.",
    "1836654": "I've shared a notebook for generic feature engineering: https://www.kaggle.com/code/lucasmorin/amex-feature-engineering\nThe idea is to work in chunks to avoid oom errors.",
    "1836922": "Dear @bmkhandu, \n\nI recommend using @raddar [parquet data with noise reduction.](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) You can check out his discussion board [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514). \n\nI used these parquet files without noise to aggregate the data and do feature engineering, which improved my score a bit more. You can find my dataset [here](https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise) and my [notebook ](https://www.kaggle.com/code/heyspaceturtle/data-aggregation)\n\nCheers!"
  },
  "source": "meta"
}