{
  "id": 331840,
  "title": "Summary: Basic pipeline for tabular competition for beginner!",
  "url": "/competitions/amex-default-prediction/discussion/331840",
  "author_name": "KhanhVD",
  "post_date": "2022-06-19T04:06:17.422000",
  "votes": 58,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I wanted to share a simple pipeline to start a tabular competition for beginners. It's been a long time since Kaggle had a true tabular competition and I wish it was easier for newbies to get started.</p>\n<h1>Step 1: Look at data + Preprocessing data: EDA, Data cleaning, Reduce data size without loss of information,..</h1>\n<p>The following references will be helpful:</p>\n<h2>Magic data format: random uniform noise added to each column</h2>\n<p>Link discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">Integer columns in the data - here you go!</a>- Raddar have found it: <strong>all float type columns have random uniform noise of [0,0.01] added to each column</strong></p>\n<ul>\n<li>originally we had 188 float/categorical type features. These were transformed into<ul>\n<li>95 np.int8/np.int16 types</li>\n<li>93 np.float32 types</li></ul></li>\n<li>Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.</li>\n<li>saved in parquet format (only 1.7GB training data!)<br>\nThis is really helpful when you look at the original data set size (50.31GB), now using Kaggle Notebook is possible and convenient for people with limited resources.</li>\n</ul>\n<h2>Step by Step: How to reduce data size</h2>\n<p>Link discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">How To Reduce Data Size</a>- Thanks Chris<br>\nRead the following helpful discussion by Chris if you want a deeper understanding of reduce data size. </p>\n<ul>\n<li>Step 1 - Reduce Data Types!: <ul>\n<li>convert int32 or int64 which only use 4 bytes or 8 bytes.</li>\n<li>Reduce 10 bytes to 3 bytes: date with time</li>\n<li>Reduce 88 bytes to 11 bytes: categorical columns: </li>\n<li>177 Numeric Columns - Reduce 1416 bytes to 353 bytes</li></ul></li>\n<li>Step 2 - Choose Your File Format</li>\n<li>Step 3 - Choose Multiple Files or Not</li>\n<li>Step 4 - Read Raddar's Discussion</li>\n</ul>\n<h2>Look at data: Missing value</h2>\n<p>Link discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756\" target=\"_blank\">The distribution of missing values over time</a><br>\n<strong>Note</strong>: The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data &gt;90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard</p>\n<h1>Step 2: Feature Engineering:</h1>\n<p>Engineering features is key to improving your LB score. We have seen some basic features from public notebook:</p>\n<ul>\n<li>Aggregated features: <a href=\"https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\" target=\"_blank\">Amex Agg Data How It Created</a></li>\n<li>Aggregated features + Rapid GPU for faster: <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">XGBoost Starter - [0.793]</a></li>\n<li>\"After-pay\" features. It makes intuitive semse that subtracting the payments from balance/spend etc provides new information about the users' behavior: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">RAPIDS cudf Feature Engineering + XGB</a></li>\n<li>Statistical features (mean, std, min, max,..): <a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">AMEX LightGBM Quickstart</a><ul>\n<li>Selected features averaged over all statements of a customer</li>\n<li>The minimum or maximum of selected features over all statements of a customer</li>\n<li>Selected features taken from the last statement of a customer</li></ul></li>\n</ul>\n<p><strong>Look on</strong></p>\n<ol>\n<li>Features with missing values</li>\n<li>Features with low variance</li>\n<li>Highly correlated features</li>\n<li>Univariate features</li>\n<li>Train/Test distributions</li>\n</ol>\n<p>Below are few note on how to engineer new features <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575\" target=\"_blank\">Feature Engineering Techniques</a> by Chris Deotte of Chris Deotte in IEEE-CIS Fraud Detection competition</p>\n<h1>Step 3: Feature Selection</h1>\n<p>You can read my discussions about feature selection here with basic feature selection techniques that everyone should know:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/optiver-realized-volatility-prediction/discussion/269283\" target=\"_blank\">Some Feature Selection Technique</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992\" target=\"_blank\">Few note about XGBoost &amp; Feature Engineering/Feature Selection</a></li>\n<li>AmbrosM's following discussion is also really helpful: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">Which is the right feature importance?</a>- In this post he compare the three methods for feature selection and show that the last one is the right one.<ul>\n<li>Split feature importance?</li>\n<li>Gain feature importance?</li>\n<li>Permutation feature importance?<br>\n<strong>Conclusion</strong>: Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.</li></ul></li>\n</ul>\n<h1>Step 4: Baseline Model &amp; Hyper-parameters optimization</h1>\n<p>There are many notebooks starting with tree-based model: LightGBM, XGBoost, CatBoost on both CPU and GPU (RAPID) you can refer.<br>\nSome note about tree-based from <a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">AMEX LightGBM Quickstart</a>: Thanks AmbrosM</p>\n<p>Preprocessing for LightGBM is much simpler than for neural networks:</p>\n<ol>\n<li>Neural networks can't process missing values; LightGBM handles them automatically.</li>\n<li>Categorical features need to be one-hot encoded for neural networks; LightGBM handles them automatically.</li>\n<li>With neural networks, you need to think about outliers; tree-based algorithms deal with outliers easily.</li>\n<li>Neural networks need scaled inputs; tree-based algorithms don't depend on scaling.</li>\n</ol>\n<p><strong>Importance</strong>: Validation score with the competition's scoring function</p>\n<p>Notebook example for using XGBoost with RAPID &amp; GPU faster: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">RAPIDS cudf Feature Engineering + XGB</a></p>\n<p><strong>XGBoost Tip &amp; Trick</strong>:  <a href=\"https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992\" target=\"_blank\">Few note about XGBoost &amp; Feature Engineering/Feature Selection</a></p>\n<h2>Hyper-parameters optimization</h2>\n<p>There are many method for Hyper-parameters optimization like:</p>\n<ul>\n<li>RandomizedSearchCV <a href=\"https://www.kaggle.com/code/funxexcel/p3-random-forest-tuning-randomizedsearchcv\" target=\"_blank\">example</a>, GridSearchCV <a href=\"https://www.kaggle.com/code/ihelon/titanic-hyperparameter-tuning-with-gridsearchcv\" target=\"_blank\">example</a></li>\n<li>Bayesian Optimization: <a href=\"https://www.kaggle.com/code/vincentlugat/ieee-lgb-bayesian-opt\" target=\"_blank\">example</a></li>\n<li>Optuna: <a href=\"https://www.kaggle.com/code/hamzaghanmi/lgbm-hyperparameter-tuning-using-optuna\" target=\"_blank\">example</a></li>\n</ul>\n<h1>Step 5: Ensemble Models</h1>\n<ul>\n<li>Weighted/Average Ensemble</li>\n<li>Stacking Model</li>\n</ul>\n<p><strong>[UPDATE…]</strong></p>",
  "messages": [
    {
      "id": 1825176,
      "postDate": "2022-06-19T04:06:17.423Z",
      "content": "<p>I wanted to share a simple pipeline to start a tabular competition for beginners. It's been a long time since Kaggle had a true tabular competition and I wish it was easier for newbies to get started.</p>\n<h1>Step 1: Look at data + Preprocessing data: EDA, Data cleaning, Reduce data size without loss of information,..</h1>\n<p>The following references will be helpful:</p>\n<h2>Magic data format: random uniform noise added to each column</h2>\n<p>Link discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">Integer columns in the data - here you go!</a>- Raddar have found it: <strong>all float type columns have random uniform noise of [0,0.01] added to each column</strong></p>\n<ul>\n<li>originally we had 188 float/categorical type features. These were transformed into<ul>\n<li>95 np.int8/np.int16 types</li>\n<li>93 np.float32 types</li></ul></li>\n<li>Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.</li>\n<li>saved in parquet format (only 1.7GB training data!)<br>\nThis is really helpful when you look at the original data set size (50.31GB), now using Kaggle Notebook is possible and convenient for people with limited resources.</li>\n</ul>\n<h2>Step by Step: How to reduce data size</h2>\n<p>Link discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">How To Reduce Data Size</a>- Thanks Chris<br>\nRead the following helpful discussion by Chris if you want a deeper understanding of reduce data size. </p>\n<ul>\n<li>Step 1 - Reduce Data Types!: <ul>\n<li>convert int32 or int64 which only use 4 bytes or 8 bytes.</li>\n<li>Reduce 10 bytes to 3 bytes: date with time</li>\n<li>Reduce 88 bytes to 11 bytes: categorical columns: </li>\n<li>177 Numeric Columns - Reduce 1416 bytes to 353 bytes</li></ul></li>\n<li>Step 2 - Choose Your File Format</li>\n<li>Step 3 - Choose Multiple Files or Not</li>\n<li>Step 4 - Read Raddar's Discussion</li>\n</ul>\n<h2>Look at data: Missing value</h2>\n<p>Link discussion: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756\" target=\"_blank\">The distribution of missing values over time</a><br>\n<strong>Note</strong>: The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data &gt;90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard</p>\n<h1>Step 2: Feature Engineering:</h1>\n<p>Engineering features is key to improving your LB score. We have seen some basic features from public notebook:</p>\n<ul>\n<li>Aggregated features: <a href=\"https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created\" target=\"_blank\">Amex Agg Data How It Created</a></li>\n<li>Aggregated features + Rapid GPU for faster: <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">XGBoost Starter - [0.793]</a></li>\n<li>\"After-pay\" features. It makes intuitive semse that subtracting the payments from balance/spend etc provides new information about the users' behavior: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">RAPIDS cudf Feature Engineering + XGB</a></li>\n<li>Statistical features (mean, std, min, max,..): <a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">AMEX LightGBM Quickstart</a><ul>\n<li>Selected features averaged over all statements of a customer</li>\n<li>The minimum or maximum of selected features over all statements of a customer</li>\n<li>Selected features taken from the last statement of a customer</li></ul></li>\n</ul>\n<p><strong>Look on</strong></p>\n<ol>\n<li>Features with missing values</li>\n<li>Features with low variance</li>\n<li>Highly correlated features</li>\n<li>Univariate features</li>\n<li>Train/Test distributions</li>\n</ol>\n<p>Below are few note on how to engineer new features <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575\" target=\"_blank\">Feature Engineering Techniques</a> by Chris Deotte of Chris Deotte in IEEE-CIS Fraud Detection competition</p>\n<h1>Step 3: Feature Selection</h1>\n<p>You can read my discussions about feature selection here with basic feature selection techniques that everyone should know:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/optiver-realized-volatility-prediction/discussion/269283\" target=\"_blank\">Some Feature Selection Technique</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992\" target=\"_blank\">Few note about XGBoost &amp; Feature Engineering/Feature Selection</a></li>\n<li>AmbrosM's following discussion is also really helpful: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131\" target=\"_blank\">Which is the right feature importance?</a>- In this post he compare the three methods for feature selection and show that the last one is the right one.<ul>\n<li>Split feature importance?</li>\n<li>Gain feature importance?</li>\n<li>Permutation feature importance?<br>\n<strong>Conclusion</strong>: Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.</li></ul></li>\n</ul>\n<h1>Step 4: Baseline Model &amp; Hyper-parameters optimization</h1>\n<p>There are many notebooks starting with tree-based model: LightGBM, XGBoost, CatBoost on both CPU and GPU (RAPID) you can refer.<br>\nSome note about tree-based from <a href=\"https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart\" target=\"_blank\">AMEX LightGBM Quickstart</a>: Thanks AmbrosM</p>\n<p>Preprocessing for LightGBM is much simpler than for neural networks:</p>\n<ol>\n<li>Neural networks can't process missing values; LightGBM handles them automatically.</li>\n<li>Categorical features need to be one-hot encoded for neural networks; LightGBM handles them automatically.</li>\n<li>With neural networks, you need to think about outliers; tree-based algorithms deal with outliers easily.</li>\n<li>Neural networks need scaled inputs; tree-based algorithms don't depend on scaling.</li>\n</ol>\n<p><strong>Importance</strong>: Validation score with the competition's scoring function</p>\n<p>Notebook example for using XGBoost with RAPID &amp; GPU faster: <a href=\"https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb\" target=\"_blank\">RAPIDS cudf Feature Engineering + XGB</a></p>\n<p><strong>XGBoost Tip &amp; Trick</strong>:  <a href=\"https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992\" target=\"_blank\">Few note about XGBoost &amp; Feature Engineering/Feature Selection</a></p>\n<h2>Hyper-parameters optimization</h2>\n<p>There are many method for Hyper-parameters optimization like:</p>\n<ul>\n<li>RandomizedSearchCV <a href=\"https://www.kaggle.com/code/funxexcel/p3-random-forest-tuning-randomizedsearchcv\" target=\"_blank\">example</a>, GridSearchCV <a href=\"https://www.kaggle.com/code/ihelon/titanic-hyperparameter-tuning-with-gridsearchcv\" target=\"_blank\">example</a></li>\n<li>Bayesian Optimization: <a href=\"https://www.kaggle.com/code/vincentlugat/ieee-lgb-bayesian-opt\" target=\"_blank\">example</a></li>\n<li>Optuna: <a href=\"https://www.kaggle.com/code/hamzaghanmi/lgbm-hyperparameter-tuning-using-optuna\" target=\"_blank\">example</a></li>\n</ul>\n<h1>Step 5: Ensemble Models</h1>\n<ul>\n<li>Weighted/Average Ensemble</li>\n<li>Stacking Model</li>\n</ul>\n<p><strong>[UPDATE…]</strong></p>",
      "rawMarkdown": "I wanted to share a simple pipeline to start a tabular competition for beginners. It's been a long time since Kaggle had a true tabular competition and I wish it was easier for newbies to get started.\n# Step 1: Look at data + Preprocessing data: EDA, Data cleaning, Reduce data size without loss of information,..\nThe following references will be helpful:\n## Magic data format: random uniform noise added to each column\nLink discussion: [Integer columns in the data - here you go!](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)- Raddar have found it: **all float type columns have random uniform noise of [0,0.01] added to each column**\n- originally we had 188 float/categorical type features. These were transformed into\n - 95 np.int8/np.int16 types\n - 93 np.float32 types\n- Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.\n- saved in parquet format (only 1.7GB training data!)\nThis is really helpful when you look at the original data set size (50.31GB), now using Kaggle Notebook is possible and convenient for people with limited resources.\n## Step by Step: How to reduce data size\nLink discussion: [How To Reduce Data Size](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054)- Thanks Chris\nRead the following helpful discussion by Chris if you want a deeper understanding of reduce data size. \n- Step 1 - Reduce Data Types!: \n - convert int32 or int64 which only use 4 bytes or 8 bytes.\n - Reduce 10 bytes to 3 bytes: date with time\n - Reduce 88 bytes to 11 bytes: categorical columns: \n - 177 Numeric Columns - Reduce 1416 bytes to 353 bytes\n- Step 2 - Choose Your File Format\n- Step 3 - Choose Multiple Files or Not\n- Step 4 - Read Raddar's Discussion\n## Look at data: Missing value\nLink discussion: [The distribution of missing values over time](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756)\n**Note**: The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data >90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard\n# Step 2: Feature Engineering: \nEngineering features is key to improving your LB score. We have seen some basic features from public notebook:\n- Aggregated features: [Amex Agg Data How It Created](https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created)\n- Aggregated features + Rapid GPU for faster: [XGBoost Starter - [0.793]](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793)\n- \"After-pay\" features. It makes intuitive semse that subtracting the payments from balance/spend etc provides new information about the users' behavior: [RAPIDS cudf Feature Engineering + XGB](https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb)\n- Statistical features (mean, std, min, max,..): [AMEX LightGBM Quickstart](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart)\n - Selected features averaged over all statements of a customer\n - The minimum or maximum of selected features over all statements of a customer\n - Selected features taken from the last statement of a customer\n\n**Look on**\n1. Features with missing values\n2. Features with low variance\n3. Highly correlated features\n4. Univariate features\n5. Train/Test distributions\n\nBelow are few note on how to engineer new features [Feature Engineering Techniques](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575) by Chris Deotte of Chris Deotte in IEEE-CIS Fraud Detection competition\n\n# Step 3: Feature Selection\nYou can read my discussions about feature selection here with basic feature selection techniques that everyone should know:\n- [Some Feature Selection Technique](https://www.kaggle.com/competitions/optiver-realized-volatility-prediction/discussion/269283)\n- [Few note about XGBoost & Feature Engineering/Feature Selection](https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992)\n- AmbrosM's following discussion is also really helpful: [Which is the right feature importance?](https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131)- In this post he compare the three methods for feature selection and show that the last one is the right one.\n - Split feature importance?\n - Gain feature importance?\n - Permutation feature importance?\n**Conclusion**: Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.\n\n# Step 4: Baseline Model & Hyper-parameters optimization\nThere are many notebooks starting with tree-based model: LightGBM, XGBoost, CatBoost on both CPU and GPU (RAPID) you can refer.\nSome note about tree-based from [AMEX LightGBM Quickstart](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart): Thanks AmbrosM\n\nPreprocessing for LightGBM is much simpler than for neural networks:\n\n1. Neural networks can't process missing values; LightGBM handles them automatically.\n2. Categorical features need to be one-hot encoded for neural networks; LightGBM handles them automatically.\n3. With neural networks, you need to think about outliers; tree-based algorithms deal with outliers easily.\n4. Neural networks need scaled inputs; tree-based algorithms don't depend on scaling.\n\n**Importance**: Validation score with the competition's scoring function\n\nNotebook example for using XGBoost with RAPID & GPU faster: [RAPIDS cudf Feature Engineering + XGB](https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb)\n\n**XGBoost Tip & Trick**:  [Few note about XGBoost & Feature Engineering/Feature Selection](https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992)\n\n## Hyper-parameters optimization\n\nThere are many method for Hyper-parameters optimization like:\n- RandomizedSearchCV [example](https://www.kaggle.com/code/funxexcel/p3-random-forest-tuning-randomizedsearchcv), GridSearchCV [example](https://www.kaggle.com/code/ihelon/titanic-hyperparameter-tuning-with-gridsearchcv)\n- Bayesian Optimization: [example](https://www.kaggle.com/code/vincentlugat/ieee-lgb-bayesian-opt)\n- Optuna: [example](https://www.kaggle.com/code/hamzaghanmi/lgbm-hyperparameter-tuning-using-optuna)\n\n# Step 5: Ensemble Models\n\n- Weighted/Average Ensemble\n- Stacking Model\n\n**[UPDATE...]**",
      "votes": 57
    },
    {
      "id": 1827200,
      "postDate": "2022-06-20T23:19:09.310Z",
      "content": "<p>Thanks for sharing this .. this will help a lot!</p>",
      "rawMarkdown": "Thanks for sharing this .. this will help a lot!"
    },
    {
      "id": 1826257,
      "postDate": "2022-06-20T07:49:51.530Z",
      "content": "<p>Nice summary. The sentence \"Categorical features need to be one-hot encoded for neural networks\" should be softened, NNs can work with any mapping of categorical features e.g. LabelEncoder but one-hot encoding is one of the best performing (because it also works well with ordinal and non ordinal features)</p>",
      "rawMarkdown": "Nice summary. The sentence \"Categorical features need to be one-hot encoded for neural networks\" should be softened, NNs can work with any mapping of categorical features e.g. LabelEncoder but one-hot encoding is one of the best performing (because it also works well with ordinal and non ordinal features)"
    },
    {
      "id": 1827249,
      "postDate": "2022-06-21T00:31:14.637Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1828172,
          "postDate": "2022-06-21T14:18:48.413Z",
          "content": "<p>Okay, i will update soon!</p>",
          "rawMarkdown": "Okay, i will update soon!"
        }
      ]
    },
    {
      "id": 1912435,
      "postDate": "2022-08-24T18:09:17.327Z",
      "content": "<p>Thanks for sharing this</p>",
      "rawMarkdown": "Thanks for sharing this"
    },
    {
      "id": 1858593,
      "postDate": "2022-07-17T03:54:07.247Z",
      "content": "<p>Very helpful. Thank you</p>",
      "rawMarkdown": "Very helpful. Thank you"
    }
  ],
  "comments": [
    {
      "id": 1827200,
      "author_name": "Chirag Desai",
      "author_url": "",
      "post_date": "2022-06-20T23:19:09.310000",
      "content": "<p>Thanks for sharing this .. this will help a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1826257,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-06-20T07:49:51.530000",
      "content": "<p>Nice summary. The sentence \"Categorical features need to be one-hot encoded for neural networks\" should be softened, NNs can work with any mapping of categorical features e.g. LabelEncoder but one-hot encoding is one of the best performing (because it also works well with ordinal and non ordinal features)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1827249,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-21T00:31:14.637000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1828172,
          "author_name": "KhanhVD",
          "author_url": "",
          "post_date": "2022-06-21T14:18:48.413000",
          "content": "<p>Okay, i will update soon!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1912435,
      "author_name": "Gaju Ahmed",
      "author_url": "",
      "post_date": "2022-08-24T18:09:17.327000",
      "content": "<p>Thanks for sharing this</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1858593,
      "author_name": "Little bird",
      "author_url": "",
      "post_date": "2022-07-17T03:54:07.247000",
      "content": "<p>Very helpful. Thank you</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1825176": "I wanted to share a simple pipeline to start a tabular competition for beginners. It's been a long time since Kaggle had a true tabular competition and I wish it was easier for newbies to get started.\n# Step 1: Look at data + Preprocessing data: EDA, Data cleaning, Reduce data size without loss of information,..\nThe following references will be helpful:\n## Magic data format: random uniform noise added to each column\nLink discussion: [Integer columns in the data - here you go!](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)- Raddar have found it: **all float type columns have random uniform noise of [0,0.01] added to each column**\n- originally we had 188 float/categorical type features. These were transformed into\n - 95 np.int8/np.int16 types\n - 93 np.float32 types\n- Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.\n- saved in parquet format (only 1.7GB training data!)\nThis is really helpful when you look at the original data set size (50.31GB), now using Kaggle Notebook is possible and convenient for people with limited resources.\n## Step by Step: How to reduce data size\nLink discussion: [How To Reduce Data Size](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054)- Thanks Chris\nRead the following helpful discussion by Chris if you want a deeper understanding of reduce data size. \n- Step 1 - Reduce Data Types!: \n - convert int32 or int64 which only use 4 bytes or 8 bytes.\n - Reduce 10 bytes to 3 bytes: date with time\n - Reduce 88 bytes to 11 bytes: categorical columns: \n - 177 Numeric Columns - Reduce 1416 bytes to 353 bytes\n- Step 2 - Choose Your File Format\n- Step 3 - Choose Multiple Files or Not\n- Step 4 - Read Raddar's Discussion\n## Look at data: Missing value\nLink discussion: [The distribution of missing values over time](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756)\n**Note**: The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data >90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard\n# Step 2: Feature Engineering: \nEngineering features is key to improving your LB score. We have seen some basic features from public notebook:\n- Aggregated features: [Amex Agg Data How It Created](https://www.kaggle.com/code/huseyincot/amex-agg-data-how-it-created)\n- Aggregated features + Rapid GPU for faster: [XGBoost Starter - [0.793]](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793)\n- \"After-pay\" features. It makes intuitive semse that subtracting the payments from balance/spend etc provides new information about the users' behavior: [RAPIDS cudf Feature Engineering + XGB](https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb)\n- Statistical features (mean, std, min, max,..): [AMEX LightGBM Quickstart](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart)\n - Selected features averaged over all statements of a customer\n - The minimum or maximum of selected features over all statements of a customer\n - Selected features taken from the last statement of a customer\n\n**Look on**\n1. Features with missing values\n2. Features with low variance\n3. Highly correlated features\n4. Univariate features\n5. Train/Test distributions\n\nBelow are few note on how to engineer new features [Feature Engineering Techniques](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/108575) by Chris Deotte of Chris Deotte in IEEE-CIS Fraud Detection competition\n\n# Step 3: Feature Selection\nYou can read my discussions about feature selection here with basic feature selection techniques that everyone should know:\n- [Some Feature Selection Technique](https://www.kaggle.com/competitions/optiver-realized-volatility-prediction/discussion/269283)\n- [Few note about XGBoost & Feature Engineering/Feature Selection](https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992)\n- AmbrosM's following discussion is also really helpful: [Which is the right feature importance?](https://www.kaggle.com/competitions/amex-default-prediction/discussion/331131)- In this post he compare the three methods for feature selection and show that the last one is the right one.\n - Split feature importance?\n - Gain feature importance?\n - Permutation feature importance?\n**Conclusion**: Although permutation feature importance is slower than the gradient boosters' built-in importance measures, it is much more useful because it measures generalization and distinguishes positive from negative importances. Features with negative importance should be dropped from the model.\n\n# Step 4: Baseline Model & Hyper-parameters optimization\nThere are many notebooks starting with tree-based model: LightGBM, XGBoost, CatBoost on both CPU and GPU (RAPID) you can refer.\nSome note about tree-based from [AMEX LightGBM Quickstart](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart): Thanks AmbrosM\n\nPreprocessing for LightGBM is much simpler than for neural networks:\n\n1. Neural networks can't process missing values; LightGBM handles them automatically.\n2. Categorical features need to be one-hot encoded for neural networks; LightGBM handles them automatically.\n3. With neural networks, you need to think about outliers; tree-based algorithms deal with outliers easily.\n4. Neural networks need scaled inputs; tree-based algorithms don't depend on scaling.\n\n**Importance**: Validation score with the competition's scoring function\n\nNotebook example for using XGBoost with RAPID & GPU faster: [RAPIDS cudf Feature Engineering + XGB](https://www.kaggle.com/code/jiweiliu/rapids-cudf-feature-engineering-xgb)\n\n**XGBoost Tip & Trick**:  [Few note about XGBoost & Feature Engineering/Feature Selection](https://www.kaggle.com/competitions/foursquare-location-matching/discussion/321992)\n\n## Hyper-parameters optimization\n\nThere are many method for Hyper-parameters optimization like:\n- RandomizedSearchCV [example](https://www.kaggle.com/code/funxexcel/p3-random-forest-tuning-randomizedsearchcv), GridSearchCV [example](https://www.kaggle.com/code/ihelon/titanic-hyperparameter-tuning-with-gridsearchcv)\n- Bayesian Optimization: [example](https://www.kaggle.com/code/vincentlugat/ieee-lgb-bayesian-opt)\n- Optuna: [example](https://www.kaggle.com/code/hamzaghanmi/lgbm-hyperparameter-tuning-using-optuna)\n\n# Step 5: Ensemble Models\n\n- Weighted/Average Ensemble\n- Stacking Model\n\n**[UPDATE...]**",
    "1827200": "Thanks for sharing this .. this will help a lot!",
    "1826257": "Nice summary. The sentence \"Categorical features need to be one-hot encoded for neural networks\" should be softened, NNs can work with any mapping of categorical features e.g. LabelEncoder but one-hot encoding is one of the best performing (because it also works well with ordinal and non ordinal features)",
    "1827249": "",
    "1912435": "Thanks for sharing this",
    "1858593": "Very helpful. Thank you"
  }
}