{
  "id": 335892,
  "title": "Tabular Classification - Tips and Tricks",
  "url": "/competitions/amex-default-prediction/discussion/335892",
  "author_name": "The Devastator",
  "post_date": "2022-07-08T10:19:39.431000",
  "votes": 358,
  "comment_count": 41,
  "views": 0,
  "content": "<h1>Tabular Classification - Tips and Tricks</h1>\n<p>Sharing a long list of tips and tricks found on previous tabular classification competitions on Kaggle.<br>\nTake this as a list of ideas you can attempt to try for improving your score!</p>\n<blockquote>\n  <p><strong>Credits:</strong> </p>\n  <ul>\n  <li>Many of the links are based on parts of <a href=\"https://neptune.ai/blog/tabular-data-binary-classification-tips-and-tricks-from-5-kaggle-competitions\" target=\"_blank\">this</a> blog post (with edits and filtering).</li>\n  <li>Including links from: <a href=\"https://www.kaggle.com/discussions/getting-started/277121\" target=\"_blank\">The best of kaggle - GBMs</a>.</li>\n  <li>Also from: <a href=\"https://www.kaggle.com/discussions/getting-started/291437\" target=\"_blank\">The best of kaggle - Ensembles</a>.</li>\n  </ul>\n</blockquote>\n<h3>Dealing with larger datasets</h3>\n<p>An issue close to our hearts! </p>\n<ul>\n<li>Faster <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/59575\" target=\"_blank\">data loading with pandas.</a></li>\n<li>Data compression techniques to <a href=\"https://www.kaggle.com/nickycan/compress-70-of-dataset\" target=\"_blank\">reduce the size of data by 70%</a>.</li>\n<li>Optimize the memory by r<a href=\"https://www.kaggle.com/shrutimechlearn/large-data-loading-trick-with-ms-malware-data\" target=\"_blank\">educing the size of some attributes.</a></li>\n<li>Use open-source libraries such as <a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\" target=\"_blank\">Dask to read and manipulate the data</a>, it performs parallel computing and saves up memory space.</li>\n<li>Use <a href=\"https://github.com/rapidsai/cudf\" target=\"_blank\">cudf</a>.</li>\n<li>Convert data to <a href=\"https://arrow.apache.org/docs/python/parquet.html\" target=\"_blank\">parquet</a> format.</li>\n<li>Converting data to <a href=\"https://medium.com/@snehotosh.banerjee/feather-a-fast-on-disk-format-for-r-and-python-data-frames-de33d0516b03\" target=\"_blank\">feather</a> format.</li>\n<li>Reducing memory usage for <a href=\"https://www.kaggle.com/mjbahmani/reducing-memory-size-for-ieee\" target=\"_blank\">optimizing RAM</a>.</li>\n</ul>\n<h3>Data exploration</h3>\n<p>Data exploration is a must on all competitions, Here are some references from past competitions use them to get some ideas on what exactly to explore.</p>\n<ul>\n<li>EDA for microsoft <a href=\"https://www.kaggle.com/youhanlee/my-eda-i-want-to-see-all\" target=\"_blank\">malware detection.</a></li>\n<li>Time Series <a href=\"https://www.kaggle.com/cdeotte/time-split-validation-malware-0-68\" target=\"_blank\">EDA for malware detection.</a></li>\n<li>Complete <a href=\"https://www.kaggle.com/codename007/home-credit-complete-eda-feature-importance\" target=\"_blank\">EDA for home credit loan prediction</a>.</li>\n<li>Complete <a href=\"https://www.kaggle.com/gpreda/santander-eda-and-prediction\" target=\"_blank\">EDA for Santader prediction.</a></li>\n<li>EDA for <a href=\"https://www.kaggle.com/go1dfish/basic-eda\" target=\"_blank\">VSB Power Line Fault Detection.</a></li>\n</ul>\n<h3>Preprocessing</h3>\n<p>Before we start to engineer features, we need to do some preprocessing.<br>\nThose invlolve removing missing values, converting categorical variables to numeric, and scaling the data, etc..<br>\nHere are some references from past competitions use them to get some ideas for preprocessing.</p>\n<ul>\n<li>Methods to <a href=\"https://www.kaggle.com/shahules/tackling-class-imbalance\" target=\"_blank\">tackle class imbalance</a>.</li>\n<li>Data augmentation by <a href=\"https://machinelearningmastery.com/smote-oversampling-for-imbalanced-classification/\" target=\"_blank\">Synthetic Minority Oversampling Technique</a>.</li>\n<li>Fast inplace <a href=\"https://www.kaggle.com/jiweiliu/fast-inplace-shuffle-for-augmentation\" target=\"_blank\">shuffle for augmentation</a>.</li>\n<li>Finding <a href=\"https://www.kaggle.com/yag320/list-of-fake-samples-and-public-private-lb-split\" target=\"_blank\">synthetic samples in the dataset.</a></li>\n<li><a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\" target=\"_blank\">Signal denoising</a> used in signal processing competitions.</li>\n<li>Finding <a href=\"https://www.kaggle.com/jpmiller/patterns-of-missing-data\" target=\"_blank\">patterns of missing data</a>.</li>\n<li>Methods to handle <a href=\"https://towardsdatascience.com/6-different-ways-to-compensate-for-missing-values-data-imputation-with-examples-6022d9ca0779\" target=\"_blank\">missing data</a>.</li>\n<li>An overview of various <a href=\"https://www.kaggle.com/shahules/an-overview-of-encoding-techniques\" target=\"_blank\">encoding techniques for categorical data.</a></li>\n<li>Building <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64598\" target=\"_blank\">model to predict missing values.</a></li>\n<li>Random <a href=\"https://www.kaggle.com/brandenkmurray/randomly-shuffled-data-also-works\" target=\"_blank\">shuffling of data</a> to create new synthetic training set.</li>\n</ul>\n<h3>Feature engineering</h3>\n<p>Now we can start to engineer features, below you can find some references from past competitions use them to get some ideas to engineer features. </p>\n<ul>\n<li>Target <a href=\"https://medium.com/@pouryaayria/k-fold-target-encoding-dfe9a594874b/\" target=\"_blank\">encoding cross validation</a> for better encoding.</li>\n<li>Entity embedding to <a href=\"https://www.kaggle.com/abhishek/entity-embeddings-to-handle-categories\" target=\"_blank\">handle categories</a>.</li>\n<li>Encoding c<a href=\"https://www.kaggle.com/avanwyk/encoding-cyclical-features-for-deep-learning\" target=\"_blank\">yclic features for deep learning.</a></li>\n<li>Manual <a href=\"https://www.kaggle.com/willkoehrsen/introduction-to-manual-feature-engineering\" target=\"_blank\">feature engineering methods</a>.</li>\n<li>Automated feature engineering techniques <a href=\"https://www.kaggle.com/willkoehrsen/automated-feature-engineering-basics\" target=\"_blank\">using featuretools</a>.</li>\n<li>Top hard crafted features used in <a href=\"https://www.kaggle.com/sanderf/7th-place-solution-microsoft-malware-prediction\" target=\"_blank\">microsoft malware detection</a>.</li>\n<li>Denoising NN for <a href=\"https://towardsdatascience.com/applied-deep-learning-part-3-autoencoders-1c083af4d798\" target=\"_blank\">feature extraction</a>.</li>\n<li>Feature engineering <a href=\"https://www.kaggle.com/cdeotte/rapids-feature-engineering-fraud-0-96/\" target=\"_blank\">using RAPIDS framework.</a></li>\n<li>Things to remember while processing f<a href=\"https://www.kaggle.com/c/ieee-fraud-detection/discussion/108575\" target=\"_blank\">eatures using LGBM.</a></li>\n<li><a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64593\" target=\"_blank\">Lag features and moving averages.</a></li>\n<li><a href=\"https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567\" target=\"_blank\">Principal component analysis</a> for dimensionality reduction.</li>\n<li>LDA for <a href=\"https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567\" target=\"_blank\">dimensionality reduction</a>.</li>\n<li>Best hand crafted LGBM features for <a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/85157\" target=\"_blank\">microsoft malware detection</a>.</li>\n<li>Generating <a href=\"https://www.kaggle.com/philippsinger/frequency-features-without-test-data-information\" target=\"_blank\">frequency features.</a></li>\n<li>Dropping variables with <a href=\"https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix\" target=\"_blank\">different train and test distribution.</a></li>\n<li><a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64593\" target=\"_blank\">Aggregate time series features</a> for home credit competition.</li>\n<li><a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64593\" target=\"_blank\">Time Series</a> features used in home credit default risk.</li>\n<li>Scale, Standardize and n<a href=\"https://towardsdatascience.com/scale-standardize-or-normalize-with-scikit-learn-6ccc7d176a02\" target=\"_blank\">ormalize with sklearn</a>.</li>\n<li>Handcrafted features for <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/57750\" target=\"_blank\">Home default risk competition.</a></li>\n<li>Handcrafted <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/89070\" target=\"_blank\">features used in Santander Transaction Prediction.</a></li>\n</ul>\n<h3>Feature selection</h3>\n<p>Now that we got our features engineered, we should check if they are useful or not.<br>\nTo do this, we need to do some feature selection.<br>\nHere are some references from past competitions use them to get some ideas for feature selection.</p>\n<ul>\n<li>Six ways to do <a href=\"https://www.kaggle.com/sz8416/6-ways-for-feature-selection\" target=\"_blank\">features selection using sklearn</a>.</li>\n<li><a href=\"https://www.kaggle.com/c/ieee-fraud-detection/discussion/107877#latest-635386\" target=\"_blank\">Permutation feature importance</a>.</li>\n<li><a href=\"https://www.kaggle.com/tunguz/adversarial-ieee/\" target=\"_blank\">Adversarial feature validation</a>.</li>\n<li>Feature selection using <a href=\"https://www.kaggle.com/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">null importances.</a></li>\n<li>Tree explainer using <a href=\"https://github.com/slundberg/shap\" target=\"_blank\">SHAP.</a></li>\n<li>DeepNN explainer using <a href=\"https://github.com/slundberg/shap\" target=\"_blank\">SHAP</a>.</li>\n</ul>\n<h3>Modeling</h3>\n<p>Below you can find <a href=\"https://www.kaggle.com/discussions/getting-started/277121\" target=\"_blank\">The best of kaggle - GBMs</a>.<br>\nThese are the top voted/viewed notebooks all over kaggle for GBMs.  <br>\n<strong>Usage rules are simple: You like it? Upvote the original!</strong></p>\n<h3>XGBoost</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/anokas/data-analysis-xgboost-starter-0-35460-lb\" target=\"_blank\">Data Analysis &amp; XGBoost Starter (0.35460 LB)</a> (137K views, 1377 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/erikbruin/house-prices-lasso-xgboost-and-a-detailed-eda\" target=\"_blank\">House prices: Lasso, XGBoost, and a detailed EDA</a> (183K views, 1317 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dlarionov/feature-engineering-xgboost\" target=\"_blank\">Feature engineering, xgboost</a> (144K views, 942 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dansbecker/xgboost\" target=\"_blank\">XGBoost</a> (173K views, 918 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/alexisbcook/xgboost\" target=\"_blank\">XGBoost</a> (266K views, 578 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hamditarek/market-prediction-xgboost-with-gpu-fit-in-1min\" target=\"_blank\">Market Prediction: XGBoost with GPU (Fit in 1min)</a> (53K views, 411 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/a-guide-on-xgboost-hyperparameters-tuning\" target=\"_blank\">A Guide on XGBoost hyperparameters tuning</a> (55K views, 358 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/titanic-wcg-xgboost-0-84688\" target=\"_blank\">Titanic WCG+XGBoost [0.84688]</a> (27K views, 314 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/shahules/xgboost-feature-selection-dsbowl\" target=\"_blank\">XGBoost &amp; Feature Selection DSBowl 🥣 🥣</a> (36K views, 299 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tqchen/understanding-xgboost-model-on-otto-data\" target=\"_blank\">Understanding XGBoost Model on Otto Data</a> (130K views, 297 Votes)*</li>\n</ul>\n<h3>LightGBM</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/jsaguiar/lightgbm-with-simple-features\" target=\"_blank\">LightGBM with Simple Features</a> (85K views, 619 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix\" target=\"_blank\">LightGBM. Baseline Model Using Sparse Matrix</a> (49K views, 409 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tommy1028/lightgbm-starter-with-feature-engineering-idea\" target=\"_blank\">LightGBM starter with feature engineering idea</a> (16K views, 370 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/erikbruin/google-analytics-eda-lightgbm-screenshots\" target=\"_blank\">Google Analytics EDA + LightGBM + Screenshots</a> (40K views, 344 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/davidcairuz/feature-engineering-lightgbm\" target=\"_blank\">Feature Engineering &amp; LightGBM</a> (26K views, 272 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending\" target=\"_blank\">Simple LightGBM without blending</a> (18K views, 263 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/samratp/lightgbm-xgboost-catboost\" target=\"_blank\">LightGBM + XGBoost + Catboost</a> (37K views, 245 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/aitude/ashrae-kfold-lightgbm-without-leak-1-08\" target=\"_blank\">ASHRAE- KFold LightGBM - without leak (1.08)</a> (15K views, 240 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/lightgbm-classifier-in-python\" target=\"_blank\">LightGBM Classifier in Python</a> (52K views, 231 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-lb-0-9680\" target=\"_blank\">LightGBM (Fixing unbalanced data)</a> (56K views, 228 Votes)*</li>\n</ul>\n<h3>CatBoost</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/mhviraf/a-new-baseline-for-dsb-2019-catboost-model\" target=\"_blank\">A new baseline for DSB 2019 - Catboost model</a> (15K views, 316 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/samratp/lightgbm-xgboost-catboost\" target=\"_blank\">LightGBM + XGBoost + Catboost</a> (37K views, 245 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/abhinand05/catboost-a-deeper-dive\" target=\"_blank\">CatBoost: A Deeper Dive</a> (8K views, 178 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/allunia/house-prices-tutorial-with-catboost\" target=\"_blank\">House Prices Tutorial with Catboost</a> (12K views, 176 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm\" target=\"_blank\">Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM</a> (25K views, 173 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/usharengaraju/widsdatathon2021-catboost-starter\" target=\"_blank\">WiDSDatathon2021-Catboost-Starter</a> (6K views, 167 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mitribunskiy/tutorial-catboost-overview\" target=\"_blank\">Tutorial: CatBoost Overview</a> (39K views, 127 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/catboost-classifier-in-python\" target=\"_blank\">CatBoost Classifier in Python</a> (28K views, 126 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/braquino/catboost-some-more-features\" target=\"_blank\">Catboost - Some more features</a> (7K views, 103 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/mlcourse-ai-fall-2019-catboost-starter\" target=\"_blank\">mlcourse.ai. Fall 2019. Catboost starter</a> (18K views, 98 Votes)*</li>\n</ul>\n<h3>Ensemble</h3>\n<p>Now that we have our model trained, we should ensemble our predictions.</p>\n<p>Below you can find <a href=\"https://www.kaggle.com/discussions/getting-started/291437\" target=\"_blank\">The best of kaggle - Ensembles</a>.<br>\nThese are the top voted/viewed notebooks all over kaggle for GBMs.<br>\n<strong>Usage rules are simple: You like it? Upvote the original!</strong></p>\n<h3>Ensembling Techniques</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\" target=\"_blank\">Introduction to Ensembling/Stacking in Python</a> (620K views, 5436 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/yassineghouzam/titanic-top-4-with-ensemble-modeling\" target=\"_blank\">Titanic Top 4% with ensemble modeling</a> (185K views, 2501 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/viveksrinivasan/eda-ensemble-model-top-10-percentile\" target=\"_blank\">EDA &amp; Ensemble Model (Top 10 Percentile)</a> (98K views, 513 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble\" target=\"_blank\">Analysis of Melanoma Metadata and EffNet Ensemble</a> (24K views, 405 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tannercarbonati/detailed-data-analysis-ensemble-modeling\" target=\"_blank\">Detailed Data Analysis &amp; Ensemble Modeling</a> (55K views, 360 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/ensemble-folds-with-median-0-153\" target=\"_blank\">Ensemble Folds with MEDIAN - [0.153]</a> (9K views, 303 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/shonenkov/wbf-approach-for-ensemble\" target=\"_blank\">WBF approach for ensemble</a> (18K views, 282 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903\" target=\"_blank\">Ensemble: Resnext50_32x4d + Efficientnet = 0.903</a> (16K views, 274 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/arthurtok/employee-attrition-via-ensemble-tree-based-methods\" target=\"_blank\">Employee attrition via Ensemble tree-based methods</a> (64K views, 253 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/isaienkov/top-3-efficient-ensembling-in-few-lines-of-code\" target=\"_blank\">Top 3%. Efficient ensembling in few lines of code</a> (30K views, 242 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/pavansanagapati/ensemble-learning-techniques-tutorial\" target=\"_blank\">Ensemble Learning Techniques Tutorial</a> (55K views, 215 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jhoward/minimal-lstm-nb-svm-baseline-ensemble\" target=\"_blank\">Minimal LSTM + NB-SVM baseline ensemble</a> (30K views, 209 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example\" target=\"_blank\">Ensemble Model: Stacked Model Example</a> (68K views, 207 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hamditarek/ensemble\" target=\"_blank\">🤗🤗 Ensemble 🤗🤗</a> (64K views, 196 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble\" target=\"_blank\">EDA&amp;Modelling of the External Data Inc. Ensemble</a> (8K views, 195 Votes)*</li>\n</ul>\n<h3>Blending</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending\" target=\"_blank\">Simple LightGBM without blending</a> (18K views, 265 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/a763337092/blending-tensorflow-and-pytorch\" target=\"_blank\">Blending tensorflow and pytorch🔥🔥🔥</a> (13K views, 220 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/speedwagon/quadratic-discriminant-analysis\" target=\"_blank\">Another model for your blending</a> (11K views, 216 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/gogo827jz/optimise-blending-weights-with-bonus-0\" target=\"_blank\">Optimise Blending Weights with Bonus :0</a> (7K views, 184 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/suicaokhoailang/blending-with-linear-regression-0-688-lb\" target=\"_blank\">Blending with Linear Regression [0.688 LB]</a> (14K views, 136 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/abhishek/blending-blending-blending\" target=\"_blank\">blending blending blending</a> (3K views, 128 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jagangupta/lessons-from-toxic-blending-is-the-new-sexy\" target=\"_blank\">Lessons from Toxic : Blending is the new sexy</a> (10K views, 128 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\" target=\"_blank\">non-blending lightGBM model LB: 0.977</a> (16K views, 127 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hengzheng/bayesian-optimization-seed-blending\" target=\"_blank\">Bayesian Optimization Seed Blending</a> (9K views, 118 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/gpreda/elo-world-high-score-without-blending\" target=\"_blank\">elo_world_high_score_without_blending</a> (8K views, 112 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/abhishek/competition-part-5-blending-101\" target=\"_blank\">competition part-5: blending 101</a> (3K views, 105 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/beginner-eda-with-feature-eng-and-blending-models\" target=\"_blank\">Beginner EDA with Feature Eng. and Blending Models</a> (3K views, 99 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors\" target=\"_blank\">Cross-validation, weighted linear blending, errors</a> (7K views, 84 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/sagarjiyani/blending-tensorflow-pytorch-th-0-4914\" target=\"_blank\">Blending TensorFlow + PyTorch (th=0.4914)</a> (5K views, 82 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/gogo827jz/blending-nn-and-lgbm-rf\" target=\"_blank\">Blending NN and LGBM/RF</a> (6K views, 81 Votes)*</li>\n</ul>\n<h3>Stacking</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard\" target=\"_blank\">Stacked Regressions : Top 4% on LeaderBoard</a> (493K views, 6236 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\" target=\"_blank\">Introduction to Ensembling/Stacking in Python</a> (620K views, 5436 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dimitreoliveira/model-stacking-feature-engineering-and-eda\" target=\"_blank\">Model stacking, feature engineering and EDA</a> (43K views, 409 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/agodwinp/stacking-house-prices-walkthrough-to-top-5\" target=\"_blank\">Stacking House Prices - Walkthrough to Top 5%</a> (32K views, 227 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mmueller/stacking-starter\" target=\"_blank\">Stacking Starter</a> (41K views, 221 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example\" target=\"_blank\">Ensemble Model: Stacked Model Example</a> (68K views, 207 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/nicapotato/titanic-voting-pipeline-stack-and-guide\" target=\"_blank\">Titanic: Voting, Pipeline, Stack, and Guide</a> (21K views, 197 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dongxu027/explore-stacking-lb-0-1463\" target=\"_blank\">Explore Stacking (LB 0.1463)</a> (18K views, 180 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm\" target=\"_blank\">Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM</a> (26K views, 176 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/top-1-approach-eda-new-models-and-stacking\" target=\"_blank\">Top 1% Approach: EDA, New Models and Stacking</a> (9K views, 172 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/yekenot/simple-stacker-lb-0-284\" target=\"_blank\">Simple Stacker LB 0.284</a> (20K views, 172 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/itslek/stack-blend-lrs-xgb-lgb-house-prices-k-v17\" target=\"_blank\">Stack&amp;Blend LRs XGB LGB {House Prices K} v17</a> (6K views, 164 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/agehsbarg/top-10-0-10943-stacking-mice-and-brutal-force\" target=\"_blank\">Top 10 (0.10943): stacking, MICE and brutal force</a> (21K views, 156 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hhstrand/oof-stacking-regime\" target=\"_blank\">OOF stacking regime</a> (14K views, 145 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tunguz/eloda-with-feature-engineering-and-stacking\" target=\"_blank\">EloDA with Feature Engineering and Stacking</a> (10K views, 141 Votes)*</li>\n</ul>\n<h3>Bagging</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/ammarnassanalhajali/riiid-lgbm-bagging2-sakt-0-781\" target=\"_blank\">Riiid LGBM bagging2 + SAKT =0.781</a> (11K views, 188 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/leadbest/sakt-riiid-lgbm-bagging2\" target=\"_blank\">SAKT + Riiid LGBM bagging2</a> (8K views, 139 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/topic-5-ensembles-part-1-bagging\" target=\"_blank\">Topic 5. Ensembles. Part 1. Bagging</a> (16K views, 137 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/bagging-vs-boosting\" target=\"_blank\">Bagging vs Boosting</a> (22K views, 117 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/julianguo/fork-of-riiid-lgbm-bagging2-1-471152\" target=\"_blank\">Fork of Riiid LGBM bagging2.1 471152</a> (8K views, 117 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2\" target=\"_blank\">Riiid! LGBM bagging2</a> (6K views, 113 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kabure/predicting-house-prices-xgb-rf-bagging-reg-pipe\" target=\"_blank\">Predicting House Prices [XGB/RF/Bagging-Reg Pipe]</a> (10K views, 89 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/danijelk/keras-starter-with-bagging-lb-1120-596\" target=\"_blank\">Keras starter with bagging (LB: 1120.596)</a> (18K views, 77 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mtinti/keras-starter-with-bagging-1111-84364\" target=\"_blank\">Keras starter with bagging 1111.84364</a> (19K views, 77 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2-1\" target=\"_blank\">Riiid LGBM bagging2.1</a> (4K views, 62 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/andypenrose/baggingregressor-rapids-ensemble\" target=\"_blank\">BaggingRegressor + RAPIDS Ensemble</a> (4K views, 59 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/fahadmehfoooz/credit-analysis-with-knn-dtree-rf-bagging-ann\" target=\"_blank\">Credit analysis with KNN/DTree/RF/Bagging/ANN🏦</a> (&lt;1K views, 55 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm\" target=\"_blank\">Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM</a> (1K views, 51 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting\" target=\"_blank\">Ensemble ML Algorithms : Bagging, Boosting, Voting</a> (8K views, 48 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/cttsai/simple-xgboost-cv-bagging\" target=\"_blank\">Simple XGBoost CV Bagging</a> (3K views, 42 Votes)*</li>\n</ul>\n<h3>Boosting</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/felipemello/boosting-creativity-towards-feature-engineering\" target=\"_blank\">Boosting creativity towards feature engineering</a> (9K views, 229 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/yashvi/vehicle-insurance-eda-and-boosting-models\" target=\"_blank\">Vehicle Insurance EDA and boosting models</a> (19K views, 200 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/topic-10-gradient-boosting\" target=\"_blank\">Topic 10. Gradient Boosting</a> (21K views, 188 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/grroverpr/gradient-boosting-simplified\" target=\"_blank\">Gradient boosting simplified</a> (48K views, 147 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/bagging-vs-boosting\" target=\"_blank\">Bagging vs Boosting</a> (22K views, 117 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/niteshyadav3103/diabetes-prediction-stacking-boosting\" target=\"_blank\">Diabetes Prediction (Stacking + Boosting)</a> (1K views, 85 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/roydatascience/elo-stack-with-goss-boosting\" target=\"_blank\">Elo Stack With Goss Boosting</a> (6K views, 78 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/assignment-10-gradient-boosting-and-flight-delays\" target=\"_blank\">Assignment 10. Gradient boosting and flight delays</a> (9K views, 71 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/lavanyashukla01/battle-of-the-boosting-algos-lgb-xgb-catboost\" target=\"_blank\">Battle of the Boosting Algos: LGB, XGB, Catboost</a> (9K views, 55 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/vinnsvinay/introduction-to-boosting-using-lgbm-lb-0-68357\" target=\"_blank\">Introduction to Boosting using LGBM(LB: 0.68357)</a> (14K views, 52 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tunguz/fe-pipeline-with-histgradientboostingregressor\" target=\"_blank\">FE-Pipeline with HistGradientBoostingRegressor</a> (3K views, 51 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm\" target=\"_blank\">Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM</a> (1K views, 51 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/parulpandey/explainable-boosting-machines-for-tabular-data\" target=\"_blank\">Explainable Boosting machines for Tabular data</a> (2K views, 49 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting\" target=\"_blank\">Ensemble ML Algorithms : Bagging, Boosting, Voting</a> (8K views, 48 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/sid321axn/house-price-prediction-gboosting-adaboost-etc\" target=\"_blank\">House Price Prediction : GBoosting,AdaBoost etc.</a> (8K views, 42 Votes)*</li>\n</ul>\n<h3>Hyperparameters Tuning</h3>\n<ul>\n<li>LGBM <a href=\"https://www.kaggle.com/mlisovyi/lightgbm-hyperparameter-optimisation-lb-0-761\" target=\"_blank\">hyperparameter tuning</a> methods.</li>\n<li>Automated <a href=\"https://www.kaggle.com/willkoehrsen/automated-model-tuning\" target=\"_blank\">model tuning</a> methods.</li>\n<li>Parameter tuning with <a href=\"https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt\" target=\"_blank\">hyper plot</a>.</li>\n<li><a href=\"http://krasserm.github.io/2018/03/21/bayesian-optimization/\" target=\"_blank\">Bayesian optimization</a> for hyperparameter tuning.</li>\n<li><a href=\"https://www.kaggle.com/nicapotato/gpyopt-hyperparameter-optimisation-gpu-lgbm\" target=\"_blank\">Gpyopt Hyperparameter Optimisation</a>.</li>\n</ul>",
  "messages": [
    {
      "id": 1848063,
      "postDate": "2022-07-08T10:19:39.430Z",
      "content": "<h1>Tabular Classification - Tips and Tricks</h1>\n<p>Sharing a long list of tips and tricks found on previous tabular classification competitions on Kaggle.<br>\nTake this as a list of ideas you can attempt to try for improving your score!</p>\n<blockquote>\n  <p><strong>Credits:</strong> </p>\n  <ul>\n  <li>Many of the links are based on parts of <a href=\"https://neptune.ai/blog/tabular-data-binary-classification-tips-and-tricks-from-5-kaggle-competitions\" target=\"_blank\">this</a> blog post (with edits and filtering).</li>\n  <li>Including links from: <a href=\"https://www.kaggle.com/discussions/getting-started/277121\" target=\"_blank\">The best of kaggle - GBMs</a>.</li>\n  <li>Also from: <a href=\"https://www.kaggle.com/discussions/getting-started/291437\" target=\"_blank\">The best of kaggle - Ensembles</a>.</li>\n  </ul>\n</blockquote>\n<h3>Dealing with larger datasets</h3>\n<p>An issue close to our hearts! </p>\n<ul>\n<li>Faster <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/59575\" target=\"_blank\">data loading with pandas.</a></li>\n<li>Data compression techniques to <a href=\"https://www.kaggle.com/nickycan/compress-70-of-dataset\" target=\"_blank\">reduce the size of data by 70%</a>.</li>\n<li>Optimize the memory by r<a href=\"https://www.kaggle.com/shrutimechlearn/large-data-loading-trick-with-ms-malware-data\" target=\"_blank\">educing the size of some attributes.</a></li>\n<li>Use open-source libraries such as <a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\" target=\"_blank\">Dask to read and manipulate the data</a>, it performs parallel computing and saves up memory space.</li>\n<li>Use <a href=\"https://github.com/rapidsai/cudf\" target=\"_blank\">cudf</a>.</li>\n<li>Convert data to <a href=\"https://arrow.apache.org/docs/python/parquet.html\" target=\"_blank\">parquet</a> format.</li>\n<li>Converting data to <a href=\"https://medium.com/@snehotosh.banerjee/feather-a-fast-on-disk-format-for-r-and-python-data-frames-de33d0516b03\" target=\"_blank\">feather</a> format.</li>\n<li>Reducing memory usage for <a href=\"https://www.kaggle.com/mjbahmani/reducing-memory-size-for-ieee\" target=\"_blank\">optimizing RAM</a>.</li>\n</ul>\n<h3>Data exploration</h3>\n<p>Data exploration is a must on all competitions, Here are some references from past competitions use them to get some ideas on what exactly to explore.</p>\n<ul>\n<li>EDA for microsoft <a href=\"https://www.kaggle.com/youhanlee/my-eda-i-want-to-see-all\" target=\"_blank\">malware detection.</a></li>\n<li>Time Series <a href=\"https://www.kaggle.com/cdeotte/time-split-validation-malware-0-68\" target=\"_blank\">EDA for malware detection.</a></li>\n<li>Complete <a href=\"https://www.kaggle.com/codename007/home-credit-complete-eda-feature-importance\" target=\"_blank\">EDA for home credit loan prediction</a>.</li>\n<li>Complete <a href=\"https://www.kaggle.com/gpreda/santander-eda-and-prediction\" target=\"_blank\">EDA for Santader prediction.</a></li>\n<li>EDA for <a href=\"https://www.kaggle.com/go1dfish/basic-eda\" target=\"_blank\">VSB Power Line Fault Detection.</a></li>\n</ul>\n<h3>Preprocessing</h3>\n<p>Before we start to engineer features, we need to do some preprocessing.<br>\nThose invlolve removing missing values, converting categorical variables to numeric, and scaling the data, etc..<br>\nHere are some references from past competitions use them to get some ideas for preprocessing.</p>\n<ul>\n<li>Methods to <a href=\"https://www.kaggle.com/shahules/tackling-class-imbalance\" target=\"_blank\">tackle class imbalance</a>.</li>\n<li>Data augmentation by <a href=\"https://machinelearningmastery.com/smote-oversampling-for-imbalanced-classification/\" target=\"_blank\">Synthetic Minority Oversampling Technique</a>.</li>\n<li>Fast inplace <a href=\"https://www.kaggle.com/jiweiliu/fast-inplace-shuffle-for-augmentation\" target=\"_blank\">shuffle for augmentation</a>.</li>\n<li>Finding <a href=\"https://www.kaggle.com/yag320/list-of-fake-samples-and-public-private-lb-split\" target=\"_blank\">synthetic samples in the dataset.</a></li>\n<li><a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\" target=\"_blank\">Signal denoising</a> used in signal processing competitions.</li>\n<li>Finding <a href=\"https://www.kaggle.com/jpmiller/patterns-of-missing-data\" target=\"_blank\">patterns of missing data</a>.</li>\n<li>Methods to handle <a href=\"https://towardsdatascience.com/6-different-ways-to-compensate-for-missing-values-data-imputation-with-examples-6022d9ca0779\" target=\"_blank\">missing data</a>.</li>\n<li>An overview of various <a href=\"https://www.kaggle.com/shahules/an-overview-of-encoding-techniques\" target=\"_blank\">encoding techniques for categorical data.</a></li>\n<li>Building <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64598\" target=\"_blank\">model to predict missing values.</a></li>\n<li>Random <a href=\"https://www.kaggle.com/brandenkmurray/randomly-shuffled-data-also-works\" target=\"_blank\">shuffling of data</a> to create new synthetic training set.</li>\n</ul>\n<h3>Feature engineering</h3>\n<p>Now we can start to engineer features, below you can find some references from past competitions use them to get some ideas to engineer features. </p>\n<ul>\n<li>Target <a href=\"https://medium.com/@pouryaayria/k-fold-target-encoding-dfe9a594874b/\" target=\"_blank\">encoding cross validation</a> for better encoding.</li>\n<li>Entity embedding to <a href=\"https://www.kaggle.com/abhishek/entity-embeddings-to-handle-categories\" target=\"_blank\">handle categories</a>.</li>\n<li>Encoding c<a href=\"https://www.kaggle.com/avanwyk/encoding-cyclical-features-for-deep-learning\" target=\"_blank\">yclic features for deep learning.</a></li>\n<li>Manual <a href=\"https://www.kaggle.com/willkoehrsen/introduction-to-manual-feature-engineering\" target=\"_blank\">feature engineering methods</a>.</li>\n<li>Automated feature engineering techniques <a href=\"https://www.kaggle.com/willkoehrsen/automated-feature-engineering-basics\" target=\"_blank\">using featuretools</a>.</li>\n<li>Top hard crafted features used in <a href=\"https://www.kaggle.com/sanderf/7th-place-solution-microsoft-malware-prediction\" target=\"_blank\">microsoft malware detection</a>.</li>\n<li>Denoising NN for <a href=\"https://towardsdatascience.com/applied-deep-learning-part-3-autoencoders-1c083af4d798\" target=\"_blank\">feature extraction</a>.</li>\n<li>Feature engineering <a href=\"https://www.kaggle.com/cdeotte/rapids-feature-engineering-fraud-0-96/\" target=\"_blank\">using RAPIDS framework.</a></li>\n<li>Things to remember while processing f<a href=\"https://www.kaggle.com/c/ieee-fraud-detection/discussion/108575\" target=\"_blank\">eatures using LGBM.</a></li>\n<li><a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64593\" target=\"_blank\">Lag features and moving averages.</a></li>\n<li><a href=\"https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567\" target=\"_blank\">Principal component analysis</a> for dimensionality reduction.</li>\n<li>LDA for <a href=\"https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567\" target=\"_blank\">dimensionality reduction</a>.</li>\n<li>Best hand crafted LGBM features for <a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/85157\" target=\"_blank\">microsoft malware detection</a>.</li>\n<li>Generating <a href=\"https://www.kaggle.com/philippsinger/frequency-features-without-test-data-information\" target=\"_blank\">frequency features.</a></li>\n<li>Dropping variables with <a href=\"https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix\" target=\"_blank\">different train and test distribution.</a></li>\n<li><a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64593\" target=\"_blank\">Aggregate time series features</a> for home credit competition.</li>\n<li><a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/64593\" target=\"_blank\">Time Series</a> features used in home credit default risk.</li>\n<li>Scale, Standardize and n<a href=\"https://towardsdatascience.com/scale-standardize-or-normalize-with-scikit-learn-6ccc7d176a02\" target=\"_blank\">ormalize with sklearn</a>.</li>\n<li>Handcrafted features for <a href=\"https://www.kaggle.com/c/home-credit-default-risk/discussion/57750\" target=\"_blank\">Home default risk competition.</a></li>\n<li>Handcrafted <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/89070\" target=\"_blank\">features used in Santander Transaction Prediction.</a></li>\n</ul>\n<h3>Feature selection</h3>\n<p>Now that we got our features engineered, we should check if they are useful or not.<br>\nTo do this, we need to do some feature selection.<br>\nHere are some references from past competitions use them to get some ideas for feature selection.</p>\n<ul>\n<li>Six ways to do <a href=\"https://www.kaggle.com/sz8416/6-ways-for-feature-selection\" target=\"_blank\">features selection using sklearn</a>.</li>\n<li><a href=\"https://www.kaggle.com/c/ieee-fraud-detection/discussion/107877#latest-635386\" target=\"_blank\">Permutation feature importance</a>.</li>\n<li><a href=\"https://www.kaggle.com/tunguz/adversarial-ieee/\" target=\"_blank\">Adversarial feature validation</a>.</li>\n<li>Feature selection using <a href=\"https://www.kaggle.com/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">null importances.</a></li>\n<li>Tree explainer using <a href=\"https://github.com/slundberg/shap\" target=\"_blank\">SHAP.</a></li>\n<li>DeepNN explainer using <a href=\"https://github.com/slundberg/shap\" target=\"_blank\">SHAP</a>.</li>\n</ul>\n<h3>Modeling</h3>\n<p>Below you can find <a href=\"https://www.kaggle.com/discussions/getting-started/277121\" target=\"_blank\">The best of kaggle - GBMs</a>.<br>\nThese are the top voted/viewed notebooks all over kaggle for GBMs.  <br>\n<strong>Usage rules are simple: You like it? Upvote the original!</strong></p>\n<h3>XGBoost</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/anokas/data-analysis-xgboost-starter-0-35460-lb\" target=\"_blank\">Data Analysis &amp; XGBoost Starter (0.35460 LB)</a> (137K views, 1377 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/erikbruin/house-prices-lasso-xgboost-and-a-detailed-eda\" target=\"_blank\">House prices: Lasso, XGBoost, and a detailed EDA</a> (183K views, 1317 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dlarionov/feature-engineering-xgboost\" target=\"_blank\">Feature engineering, xgboost</a> (144K views, 942 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dansbecker/xgboost\" target=\"_blank\">XGBoost</a> (173K views, 918 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/alexisbcook/xgboost\" target=\"_blank\">XGBoost</a> (266K views, 578 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hamditarek/market-prediction-xgboost-with-gpu-fit-in-1min\" target=\"_blank\">Market Prediction: XGBoost with GPU (Fit in 1min)</a> (53K views, 411 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/a-guide-on-xgboost-hyperparameters-tuning\" target=\"_blank\">A Guide on XGBoost hyperparameters tuning</a> (55K views, 358 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/titanic-wcg-xgboost-0-84688\" target=\"_blank\">Titanic WCG+XGBoost [0.84688]</a> (27K views, 314 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/shahules/xgboost-feature-selection-dsbowl\" target=\"_blank\">XGBoost &amp; Feature Selection DSBowl 🥣 🥣</a> (36K views, 299 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tqchen/understanding-xgboost-model-on-otto-data\" target=\"_blank\">Understanding XGBoost Model on Otto Data</a> (130K views, 297 Votes)*</li>\n</ul>\n<h3>LightGBM</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/jsaguiar/lightgbm-with-simple-features\" target=\"_blank\">LightGBM with Simple Features</a> (85K views, 619 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix\" target=\"_blank\">LightGBM. Baseline Model Using Sparse Matrix</a> (49K views, 409 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tommy1028/lightgbm-starter-with-feature-engineering-idea\" target=\"_blank\">LightGBM starter with feature engineering idea</a> (16K views, 370 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/erikbruin/google-analytics-eda-lightgbm-screenshots\" target=\"_blank\">Google Analytics EDA + LightGBM + Screenshots</a> (40K views, 344 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/davidcairuz/feature-engineering-lightgbm\" target=\"_blank\">Feature Engineering &amp; LightGBM</a> (26K views, 272 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending\" target=\"_blank\">Simple LightGBM without blending</a> (18K views, 263 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/samratp/lightgbm-xgboost-catboost\" target=\"_blank\">LightGBM + XGBoost + Catboost</a> (37K views, 245 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/aitude/ashrae-kfold-lightgbm-without-leak-1-08\" target=\"_blank\">ASHRAE- KFold LightGBM - without leak (1.08)</a> (15K views, 240 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/lightgbm-classifier-in-python\" target=\"_blank\">LightGBM Classifier in Python</a> (52K views, 231 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-lb-0-9680\" target=\"_blank\">LightGBM (Fixing unbalanced data)</a> (56K views, 228 Votes)*</li>\n</ul>\n<h3>CatBoost</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/mhviraf/a-new-baseline-for-dsb-2019-catboost-model\" target=\"_blank\">A new baseline for DSB 2019 - Catboost model</a> (15K views, 316 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/samratp/lightgbm-xgboost-catboost\" target=\"_blank\">LightGBM + XGBoost + Catboost</a> (37K views, 245 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/abhinand05/catboost-a-deeper-dive\" target=\"_blank\">CatBoost: A Deeper Dive</a> (8K views, 178 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/allunia/house-prices-tutorial-with-catboost\" target=\"_blank\">House Prices Tutorial with Catboost</a> (12K views, 176 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm\" target=\"_blank\">Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM</a> (25K views, 173 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/usharengaraju/widsdatathon2021-catboost-starter\" target=\"_blank\">WiDSDatathon2021-Catboost-Starter</a> (6K views, 167 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mitribunskiy/tutorial-catboost-overview\" target=\"_blank\">Tutorial: CatBoost Overview</a> (39K views, 127 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/catboost-classifier-in-python\" target=\"_blank\">CatBoost Classifier in Python</a> (28K views, 126 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/braquino/catboost-some-more-features\" target=\"_blank\">Catboost - Some more features</a> (7K views, 103 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/mlcourse-ai-fall-2019-catboost-starter\" target=\"_blank\">mlcourse.ai. Fall 2019. Catboost starter</a> (18K views, 98 Votes)*</li>\n</ul>\n<h3>Ensemble</h3>\n<p>Now that we have our model trained, we should ensemble our predictions.</p>\n<p>Below you can find <a href=\"https://www.kaggle.com/discussions/getting-started/291437\" target=\"_blank\">The best of kaggle - Ensembles</a>.<br>\nThese are the top voted/viewed notebooks all over kaggle for GBMs.<br>\n<strong>Usage rules are simple: You like it? Upvote the original!</strong></p>\n<h3>Ensembling Techniques</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\" target=\"_blank\">Introduction to Ensembling/Stacking in Python</a> (620K views, 5436 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/yassineghouzam/titanic-top-4-with-ensemble-modeling\" target=\"_blank\">Titanic Top 4% with ensemble modeling</a> (185K views, 2501 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/viveksrinivasan/eda-ensemble-model-top-10-percentile\" target=\"_blank\">EDA &amp; Ensemble Model (Top 10 Percentile)</a> (98K views, 513 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble\" target=\"_blank\">Analysis of Melanoma Metadata and EffNet Ensemble</a> (24K views, 405 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tannercarbonati/detailed-data-analysis-ensemble-modeling\" target=\"_blank\">Detailed Data Analysis &amp; Ensemble Modeling</a> (55K views, 360 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/cdeotte/ensemble-folds-with-median-0-153\" target=\"_blank\">Ensemble Folds with MEDIAN - [0.153]</a> (9K views, 303 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/shonenkov/wbf-approach-for-ensemble\" target=\"_blank\">WBF approach for ensemble</a> (18K views, 282 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903\" target=\"_blank\">Ensemble: Resnext50_32x4d + Efficientnet = 0.903</a> (16K views, 274 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/arthurtok/employee-attrition-via-ensemble-tree-based-methods\" target=\"_blank\">Employee attrition via Ensemble tree-based methods</a> (64K views, 253 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/isaienkov/top-3-efficient-ensembling-in-few-lines-of-code\" target=\"_blank\">Top 3%. Efficient ensembling in few lines of code</a> (30K views, 242 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/pavansanagapati/ensemble-learning-techniques-tutorial\" target=\"_blank\">Ensemble Learning Techniques Tutorial</a> (55K views, 215 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jhoward/minimal-lstm-nb-svm-baseline-ensemble\" target=\"_blank\">Minimal LSTM + NB-SVM baseline ensemble</a> (30K views, 209 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example\" target=\"_blank\">Ensemble Model: Stacked Model Example</a> (68K views, 207 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hamditarek/ensemble\" target=\"_blank\">🤗🤗 Ensemble 🤗🤗</a> (64K views, 196 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble\" target=\"_blank\">EDA&amp;Modelling of the External Data Inc. Ensemble</a> (8K views, 195 Votes)*</li>\n</ul>\n<h3>Blending</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending\" target=\"_blank\">Simple LightGBM without blending</a> (18K views, 265 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/a763337092/blending-tensorflow-and-pytorch\" target=\"_blank\">Blending tensorflow and pytorch🔥🔥🔥</a> (13K views, 220 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/speedwagon/quadratic-discriminant-analysis\" target=\"_blank\">Another model for your blending</a> (11K views, 216 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/gogo827jz/optimise-blending-weights-with-bonus-0\" target=\"_blank\">Optimise Blending Weights with Bonus :0</a> (7K views, 184 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/suicaokhoailang/blending-with-linear-regression-0-688-lb\" target=\"_blank\">Blending with Linear Regression [0.688 LB]</a> (14K views, 136 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/abhishek/blending-blending-blending\" target=\"_blank\">blending blending blending</a> (3K views, 128 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jagangupta/lessons-from-toxic-blending-is-the-new-sexy\" target=\"_blank\">Lessons from Toxic : Blending is the new sexy</a> (10K views, 128 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\" target=\"_blank\">non-blending lightGBM model LB: 0.977</a> (16K views, 127 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hengzheng/bayesian-optimization-seed-blending\" target=\"_blank\">Bayesian Optimization Seed Blending</a> (9K views, 118 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/gpreda/elo-world-high-score-without-blending\" target=\"_blank\">elo_world_high_score_without_blending</a> (8K views, 112 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/abhishek/competition-part-5-blending-101\" target=\"_blank\">competition part-5: blending 101</a> (3K views, 105 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/beginner-eda-with-feature-eng-and-blending-models\" target=\"_blank\">Beginner EDA with Feature Eng. and Blending Models</a> (3K views, 99 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors\" target=\"_blank\">Cross-validation, weighted linear blending, errors</a> (7K views, 84 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/sagarjiyani/blending-tensorflow-pytorch-th-0-4914\" target=\"_blank\">Blending TensorFlow + PyTorch (th=0.4914)</a> (5K views, 82 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/gogo827jz/blending-nn-and-lgbm-rf\" target=\"_blank\">Blending NN and LGBM/RF</a> (6K views, 81 Votes)*</li>\n</ul>\n<h3>Stacking</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard\" target=\"_blank\">Stacked Regressions : Top 4% on LeaderBoard</a> (493K views, 6236 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\" target=\"_blank\">Introduction to Ensembling/Stacking in Python</a> (620K views, 5436 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dimitreoliveira/model-stacking-feature-engineering-and-eda\" target=\"_blank\">Model stacking, feature engineering and EDA</a> (43K views, 409 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/agodwinp/stacking-house-prices-walkthrough-to-top-5\" target=\"_blank\">Stacking House Prices - Walkthrough to Top 5%</a> (32K views, 227 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mmueller/stacking-starter\" target=\"_blank\">Stacking Starter</a> (41K views, 221 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example\" target=\"_blank\">Ensemble Model: Stacked Model Example</a> (68K views, 207 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/nicapotato/titanic-voting-pipeline-stack-and-guide\" target=\"_blank\">Titanic: Voting, Pipeline, Stack, and Guide</a> (21K views, 197 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/dongxu027/explore-stacking-lb-0-1463\" target=\"_blank\">Explore Stacking (LB 0.1463)</a> (18K views, 180 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm\" target=\"_blank\">Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM</a> (26K views, 176 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/datafan07/top-1-approach-eda-new-models-and-stacking\" target=\"_blank\">Top 1% Approach: EDA, New Models and Stacking</a> (9K views, 172 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/yekenot/simple-stacker-lb-0-284\" target=\"_blank\">Simple Stacker LB 0.284</a> (20K views, 172 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/itslek/stack-blend-lrs-xgb-lgb-house-prices-k-v17\" target=\"_blank\">Stack&amp;Blend LRs XGB LGB {House Prices K} v17</a> (6K views, 164 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/agehsbarg/top-10-0-10943-stacking-mice-and-brutal-force\" target=\"_blank\">Top 10 (0.10943): stacking, MICE and brutal force</a> (21K views, 156 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/hhstrand/oof-stacking-regime\" target=\"_blank\">OOF stacking regime</a> (14K views, 145 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tunguz/eloda-with-feature-engineering-and-stacking\" target=\"_blank\">EloDA with Feature Engineering and Stacking</a> (10K views, 141 Votes)*</li>\n</ul>\n<h3>Bagging</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/ammarnassanalhajali/riiid-lgbm-bagging2-sakt-0-781\" target=\"_blank\">Riiid LGBM bagging2 + SAKT =0.781</a> (11K views, 188 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/leadbest/sakt-riiid-lgbm-bagging2\" target=\"_blank\">SAKT + Riiid LGBM bagging2</a> (8K views, 139 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/topic-5-ensembles-part-1-bagging\" target=\"_blank\">Topic 5. Ensembles. Part 1. Bagging</a> (16K views, 137 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/bagging-vs-boosting\" target=\"_blank\">Bagging vs Boosting</a> (22K views, 117 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/julianguo/fork-of-riiid-lgbm-bagging2-1-471152\" target=\"_blank\">Fork of Riiid LGBM bagging2.1 471152</a> (8K views, 117 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2\" target=\"_blank\">Riiid! LGBM bagging2</a> (6K views, 113 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kabure/predicting-house-prices-xgb-rf-bagging-reg-pipe\" target=\"_blank\">Predicting House Prices [XGB/RF/Bagging-Reg Pipe]</a> (10K views, 89 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/danijelk/keras-starter-with-bagging-lb-1120-596\" target=\"_blank\">Keras starter with bagging (LB: 1120.596)</a> (18K views, 77 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/mtinti/keras-starter-with-bagging-1111-84364\" target=\"_blank\">Keras starter with bagging 1111.84364</a> (19K views, 77 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2-1\" target=\"_blank\">Riiid LGBM bagging2.1</a> (4K views, 62 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/andypenrose/baggingregressor-rapids-ensemble\" target=\"_blank\">BaggingRegressor + RAPIDS Ensemble</a> (4K views, 59 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/fahadmehfoooz/credit-analysis-with-knn-dtree-rf-bagging-ann\" target=\"_blank\">Credit analysis with KNN/DTree/RF/Bagging/ANN🏦</a> (&lt;1K views, 55 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm\" target=\"_blank\">Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM</a> (1K views, 51 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting\" target=\"_blank\">Ensemble ML Algorithms : Bagging, Boosting, Voting</a> (8K views, 48 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/cttsai/simple-xgboost-cv-bagging\" target=\"_blank\">Simple XGBoost CV Bagging</a> (3K views, 42 Votes)*</li>\n</ul>\n<h3>Boosting</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/felipemello/boosting-creativity-towards-feature-engineering\" target=\"_blank\">Boosting creativity towards feature engineering</a> (9K views, 229 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/yashvi/vehicle-insurance-eda-and-boosting-models\" target=\"_blank\">Vehicle Insurance EDA and boosting models</a> (19K views, 200 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/topic-10-gradient-boosting\" target=\"_blank\">Topic 10. Gradient Boosting</a> (21K views, 188 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/grroverpr/gradient-boosting-simplified\" target=\"_blank\">Gradient boosting simplified</a> (48K views, 147 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/prashant111/bagging-vs-boosting\" target=\"_blank\">Bagging vs Boosting</a> (22K views, 117 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/niteshyadav3103/diabetes-prediction-stacking-boosting\" target=\"_blank\">Diabetes Prediction (Stacking + Boosting)</a> (1K views, 85 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/roydatascience/elo-stack-with-goss-boosting\" target=\"_blank\">Elo Stack With Goss Boosting</a> (6K views, 78 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/kashnitsky/assignment-10-gradient-boosting-and-flight-delays\" target=\"_blank\">Assignment 10. Gradient boosting and flight delays</a> (9K views, 71 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/lavanyashukla01/battle-of-the-boosting-algos-lgb-xgb-catboost\" target=\"_blank\">Battle of the Boosting Algos: LGB, XGB, Catboost</a> (9K views, 55 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/vinnsvinay/introduction-to-boosting-using-lgbm-lb-0-68357\" target=\"_blank\">Introduction to Boosting using LGBM(LB: 0.68357)</a> (14K views, 52 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/tunguz/fe-pipeline-with-histgradientboostingregressor\" target=\"_blank\">FE-Pipeline with HistGradientBoostingRegressor</a> (3K views, 51 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm\" target=\"_blank\">Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM</a> (1K views, 51 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/parulpandey/explainable-boosting-machines-for-tabular-data\" target=\"_blank\">Explainable Boosting machines for Tabular data</a> (2K views, 49 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting\" target=\"_blank\">Ensemble ML Algorithms : Bagging, Boosting, Voting</a> (8K views, 48 Votes)*</li>\n<li><a href=\"https://www.kaggle.com/sid321axn/house-price-prediction-gboosting-adaboost-etc\" target=\"_blank\">House Price Prediction : GBoosting,AdaBoost etc.</a> (8K views, 42 Votes)*</li>\n</ul>\n<h3>Hyperparameters Tuning</h3>\n<ul>\n<li>LGBM <a href=\"https://www.kaggle.com/mlisovyi/lightgbm-hyperparameter-optimisation-lb-0-761\" target=\"_blank\">hyperparameter tuning</a> methods.</li>\n<li>Automated <a href=\"https://www.kaggle.com/willkoehrsen/automated-model-tuning\" target=\"_blank\">model tuning</a> methods.</li>\n<li>Parameter tuning with <a href=\"https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt\" target=\"_blank\">hyper plot</a>.</li>\n<li><a href=\"http://krasserm.github.io/2018/03/21/bayesian-optimization/\" target=\"_blank\">Bayesian optimization</a> for hyperparameter tuning.</li>\n<li><a href=\"https://www.kaggle.com/nicapotato/gpyopt-hyperparameter-optimisation-gpu-lgbm\" target=\"_blank\">Gpyopt Hyperparameter Optimisation</a>.</li>\n</ul>",
      "rawMarkdown": "# Tabular Classification - Tips and Tricks\n\n\nSharing a long list of tips and tricks found on previous tabular classification competitions on Kaggle.\nTake this as a list of ideas you can attempt to try for improving your score!\n\n> **Credits:** \n> - Many of the links are based on parts of [this](https://neptune.ai/blog/tabular-data-binary-classification-tips-and-tricks-from-5-kaggle-competitions) blog post (with edits and filtering).\n> - Including links from: [The best of kaggle - GBMs](https://www.kaggle.com/discussions/getting-started/277121).\n> - Also from: [The best of kaggle - Ensembles](https://www.kaggle.com/discussions/getting-started/291437).\n\n### Dealing with larger datasets\n\nAn issue close to our hearts! \n\n* Faster [data loading with pandas.](https://www.kaggle.com/c/home-credit-default-risk/discussion/59575)\n* Data compression techniques to [reduce the size of data by 70%](https://www.kaggle.com/nickycan/compress-70-of-dataset).\n* Optimize the memory by r[educing the size of some attributes.](https://www.kaggle.com/shrutimechlearn/large-data-loading-trick-with-ms-malware-data)\n* Use open-source libraries such as [Dask to read and manipulate the data](https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask), it performs parallel computing and saves up memory space.\n* Use [cudf](https://github.com/rapidsai/cudf).\n* Convert data to [parquet](https://arrow.apache.org/docs/python/parquet.html) format.\n* Converting data to [feather](https://medium.com/@snehotosh.banerjee/feather-a-fast-on-disk-format-for-r-and-python-data-frames-de33d0516b03) format.\n* Reducing memory usage for [optimizing RAM](https://www.kaggle.com/mjbahmani/reducing-memory-size-for-ieee).\n\n### Data exploration\n\nData exploration is a must on all competitions, Here are some references from past competitions use them to get some ideas on what exactly to explore.\n\n* EDA for microsoft [malware detection.](https://www.kaggle.com/youhanlee/my-eda-i-want-to-see-all)\n* Time Series [EDA for malware detection.](https://www.kaggle.com/cdeotte/time-split-validation-malware-0-68)\n* Complete [EDA for home credit loan prediction](https://www.kaggle.com/codename007/home-credit-complete-eda-feature-importance).\n* Complete [EDA for Santader prediction.](https://www.kaggle.com/gpreda/santander-eda-and-prediction)\n* EDA for [VSB Power Line Fault Detection.](https://www.kaggle.com/go1dfish/basic-eda)\n\n### Preprocessing\n\nBefore we start to engineer features, we need to do some preprocessing.\nThose invlolve removing missing values, converting categorical variables to numeric, and scaling the data, etc..\nHere are some references from past competitions use them to get some ideas for preprocessing.\n\n* Methods to [tackle class imbalance](https://www.kaggle.com/shahules/tackling-class-imbalance).\n* Data augmentation by [Synthetic Minority Oversampling Technique](https://machinelearningmastery.com/smote-oversampling-for-imbalanced-classification/).\n* Fast inplace [shuffle for augmentation](https://www.kaggle.com/jiweiliu/fast-inplace-shuffle-for-augmentation).\n* Finding [synthetic samples in the dataset.](https://www.kaggle.com/yag320/list-of-fake-samples-and-public-private-lb-split)\n* [Signal denoising](https://www.kaggle.com/jackvial/dwt-signal-denoising) used in signal processing competitions.\n* Finding [patterns of missing data](https://www.kaggle.com/jpmiller/patterns-of-missing-data).\n* Methods to handle [missing data](https://towardsdatascience.com/6-different-ways-to-compensate-for-missing-values-data-imputation-with-examples-6022d9ca0779).\n* An overview of various [encoding techniques for categorical data.](https://www.kaggle.com/shahules/an-overview-of-encoding-techniques)\n* Building [model to predict missing values.](https://www.kaggle.com/c/home-credit-default-risk/discussion/64598)\n* Random [shuffling of data](https://www.kaggle.com/brandenkmurray/randomly-shuffled-data-also-works) to create new synthetic training set.\n\n### Feature engineering\n\nNow we can start to engineer features, below you can find some references from past competitions use them to get some ideas to engineer features. \n\n* Target [encoding cross validation](https://medium.com/@pouryaayria/k-fold-target-encoding-dfe9a594874b/) for better encoding.\n* Entity embedding to [handle categories](https://www.kaggle.com/abhishek/entity-embeddings-to-handle-categories).\n* Encoding c[yclic features for deep learning.](https://www.kaggle.com/avanwyk/encoding-cyclical-features-for-deep-learning)\n* Manual [feature engineering methods](https://www.kaggle.com/willkoehrsen/introduction-to-manual-feature-engineering).\n* Automated feature engineering techniques [using featuretools](https://www.kaggle.com/willkoehrsen/automated-feature-engineering-basics).\n* Top hard crafted features used in [microsoft malware detection](https://www.kaggle.com/sanderf/7th-place-solution-microsoft-malware-prediction).\n* Denoising NN for [feature extraction](https://towardsdatascience.com/applied-deep-learning-part-3-autoencoders-1c083af4d798).\n* Feature engineering [using RAPIDS framework.](https://www.kaggle.com/cdeotte/rapids-feature-engineering-fraud-0-96/)\n* Things to remember while processing f[eatures using LGBM.](https://www.kaggle.com/c/ieee-fraud-detection/discussion/108575)\n* [Lag features and moving averages.](https://www.kaggle.com/c/home-credit-default-risk/discussion/64593)\n* [Principal component analysis](https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567) for dimensionality reduction.\n* LDA for [dimensionality reduction](https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567).\n* Best hand crafted LGBM features for [microsoft malware detection](https://www.kaggle.com/c/microsoft-malware-prediction/discussion/85157).\n* Generating [frequency features.](https://www.kaggle.com/philippsinger/frequency-features-without-test-data-information)\n* Dropping variables with [different train and test distribution.](https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix)\n* [Aggregate time series features](https://www.kaggle.com/c/home-credit-default-risk/discussion/64593) for home credit competition.\n* [Time Series](https://www.kaggle.com/c/home-credit-default-risk/discussion/64593) features used in home credit default risk.\n* Scale, Standardize and n[ormalize with sklearn](https://towardsdatascience.com/scale-standardize-or-normalize-with-scikit-learn-6ccc7d176a02).\n* Handcrafted features for [Home default risk competition.](https://www.kaggle.com/c/home-credit-default-risk/discussion/57750)\n* Handcrafted [features used in Santander Transaction Prediction.](https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/89070)\n\n### Feature selection\n\nNow that we got our features engineered, we should check if they are useful or not.\nTo do this, we need to do some feature selection.\nHere are some references from past competitions use them to get some ideas for feature selection.\n\n* Six ways to do [features selection using sklearn](https://www.kaggle.com/sz8416/6-ways-for-feature-selection).\n* [Permutation feature importance](https://www.kaggle.com/c/ieee-fraud-detection/discussion/107877#latest-635386).\n* [Adversarial feature validation](https://www.kaggle.com/tunguz/adversarial-ieee/).\n* Feature selection using [null importances.](https://www.kaggle.com/ogrellier/feature-selection-with-null-importances)\n* Tree explainer using [SHAP.](https://github.com/slundberg/shap)\n* DeepNN explainer using [SHAP](https://github.com/slundberg/shap).\n\n\n### Modeling\n\nBelow you can find [The best of kaggle - GBMs](https://www.kaggle.com/discussions/getting-started/277121).\nThese are the top voted/viewed notebooks all over kaggle for GBMs.  \n**Usage rules are simple: You like it? Upvote the original!**\n\n### XGBoost\n \n* [Data Analysis & XGBoost Starter (0.35460 LB)](https://www.kaggle.com/anokas/data-analysis-xgboost-starter-0-35460-lb) (137K views, 1377 Votes)*\n* [House prices: Lasso, XGBoost, and a detailed EDA](https://www.kaggle.com/erikbruin/house-prices-lasso-xgboost-and-a-detailed-eda) (183K views, 1317 Votes)*\n* [Feature engineering, xgboost](https://www.kaggle.com/dlarionov/feature-engineering-xgboost) (144K views, 942 Votes)*\n* [XGBoost](https://www.kaggle.com/dansbecker/xgboost) (173K views, 918 Votes)*\n* [XGBoost](https://www.kaggle.com/alexisbcook/xgboost) (266K views, 578 Votes)*\n* [Market Prediction: XGBoost with GPU (Fit in 1min)](https://www.kaggle.com/hamditarek/market-prediction-xgboost-with-gpu-fit-in-1min) (53K views, 411 Votes)*\n* [A Guide on XGBoost hyperparameters tuning](https://www.kaggle.com/prashant111/a-guide-on-xgboost-hyperparameters-tuning) (55K views, 358 Votes)*\n* [Titanic WCG+XGBoost [0.84688]](https://www.kaggle.com/cdeotte/titanic-wcg-xgboost-0-84688) (27K views, 314 Votes)*\n* [XGBoost & Feature Selection DSBowl 🥣 🥣](https://www.kaggle.com/shahules/xgboost-feature-selection-dsbowl) (36K views, 299 Votes)*\n* [Understanding XGBoost Model on Otto Data](https://www.kaggle.com/tqchen/understanding-xgboost-model-on-otto-data) (130K views, 297 Votes)*\n\n### LightGBM\n \n* [LightGBM with Simple Features](https://www.kaggle.com/jsaguiar/lightgbm-with-simple-features) (85K views, 619 Votes)*\n* [LightGBM. Baseline Model Using Sparse Matrix](https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix) (49K views, 409 Votes)*\n* [LightGBM starter with feature engineering idea](https://www.kaggle.com/tommy1028/lightgbm-starter-with-feature-engineering-idea) (16K views, 370 Votes)*\n* [Google Analytics EDA + LightGBM + Screenshots](https://www.kaggle.com/erikbruin/google-analytics-eda-lightgbm-screenshots) (40K views, 344 Votes)*\n* [Feature Engineering & LightGBM](https://www.kaggle.com/davidcairuz/feature-engineering-lightgbm) (26K views, 272 Votes)*\n* [Simple LightGBM without blending](https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending) (18K views, 263 Votes)*\n* [LightGBM + XGBoost + Catboost](https://www.kaggle.com/samratp/lightgbm-xgboost-catboost) (37K views, 245 Votes)*\n* [ASHRAE- KFold LightGBM - without leak (1.08)](https://www.kaggle.com/aitude/ashrae-kfold-lightgbm-without-leak-1-08) (15K views, 240 Votes)*\n* [LightGBM Classifier in Python](https://www.kaggle.com/prashant111/lightgbm-classifier-in-python) (52K views, 231 Votes)*\n* [LightGBM (Fixing unbalanced data)](https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-lb-0-9680) (56K views, 228 Votes)*\n\n### CatBoost\n \n* [A new baseline for DSB 2019 - Catboost model](https://www.kaggle.com/mhviraf/a-new-baseline-for-dsb-2019-catboost-model) (15K views, 316 Votes)*\n* [LightGBM + XGBoost + Catboost](https://www.kaggle.com/samratp/lightgbm-xgboost-catboost) (37K views, 245 Votes)*\n* [CatBoost: A Deeper Dive](https://www.kaggle.com/abhinand05/catboost-a-deeper-dive) (8K views, 178 Votes)*\n* [House Prices Tutorial with Catboost](https://www.kaggle.com/allunia/house-prices-tutorial-with-catboost) (12K views, 176 Votes)*\n* [Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM](https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm) (25K views, 173 Votes)*\n* [WiDSDatathon2021-Catboost-Starter](https://www.kaggle.com/usharengaraju/widsdatathon2021-catboost-starter) (6K views, 167 Votes)*\n* [Tutorial: CatBoost Overview](https://www.kaggle.com/mitribunskiy/tutorial-catboost-overview) (39K views, 127 Votes)*\n* [CatBoost Classifier in Python](https://www.kaggle.com/prashant111/catboost-classifier-in-python) (28K views, 126 Votes)*\n* [Catboost - Some more features](https://www.kaggle.com/braquino/catboost-some-more-features) (7K views, 103 Votes)*\n* [mlcourse.ai. Fall 2019. Catboost starter](https://www.kaggle.com/kashnitsky/mlcourse-ai-fall-2019-catboost-starter) (18K views, 98 Votes)*\n\n\n### Ensemble\n\nNow that we have our model trained, we should ensemble our predictions.\n\nBelow you can find [The best of kaggle - Ensembles](https://www.kaggle.com/discussions/getting-started/291437).\nThese are the top voted/viewed notebooks all over kaggle for GBMs.\n**Usage rules are simple: You like it? Upvote the original!**\n\n### Ensembling Techniques\n\n* [Introduction to Ensembling/Stacking in Python](https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python) (620K views, 5436 Votes)*\n* [Titanic Top 4% with ensemble modeling](https://www.kaggle.com/yassineghouzam/titanic-top-4-with-ensemble-modeling) (185K views, 2501 Votes)*\n* [EDA & Ensemble Model (Top 10 Percentile)](https://www.kaggle.com/viveksrinivasan/eda-ensemble-model-top-10-percentile) (98K views, 513 Votes)*\n* [Analysis of Melanoma Metadata and EffNet Ensemble](https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble) (24K views, 405 Votes)*\n* [Detailed Data Analysis & Ensemble Modeling](https://www.kaggle.com/tannercarbonati/detailed-data-analysis-ensemble-modeling) (55K views, 360 Votes)*\n* [Ensemble Folds with MEDIAN - [0.153]](https://www.kaggle.com/cdeotte/ensemble-folds-with-median-0-153) (9K views, 303 Votes)*\n* [WBF approach for ensemble](https://www.kaggle.com/shonenkov/wbf-approach-for-ensemble) (18K views, 282 Votes)*\n* [Ensemble: Resnext50\\_32x4d + Efficientnet = 0.903](https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903) (16K views, 274 Votes)*\n* [Employee attrition via Ensemble tree-based methods](https://www.kaggle.com/arthurtok/employee-attrition-via-ensemble-tree-based-methods) (64K views, 253 Votes)*\n* [Top 3%. Efficient ensembling in few lines of code](https://www.kaggle.com/isaienkov/top-3-efficient-ensembling-in-few-lines-of-code) (30K views, 242 Votes)*\n* [Ensemble Learning Techniques Tutorial](https://www.kaggle.com/pavansanagapati/ensemble-learning-techniques-tutorial) (55K views, 215 Votes)*\n* [Minimal LSTM + NB-SVM baseline ensemble](https://www.kaggle.com/jhoward/minimal-lstm-nb-svm-baseline-ensemble) (30K views, 209 Votes)*\n* [Ensemble Model: Stacked Model Example](https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example) (68K views, 207 Votes)*\n* [🤗🤗 Ensemble 🤗🤗](https://www.kaggle.com/hamditarek/ensemble) (64K views, 196 Votes)*\n* [EDA&Modelling of the External Data Inc. Ensemble](https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble) (8K views, 195 Votes)*\n\n### Blending\n \n* [Simple LightGBM without blending](https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending) (18K views, 265 Votes)*\n* [Blending tensorflow and pytorch🔥🔥🔥](https://www.kaggle.com/a763337092/blending-tensorflow-and-pytorch) (13K views, 220 Votes)*\n* [Another model for your blending](https://www.kaggle.com/speedwagon/quadratic-discriminant-analysis) (11K views, 216 Votes)*\n* [Optimise Blending Weights with Bonus :0](https://www.kaggle.com/gogo827jz/optimise-blending-weights-with-bonus-0) (7K views, 184 Votes)*\n* [Blending with Linear Regression [0.688 LB]](https://www.kaggle.com/suicaokhoailang/blending-with-linear-regression-0-688-lb) (14K views, 136 Votes)*\n* [blending blending blending](https://www.kaggle.com/abhishek/blending-blending-blending) (3K views, 128 Votes)*\n* [Lessons from Toxic : Blending is the new sexy](https://www.kaggle.com/jagangupta/lessons-from-toxic-blending-is-the-new-sexy) (10K views, 128 Votes)*\n* [non-blending lightGBM model LB: 0.977](https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977) (16K views, 127 Votes)*\n* [Bayesian Optimization Seed Blending](https://www.kaggle.com/hengzheng/bayesian-optimization-seed-blending) (9K views, 118 Votes)*\n* [elo\\_world\\_high\\_score\\_without\\_blending](https://www.kaggle.com/gpreda/elo-world-high-score-without-blending) (8K views, 112 Votes)*\n* [competition part-5: blending 101](https://www.kaggle.com/abhishek/competition-part-5-blending-101) (3K views, 105 Votes)*\n* [Beginner EDA with Feature Eng. and Blending Models](https://www.kaggle.com/datafan07/beginner-eda-with-feature-eng-and-blending-models) (3K views, 99 Votes)*\n* [Cross-validation, weighted linear blending, errors](https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors) (7K views, 84 Votes)*\n* [Blending TensorFlow + PyTorch (th=0.4914)](https://www.kaggle.com/sagarjiyani/blending-tensorflow-pytorch-th-0-4914) (5K views, 82 Votes)*\n* [Blending NN and LGBM/RF](https://www.kaggle.com/gogo827jz/blending-nn-and-lgbm-rf) (6K views, 81 Votes)*\n\n### Stacking\n \n* [Stacked Regressions : Top 4% on LeaderBoard](https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard) (493K views, 6236 Votes)*\n* [Introduction to Ensembling/Stacking in Python](https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python) (620K views, 5436 Votes)*\n* [Model stacking, feature engineering and EDA](https://www.kaggle.com/dimitreoliveira/model-stacking-feature-engineering-and-eda) (43K views, 409 Votes)*\n* [Stacking House Prices - Walkthrough to Top 5%](https://www.kaggle.com/agodwinp/stacking-house-prices-walkthrough-to-top-5) (32K views, 227 Votes)*\n* [Stacking Starter](https://www.kaggle.com/mmueller/stacking-starter) (41K views, 221 Votes)*\n* [Ensemble Model: Stacked Model Example](https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example) (68K views, 207 Votes)*\n* [Titanic: Voting, Pipeline, Stack, and Guide](https://www.kaggle.com/nicapotato/titanic-voting-pipeline-stack-and-guide) (21K views, 197 Votes)*\n* [Explore Stacking (LB 0.1463)](https://www.kaggle.com/dongxu027/explore-stacking-lb-0-1463) (18K views, 180 Votes)*\n* [Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM](https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm) (26K views, 176 Votes)*\n* [Top 1% Approach: EDA, New Models and Stacking](https://www.kaggle.com/datafan07/top-1-approach-eda-new-models-and-stacking) (9K views, 172 Votes)*\n* [Simple Stacker LB 0.284](https://www.kaggle.com/yekenot/simple-stacker-lb-0-284) (20K views, 172 Votes)*\n* [Stack&Blend LRs XGB LGB {House Prices K} v17](https://www.kaggle.com/itslek/stack-blend-lrs-xgb-lgb-house-prices-k-v17) (6K views, 164 Votes)*\n* [Top 10 (0.10943): stacking, MICE and brutal force](https://www.kaggle.com/agehsbarg/top-10-0-10943-stacking-mice-and-brutal-force) (21K views, 156 Votes)*\n* [OOF stacking regime](https://www.kaggle.com/hhstrand/oof-stacking-regime) (14K views, 145 Votes)*\n* [EloDA with Feature Engineering and Stacking](https://www.kaggle.com/tunguz/eloda-with-feature-engineering-and-stacking) (10K views, 141 Votes)*\n\n### Bagging\n \n* [Riiid LGBM bagging2 + SAKT =0.781](https://www.kaggle.com/ammarnassanalhajali/riiid-lgbm-bagging2-sakt-0-781) (11K views, 188 Votes)*\n* [SAKT + Riiid LGBM bagging2](https://www.kaggle.com/leadbest/sakt-riiid-lgbm-bagging2) (8K views, 139 Votes)*\n* [Topic 5. Ensembles. Part 1. Bagging](https://www.kaggle.com/kashnitsky/topic-5-ensembles-part-1-bagging) (16K views, 137 Votes)*\n* [Bagging vs Boosting](https://www.kaggle.com/prashant111/bagging-vs-boosting) (22K views, 117 Votes)*\n* [Fork of Riiid LGBM bagging2.1 471152](https://www.kaggle.com/julianguo/fork-of-riiid-lgbm-bagging2-1-471152) (8K views, 117 Votes)*\n* [Riiid! LGBM bagging2](https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2) (6K views, 113 Votes)*\n* [Predicting House Prices [XGB/RF/Bagging-Reg Pipe]](https://www.kaggle.com/kabure/predicting-house-prices-xgb-rf-bagging-reg-pipe) (10K views, 89 Votes)*\n* [Keras starter with bagging (LB: 1120.596)](https://www.kaggle.com/danijelk/keras-starter-with-bagging-lb-1120-596) (18K views, 77 Votes)*\n* [Keras starter with bagging 1111.84364](https://www.kaggle.com/mtinti/keras-starter-with-bagging-1111-84364) (19K views, 77 Votes)*\n* [Riiid LGBM bagging2.1](https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2-1) (4K views, 62 Votes)*\n* [BaggingRegressor + RAPIDS Ensemble](https://www.kaggle.com/andypenrose/baggingregressor-rapids-ensemble) (4K views, 59 Votes)*\n* [Credit analysis with KNN/DTree/RF/Bagging/ANN🏦](https://www.kaggle.com/fahadmehfoooz/credit-analysis-with-knn-dtree-rf-bagging-ann) (<1K views, 55 Votes)*\n* [Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM](https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm) (1K views, 51 Votes)*\n* [Ensemble ML Algorithms : Bagging, Boosting, Voting](https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting) (8K views, 48 Votes)*\n* [Simple XGBoost CV Bagging](https://www.kaggle.com/cttsai/simple-xgboost-cv-bagging) (3K views, 42 Votes)*\n\n### Boosting\n \n* [Boosting creativity towards feature engineering](https://www.kaggle.com/felipemello/boosting-creativity-towards-feature-engineering) (9K views, 229 Votes)*\n* [Vehicle Insurance EDA and boosting models](https://www.kaggle.com/yashvi/vehicle-insurance-eda-and-boosting-models) (19K views, 200 Votes)*\n* [Topic 10. Gradient Boosting](https://www.kaggle.com/kashnitsky/topic-10-gradient-boosting) (21K views, 188 Votes)*\n* [Gradient boosting simplified](https://www.kaggle.com/grroverpr/gradient-boosting-simplified) (48K views, 147 Votes)*\n* [Bagging vs Boosting](https://www.kaggle.com/prashant111/bagging-vs-boosting) (22K views, 117 Votes)*\n* [Diabetes Prediction (Stacking + Boosting)](https://www.kaggle.com/niteshyadav3103/diabetes-prediction-stacking-boosting) (1K views, 85 Votes)*\n* [Elo Stack With Goss Boosting](https://www.kaggle.com/roydatascience/elo-stack-with-goss-boosting) (6K views, 78 Votes)*\n* [Assignment 10. Gradient boosting and flight delays](https://www.kaggle.com/kashnitsky/assignment-10-gradient-boosting-and-flight-delays) (9K views, 71 Votes)*\n* [Battle of the Boosting Algos: LGB, XGB, Catboost](https://www.kaggle.com/lavanyashukla01/battle-of-the-boosting-algos-lgb-xgb-catboost) (9K views, 55 Votes)*\n* [Introduction to Boosting using LGBM(LB: 0.68357)](https://www.kaggle.com/vinnsvinay/introduction-to-boosting-using-lgbm-lb-0-68357) (14K views, 52 Votes)*\n* [FE-Pipeline with HistGradientBoostingRegressor](https://www.kaggle.com/tunguz/fe-pipeline-with-histgradientboostingregressor) (3K views, 51 Votes)*\n* [Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM](https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm) (1K views, 51 Votes)*\n* [Explainable Boosting machines for Tabular data](https://www.kaggle.com/parulpandey/explainable-boosting-machines-for-tabular-data) (2K views, 49 Votes)*\n* [Ensemble ML Algorithms : Bagging, Boosting, Voting](https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting) (8K views, 48 Votes)*\n* [House Price Prediction : GBoosting,AdaBoost etc.](https://www.kaggle.com/sid321axn/house-price-prediction-gboosting-adaboost-etc) (8K views, 42 Votes)*\n\n\n### Hyperparameters Tuning\n\n* LGBM [hyperparameter tuning](https://www.kaggle.com/mlisovyi/lightgbm-hyperparameter-optimisation-lb-0-761) methods.\n* Automated [model tuning](https://www.kaggle.com/willkoehrsen/automated-model-tuning) methods.\n* Parameter tuning with [hyper plot](https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt).\n* [Bayesian optimization](http://krasserm.github.io/2018/03/21/bayesian-optimization/) for hyperparameter tuning.\n* [Gpyopt Hyperparameter Optimisation](https://www.kaggle.com/nicapotato/gpyopt-hyperparameter-optimisation-gpu-lgbm).",
      "votes": 357
    },
    {
      "id": 2325446,
      "postDate": "2023-07-01T10:55:05.263Z",
      "content": "<p>This is really very helpful.</p>",
      "rawMarkdown": "This is really very helpful."
    },
    {
      "id": 2192622,
      "postDate": "2023-03-22T19:27:42.413Z",
      "content": "<p>This is awesome resource , Keep up the great work.👍👍</p>",
      "rawMarkdown": "This is awesome resource , Keep up the great work.👍👍"
    },
    {
      "id": 2064823,
      "postDate": "2022-12-14T06:23:32.787Z",
      "content": "<p>Fantastic resource for beginner, extremely grateful to you! <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "rawMarkdown": "Fantastic resource for beginner, extremely grateful to you! @thedevastator "
    },
    {
      "id": 2050768,
      "postDate": "2022-12-01T01:35:48.947Z",
      "content": "<p>A fantastic resource all in one spot. Thank you</p>",
      "rawMarkdown": "A fantastic resource all in one spot. Thank you"
    },
    {
      "id": 1911840,
      "postDate": "2022-08-24T10:53:45.570Z",
      "content": "<p>Thank you for taking the time to put this together. Very helpful!</p>",
      "rawMarkdown": "Thank you for taking the time to put this together. Very helpful!"
    },
    {
      "id": 1897431,
      "postDate": "2022-08-13T18:02:28.747Z",
      "content": "<p>This is very useful </p>",
      "rawMarkdown": "This is very useful "
    },
    {
      "id": 1879277,
      "postDate": "2022-08-01T01:02:11.613Z",
      "content": "<p>very helpful.</p>",
      "rawMarkdown": "very helpful."
    },
    {
      "id": 1879239,
      "postDate": "2022-07-31T23:54:45.377Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, so much knowledge, thanks for sharing ✨</p>",
      "rawMarkdown": "Hello @thedevastator, so much knowledge, thanks for sharing ✨"
    },
    {
      "id": 1864977,
      "postDate": "2022-07-21T12:15:32.993Z",
      "content": "<p>tttttthhhhhhhhhhhhaaaaaaaaaaaannnnnnnnnnnnnnnnnk</p>",
      "rawMarkdown": "tttttthhhhhhhhhhhhaaaaaaaaaaaannnnnnnnnnnnnnnnnk"
    },
    {
      "id": 1861816,
      "postDate": "2022-07-19T09:04:36.197Z",
      "content": "<p>Thank you so much for sharing these tips</p>",
      "rawMarkdown": "Thank you so much for sharing these tips"
    },
    {
      "id": 1860925,
      "postDate": "2022-07-18T16:49:56.227Z",
      "content": "<p>Thank you so much for this! I'm barely starting with Kaggle competitions but I'm saving this one for later:)</p>",
      "rawMarkdown": "Thank you so much for this! I'm barely starting with Kaggle competitions but I'm saving this one for later:)"
    },
    {
      "id": 1858544,
      "postDate": "2022-07-17T02:42:41.103Z",
      "content": "<p>Thanks so much for putting this together!</p>",
      "rawMarkdown": "Thanks so much for putting this together!"
    },
    {
      "id": 1857191,
      "postDate": "2022-07-15T23:56:53.627Z",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, thanks for sharing a massive amount of information to keep improving</p>",
      "rawMarkdown": "Hello, @thedevastator, thanks for sharing a massive amount of information to keep improving"
    },
    {
      "id": 1857146,
      "postDate": "2022-07-15T22:09:31.913Z",
      "content": "<p>Wow !!<br>\nThank you very much for the amount of time you have contributed to make this…</p>",
      "rawMarkdown": "Wow !!\nThank you very much for the amount of time you have contributed to make this..."
    },
    {
      "id": 1855775,
      "postDate": "2022-07-14T22:37:56.487Z",
      "content": "<p>wow great stuff, thank you for sharing .</p>",
      "rawMarkdown": "wow great stuff, thank you for sharing ."
    },
    {
      "id": 1855526,
      "postDate": "2022-07-14T16:56:47.500Z",
      "content": "<p>Nice post, thanks for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a></p>",
      "rawMarkdown": "Nice post, thanks for sharing @thedevastator"
    },
    {
      "id": 1852313,
      "postDate": "2022-07-12T01:52:29.773Z",
      "content": "<p>This is remarkable summary!<br>\nGreat!</p>",
      "rawMarkdown": "This is remarkable summary!\nGreat!"
    },
    {
      "id": 1849096,
      "postDate": "2022-07-09T08:04:18.123Z",
      "content": "<p>very helpful.</p>",
      "rawMarkdown": "very helpful."
    },
    {
      "id": 1848827,
      "postDate": "2022-07-08T23:57:45.223Z",
      "content": "<p>Nice post, thanks for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "rawMarkdown": "Nice post, thanks for sharing @thedevastator "
    },
    {
      "id": 1848622,
      "postDate": "2022-07-08T19:02:11.103Z",
      "content": "<p>You have done a aggregate of aggregates. Thanks somethings I was just looking for found here . I think you should put this in a github  repo.. </p>",
      "rawMarkdown": "You have done a aggregate of aggregates. Thanks somethings I was just looking for found here . I think you should put this in a github  repo.. "
    },
    {
      "id": 1897691,
      "postDate": "2022-08-14T02:35:51.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1883154,
      "postDate": "2022-08-03T16:02:13.380Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1848917,
      "postDate": "2022-07-09T03:43:09.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1848302,
      "postDate": "2022-07-08T14:01:53.160Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1848077,
      "postDate": "2022-07-08T10:25:17.673Z",
      "content": "<p>Wow! Very useful! Thank you!!!</p>",
      "rawMarkdown": "Wow! Very useful! Thank you!!!",
      "votes": 1
    },
    {
      "id": 2012013,
      "postDate": "2022-11-01T02:37:48.443Z",
      "content": "<p>Super useful! Thank you! </p>",
      "rawMarkdown": "Super useful! Thank you! "
    },
    {
      "id": 2006418,
      "postDate": "2022-10-27T15:23:37.077Z",
      "content": "<p>Thanks for the insight!</p>",
      "rawMarkdown": "Thanks for the insight!"
    },
    {
      "id": 1992926,
      "postDate": "2022-10-18T02:48:47.557Z",
      "content": "<p>Thanks for the insight!</p>",
      "rawMarkdown": "Thanks for the insight!"
    },
    {
      "id": 1913373,
      "postDate": "2022-08-25T09:54:14.937Z",
      "content": "<p>Thanks for the insight!</p>",
      "rawMarkdown": "Thanks for the insight!"
    },
    {
      "id": 1910876,
      "postDate": "2022-08-23T18:43:31.537Z",
      "content": "<p>This is excellent. Thanks for sharing</p>",
      "rawMarkdown": "This is excellent. Thanks for sharing"
    },
    {
      "id": 1888314,
      "postDate": "2022-08-07T13:40:11.140Z",
      "content": "<p>Thank you. This is very useful 👍</p>",
      "rawMarkdown": "Thank you. This is very useful 👍"
    },
    {
      "id": 1885318,
      "postDate": "2022-08-05T05:08:07.880Z",
      "content": "<p>Thank you. This is very useful 👍</p>",
      "rawMarkdown": "Thank you. This is very useful 👍"
    },
    {
      "id": 1881383,
      "postDate": "2022-08-02T13:17:48.850Z",
      "content": "<p>Very helpful, Thank you!</p>",
      "rawMarkdown": "Very helpful, Thank you!"
    },
    {
      "id": 1880525,
      "postDate": "2022-08-01T19:37:07.977Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "rawMarkdown": "Thanks for sharing @thedevastator "
    },
    {
      "id": 1864373,
      "postDate": "2022-07-21T02:33:00.050Z",
      "content": "<p>very good,thanks for sharing</p>",
      "rawMarkdown": "very good,thanks for sharing"
    },
    {
      "id": 1862027,
      "postDate": "2022-07-19T12:06:02.943Z",
      "content": "<p>Nice post, thanks for sharing</p>",
      "rawMarkdown": "Nice post, thanks for sharing"
    },
    {
      "id": 1858651,
      "postDate": "2022-07-17T05:34:01.190Z",
      "content": "<p>Thanks for the tips</p>",
      "rawMarkdown": "Thanks for the tips"
    },
    {
      "id": 1856808,
      "postDate": "2022-07-15T16:02:23.300Z",
      "content": "<p>Wow! Very useful! Thank you!!!</p>",
      "rawMarkdown": "Wow! Very useful! Thank you!!!"
    },
    {
      "id": 1855356,
      "postDate": "2022-07-14T14:05:09.993Z",
      "content": "<p>Great summary, thanks for sharing!!</p>",
      "rawMarkdown": "Great summary, thanks for sharing!!"
    },
    {
      "id": 1851762,
      "postDate": "2022-07-11T14:05:58.470Z",
      "content": "<p>Thanks for the information</p>",
      "rawMarkdown": "Thanks for the information"
    },
    {
      "id": 1851541,
      "postDate": "2022-07-11T10:51:22.347Z",
      "content": "<p>Thank a lost. marked.</p>",
      "rawMarkdown": "Thank a lost. marked."
    }
  ],
  "comments": [
    {
      "id": 2325446,
      "author_name": "Koushikphy",
      "author_url": "",
      "post_date": "2023-07-01T10:55:05.263000",
      "content": "<p>This is really very helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2192622,
      "author_name": "ShantamVijayputra",
      "author_url": "",
      "post_date": "2023-03-22T19:27:42.413000",
      "content": "<p>This is awesome resource , Keep up the great work.👍👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2064823,
      "author_name": "Steve Chen",
      "author_url": "",
      "post_date": "2022-12-14T06:23:32.787000",
      "content": "<p>Fantastic resource for beginner, extremely grateful to you! <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2050768,
      "author_name": "Russell Healy",
      "author_url": "",
      "post_date": "2022-12-01T01:35:48.947000",
      "content": "<p>A fantastic resource all in one spot. Thank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1911840,
      "author_name": "jeannief",
      "author_url": "",
      "post_date": "2022-08-24T10:53:45.570000",
      "content": "<p>Thank you for taking the time to put this together. Very helpful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1897431,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-13T18:02:28.747000",
      "content": "<p>This is very useful </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1879277,
      "author_name": "dean1977",
      "author_url": "",
      "post_date": "2022-08-01T01:02:11.613000",
      "content": "<p>very helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1879239,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-07-31T23:54:45.377000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, so much knowledge, thanks for sharing ✨</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1864977,
      "author_name": "LIPING11",
      "author_url": "",
      "post_date": "2022-07-21T12:15:32.993000",
      "content": "<p>tttttthhhhhhhhhhhhaaaaaaaaaaaannnnnnnnnnnnnnnnnk</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1861816,
      "author_name": "pluto_21",
      "author_url": "",
      "post_date": "2022-07-19T09:04:36.197000",
      "content": "<p>Thank you so much for sharing these tips</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1860925,
      "author_name": "Lucia Constantino",
      "author_url": "",
      "post_date": "2022-07-18T16:49:56.227000",
      "content": "<p>Thank you so much for this! I'm barely starting with Kaggle competitions but I'm saving this one for later:)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1858544,
      "author_name": "Jonathan Bown",
      "author_url": "",
      "post_date": "2022-07-17T02:42:41.103000",
      "content": "<p>Thanks so much for putting this together!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1857191,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-07-15T23:56:53.627000",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, thanks for sharing a massive amount of information to keep improving</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1857146,
      "author_name": "SHANKAR T",
      "author_url": "",
      "post_date": "2022-07-15T22:09:31.913000",
      "content": "<p>Wow !!<br>\nThank you very much for the amount of time you have contributed to make this…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1855775,
      "author_name": "Professor Kgwadi",
      "author_url": "",
      "post_date": "2022-07-14T22:37:56.487000",
      "content": "<p>wow great stuff, thank you for sharing .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1855526,
      "author_name": "Bilal Suppal",
      "author_url": "",
      "post_date": "2022-07-14T16:56:47.500000",
      "content": "<p>Nice post, thanks for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1852313,
      "author_name": "Surdarla",
      "author_url": "",
      "post_date": "2022-07-12T01:52:29.773000",
      "content": "<p>This is remarkable summary!<br>\nGreat!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1849096,
      "author_name": "jxlijunhao",
      "author_url": "",
      "post_date": "2022-07-09T08:04:18.123000",
      "content": "<p>very helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1848827,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2022-07-08T23:57:45.223000",
      "content": "<p>Nice post, thanks for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1848622,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2022-07-08T19:02:11.103000",
      "content": "<p>You have done a aggregate of aggregates. Thanks somethings I was just looking for found here . I think you should put this in a github  repo.. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1897691,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T02:35:51.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1883154,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-03T16:02:13.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1848917,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-09T03:43:09.170000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1848302,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-08T14:01:53.160000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1848077,
      "author_name": "Hongbin Na",
      "author_url": "",
      "post_date": "2022-07-08T10:25:17.673000",
      "content": "<p>Wow! Very useful! Thank you!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2012013,
      "author_name": "Darrylkwok",
      "author_url": "",
      "post_date": "2022-11-01T02:37:48.443000",
      "content": "<p>Super useful! Thank you! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2006418,
      "author_name": "wcq_glhf",
      "author_url": "",
      "post_date": "2022-10-27T15:23:37.077000",
      "content": "<p>Thanks for the insight!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1992926,
      "author_name": "kt0011",
      "author_url": "",
      "post_date": "2022-10-18T02:48:47.557000",
      "content": "<p>Thanks for the insight!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913373,
      "author_name": "Sam Forrest",
      "author_url": "",
      "post_date": "2022-08-25T09:54:14.937000",
      "content": "<p>Thanks for the insight!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1910876,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-23T18:43:31.537000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1888314,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-07T13:40:11.140000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1885318,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-05T05:08:07.880000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1881383,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-02T13:17:48.850000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1880525,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-01T19:37:07.977000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1864373,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-21T02:33:00.050000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1862027,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-19T12:06:02.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1858651,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-17T05:34:01.190000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1856808,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-15T16:02:23.300000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1855356,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-14T14:05:09.993000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1851762,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-11T14:05:58.470000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1851541,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-11T10:51:22.347000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1848063": "# Tabular Classification - Tips and Tricks\n\n\nSharing a long list of tips and tricks found on previous tabular classification competitions on Kaggle.\nTake this as a list of ideas you can attempt to try for improving your score!\n\n> **Credits:** \n> - Many of the links are based on parts of [this](https://neptune.ai/blog/tabular-data-binary-classification-tips-and-tricks-from-5-kaggle-competitions) blog post (with edits and filtering).\n> - Including links from: [The best of kaggle - GBMs](https://www.kaggle.com/discussions/getting-started/277121).\n> - Also from: [The best of kaggle - Ensembles](https://www.kaggle.com/discussions/getting-started/291437).\n\n### Dealing with larger datasets\n\nAn issue close to our hearts! \n\n* Faster [data loading with pandas.](https://www.kaggle.com/c/home-credit-default-risk/discussion/59575)\n* Data compression techniques to [reduce the size of data by 70%](https://www.kaggle.com/nickycan/compress-70-of-dataset).\n* Optimize the memory by r[educing the size of some attributes.](https://www.kaggle.com/shrutimechlearn/large-data-loading-trick-with-ms-malware-data)\n* Use open-source libraries such as [Dask to read and manipulate the data](https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask), it performs parallel computing and saves up memory space.\n* Use [cudf](https://github.com/rapidsai/cudf).\n* Convert data to [parquet](https://arrow.apache.org/docs/python/parquet.html) format.\n* Converting data to [feather](https://medium.com/@snehotosh.banerjee/feather-a-fast-on-disk-format-for-r-and-python-data-frames-de33d0516b03) format.\n* Reducing memory usage for [optimizing RAM](https://www.kaggle.com/mjbahmani/reducing-memory-size-for-ieee).\n\n### Data exploration\n\nData exploration is a must on all competitions, Here are some references from past competitions use them to get some ideas on what exactly to explore.\n\n* EDA for microsoft [malware detection.](https://www.kaggle.com/youhanlee/my-eda-i-want-to-see-all)\n* Time Series [EDA for malware detection.](https://www.kaggle.com/cdeotte/time-split-validation-malware-0-68)\n* Complete [EDA for home credit loan prediction](https://www.kaggle.com/codename007/home-credit-complete-eda-feature-importance).\n* Complete [EDA for Santader prediction.](https://www.kaggle.com/gpreda/santander-eda-and-prediction)\n* EDA for [VSB Power Line Fault Detection.](https://www.kaggle.com/go1dfish/basic-eda)\n\n### Preprocessing\n\nBefore we start to engineer features, we need to do some preprocessing.\nThose invlolve removing missing values, converting categorical variables to numeric, and scaling the data, etc..\nHere are some references from past competitions use them to get some ideas for preprocessing.\n\n* Methods to [tackle class imbalance](https://www.kaggle.com/shahules/tackling-class-imbalance).\n* Data augmentation by [Synthetic Minority Oversampling Technique](https://machinelearningmastery.com/smote-oversampling-for-imbalanced-classification/).\n* Fast inplace [shuffle for augmentation](https://www.kaggle.com/jiweiliu/fast-inplace-shuffle-for-augmentation).\n* Finding [synthetic samples in the dataset.](https://www.kaggle.com/yag320/list-of-fake-samples-and-public-private-lb-split)\n* [Signal denoising](https://www.kaggle.com/jackvial/dwt-signal-denoising) used in signal processing competitions.\n* Finding [patterns of missing data](https://www.kaggle.com/jpmiller/patterns-of-missing-data).\n* Methods to handle [missing data](https://towardsdatascience.com/6-different-ways-to-compensate-for-missing-values-data-imputation-with-examples-6022d9ca0779).\n* An overview of various [encoding techniques for categorical data.](https://www.kaggle.com/shahules/an-overview-of-encoding-techniques)\n* Building [model to predict missing values.](https://www.kaggle.com/c/home-credit-default-risk/discussion/64598)\n* Random [shuffling of data](https://www.kaggle.com/brandenkmurray/randomly-shuffled-data-also-works) to create new synthetic training set.\n\n### Feature engineering\n\nNow we can start to engineer features, below you can find some references from past competitions use them to get some ideas to engineer features. \n\n* Target [encoding cross validation](https://medium.com/@pouryaayria/k-fold-target-encoding-dfe9a594874b/) for better encoding.\n* Entity embedding to [handle categories](https://www.kaggle.com/abhishek/entity-embeddings-to-handle-categories).\n* Encoding c[yclic features for deep learning.](https://www.kaggle.com/avanwyk/encoding-cyclical-features-for-deep-learning)\n* Manual [feature engineering methods](https://www.kaggle.com/willkoehrsen/introduction-to-manual-feature-engineering).\n* Automated feature engineering techniques [using featuretools](https://www.kaggle.com/willkoehrsen/automated-feature-engineering-basics).\n* Top hard crafted features used in [microsoft malware detection](https://www.kaggle.com/sanderf/7th-place-solution-microsoft-malware-prediction).\n* Denoising NN for [feature extraction](https://towardsdatascience.com/applied-deep-learning-part-3-autoencoders-1c083af4d798).\n* Feature engineering [using RAPIDS framework.](https://www.kaggle.com/cdeotte/rapids-feature-engineering-fraud-0-96/)\n* Things to remember while processing f[eatures using LGBM.](https://www.kaggle.com/c/ieee-fraud-detection/discussion/108575)\n* [Lag features and moving averages.](https://www.kaggle.com/c/home-credit-default-risk/discussion/64593)\n* [Principal component analysis](https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567) for dimensionality reduction.\n* LDA for [dimensionality reduction](https://medium.com/machine-learning-researcher/dimensionality-reduction-pca-and-lda-6be91734f567).\n* Best hand crafted LGBM features for [microsoft malware detection](https://www.kaggle.com/c/microsoft-malware-prediction/discussion/85157).\n* Generating [frequency features.](https://www.kaggle.com/philippsinger/frequency-features-without-test-data-information)\n* Dropping variables with [different train and test distribution.](https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix)\n* [Aggregate time series features](https://www.kaggle.com/c/home-credit-default-risk/discussion/64593) for home credit competition.\n* [Time Series](https://www.kaggle.com/c/home-credit-default-risk/discussion/64593) features used in home credit default risk.\n* Scale, Standardize and n[ormalize with sklearn](https://towardsdatascience.com/scale-standardize-or-normalize-with-scikit-learn-6ccc7d176a02).\n* Handcrafted features for [Home default risk competition.](https://www.kaggle.com/c/home-credit-default-risk/discussion/57750)\n* Handcrafted [features used in Santander Transaction Prediction.](https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/89070)\n\n### Feature selection\n\nNow that we got our features engineered, we should check if they are useful or not.\nTo do this, we need to do some feature selection.\nHere are some references from past competitions use them to get some ideas for feature selection.\n\n* Six ways to do [features selection using sklearn](https://www.kaggle.com/sz8416/6-ways-for-feature-selection).\n* [Permutation feature importance](https://www.kaggle.com/c/ieee-fraud-detection/discussion/107877#latest-635386).\n* [Adversarial feature validation](https://www.kaggle.com/tunguz/adversarial-ieee/).\n* Feature selection using [null importances.](https://www.kaggle.com/ogrellier/feature-selection-with-null-importances)\n* Tree explainer using [SHAP.](https://github.com/slundberg/shap)\n* DeepNN explainer using [SHAP](https://github.com/slundberg/shap).\n\n\n### Modeling\n\nBelow you can find [The best of kaggle - GBMs](https://www.kaggle.com/discussions/getting-started/277121).\nThese are the top voted/viewed notebooks all over kaggle for GBMs.  \n**Usage rules are simple: You like it? Upvote the original!**\n\n### XGBoost\n \n* [Data Analysis & XGBoost Starter (0.35460 LB)](https://www.kaggle.com/anokas/data-analysis-xgboost-starter-0-35460-lb) (137K views, 1377 Votes)*\n* [House prices: Lasso, XGBoost, and a detailed EDA](https://www.kaggle.com/erikbruin/house-prices-lasso-xgboost-and-a-detailed-eda) (183K views, 1317 Votes)*\n* [Feature engineering, xgboost](https://www.kaggle.com/dlarionov/feature-engineering-xgboost) (144K views, 942 Votes)*\n* [XGBoost](https://www.kaggle.com/dansbecker/xgboost) (173K views, 918 Votes)*\n* [XGBoost](https://www.kaggle.com/alexisbcook/xgboost) (266K views, 578 Votes)*\n* [Market Prediction: XGBoost with GPU (Fit in 1min)](https://www.kaggle.com/hamditarek/market-prediction-xgboost-with-gpu-fit-in-1min) (53K views, 411 Votes)*\n* [A Guide on XGBoost hyperparameters tuning](https://www.kaggle.com/prashant111/a-guide-on-xgboost-hyperparameters-tuning) (55K views, 358 Votes)*\n* [Titanic WCG+XGBoost [0.84688]](https://www.kaggle.com/cdeotte/titanic-wcg-xgboost-0-84688) (27K views, 314 Votes)*\n* [XGBoost & Feature Selection DSBowl 🥣 🥣](https://www.kaggle.com/shahules/xgboost-feature-selection-dsbowl) (36K views, 299 Votes)*\n* [Understanding XGBoost Model on Otto Data](https://www.kaggle.com/tqchen/understanding-xgboost-model-on-otto-data) (130K views, 297 Votes)*\n\n### LightGBM\n \n* [LightGBM with Simple Features](https://www.kaggle.com/jsaguiar/lightgbm-with-simple-features) (85K views, 619 Votes)*\n* [LightGBM. Baseline Model Using Sparse Matrix](https://www.kaggle.com/bogorodvo/lightgbm-baseline-model-using-sparse-matrix) (49K views, 409 Votes)*\n* [LightGBM starter with feature engineering idea](https://www.kaggle.com/tommy1028/lightgbm-starter-with-feature-engineering-idea) (16K views, 370 Votes)*\n* [Google Analytics EDA + LightGBM + Screenshots](https://www.kaggle.com/erikbruin/google-analytics-eda-lightgbm-screenshots) (40K views, 344 Votes)*\n* [Feature Engineering & LightGBM](https://www.kaggle.com/davidcairuz/feature-engineering-lightgbm) (26K views, 272 Votes)*\n* [Simple LightGBM without blending](https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending) (18K views, 263 Votes)*\n* [LightGBM + XGBoost + Catboost](https://www.kaggle.com/samratp/lightgbm-xgboost-catboost) (37K views, 245 Votes)*\n* [ASHRAE- KFold LightGBM - without leak (1.08)](https://www.kaggle.com/aitude/ashrae-kfold-lightgbm-without-leak-1-08) (15K views, 240 Votes)*\n* [LightGBM Classifier in Python](https://www.kaggle.com/prashant111/lightgbm-classifier-in-python) (52K views, 231 Votes)*\n* [LightGBM (Fixing unbalanced data)](https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-lb-0-9680) (56K views, 228 Votes)*\n\n### CatBoost\n \n* [A new baseline for DSB 2019 - Catboost model](https://www.kaggle.com/mhviraf/a-new-baseline-for-dsb-2019-catboost-model) (15K views, 316 Votes)*\n* [LightGBM + XGBoost + Catboost](https://www.kaggle.com/samratp/lightgbm-xgboost-catboost) (37K views, 245 Votes)*\n* [CatBoost: A Deeper Dive](https://www.kaggle.com/abhinand05/catboost-a-deeper-dive) (8K views, 178 Votes)*\n* [House Prices Tutorial with Catboost](https://www.kaggle.com/allunia/house-prices-tutorial-with-catboost) (12K views, 176 Votes)*\n* [Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM](https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm) (25K views, 173 Votes)*\n* [WiDSDatathon2021-Catboost-Starter](https://www.kaggle.com/usharengaraju/widsdatathon2021-catboost-starter) (6K views, 167 Votes)*\n* [Tutorial: CatBoost Overview](https://www.kaggle.com/mitribunskiy/tutorial-catboost-overview) (39K views, 127 Votes)*\n* [CatBoost Classifier in Python](https://www.kaggle.com/prashant111/catboost-classifier-in-python) (28K views, 126 Votes)*\n* [Catboost - Some more features](https://www.kaggle.com/braquino/catboost-some-more-features) (7K views, 103 Votes)*\n* [mlcourse.ai. Fall 2019. Catboost starter](https://www.kaggle.com/kashnitsky/mlcourse-ai-fall-2019-catboost-starter) (18K views, 98 Votes)*\n\n\n### Ensemble\n\nNow that we have our model trained, we should ensemble our predictions.\n\nBelow you can find [The best of kaggle - Ensembles](https://www.kaggle.com/discussions/getting-started/291437).\nThese are the top voted/viewed notebooks all over kaggle for GBMs.\n**Usage rules are simple: You like it? Upvote the original!**\n\n### Ensembling Techniques\n\n* [Introduction to Ensembling/Stacking in Python](https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python) (620K views, 5436 Votes)*\n* [Titanic Top 4% with ensemble modeling](https://www.kaggle.com/yassineghouzam/titanic-top-4-with-ensemble-modeling) (185K views, 2501 Votes)*\n* [EDA & Ensemble Model (Top 10 Percentile)](https://www.kaggle.com/viveksrinivasan/eda-ensemble-model-top-10-percentile) (98K views, 513 Votes)*\n* [Analysis of Melanoma Metadata and EffNet Ensemble](https://www.kaggle.com/datafan07/analysis-of-melanoma-metadata-and-effnet-ensemble) (24K views, 405 Votes)*\n* [Detailed Data Analysis & Ensemble Modeling](https://www.kaggle.com/tannercarbonati/detailed-data-analysis-ensemble-modeling) (55K views, 360 Votes)*\n* [Ensemble Folds with MEDIAN - [0.153]](https://www.kaggle.com/cdeotte/ensemble-folds-with-median-0-153) (9K views, 303 Votes)*\n* [WBF approach for ensemble](https://www.kaggle.com/shonenkov/wbf-approach-for-ensemble) (18K views, 282 Votes)*\n* [Ensemble: Resnext50\\_32x4d + Efficientnet = 0.903](https://www.kaggle.com/japandata509/ensemble-resnext50-32x4d-efficientnet-0-903) (16K views, 274 Votes)*\n* [Employee attrition via Ensemble tree-based methods](https://www.kaggle.com/arthurtok/employee-attrition-via-ensemble-tree-based-methods) (64K views, 253 Votes)*\n* [Top 3%. Efficient ensembling in few lines of code](https://www.kaggle.com/isaienkov/top-3-efficient-ensembling-in-few-lines-of-code) (30K views, 242 Votes)*\n* [Ensemble Learning Techniques Tutorial](https://www.kaggle.com/pavansanagapati/ensemble-learning-techniques-tutorial) (55K views, 215 Votes)*\n* [Minimal LSTM + NB-SVM baseline ensemble](https://www.kaggle.com/jhoward/minimal-lstm-nb-svm-baseline-ensemble) (30K views, 209 Votes)*\n* [Ensemble Model: Stacked Model Example](https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example) (68K views, 207 Votes)*\n* [🤗🤗 Ensemble 🤗🤗](https://www.kaggle.com/hamditarek/ensemble) (64K views, 196 Votes)*\n* [EDA&Modelling of the External Data Inc. Ensemble](https://www.kaggle.com/datafan07/eda-modelling-of-the-external-data-inc-ensemble) (8K views, 195 Votes)*\n\n### Blending\n \n* [Simple LightGBM without blending](https://www.kaggle.com/mfjwr1/simple-lightgbm-without-blending) (18K views, 265 Votes)*\n* [Blending tensorflow and pytorch🔥🔥🔥](https://www.kaggle.com/a763337092/blending-tensorflow-and-pytorch) (13K views, 220 Votes)*\n* [Another model for your blending](https://www.kaggle.com/speedwagon/quadratic-discriminant-analysis) (11K views, 216 Votes)*\n* [Optimise Blending Weights with Bonus :0](https://www.kaggle.com/gogo827jz/optimise-blending-weights-with-bonus-0) (7K views, 184 Votes)*\n* [Blending with Linear Regression [0.688 LB]](https://www.kaggle.com/suicaokhoailang/blending-with-linear-regression-0-688-lb) (14K views, 136 Votes)*\n* [blending blending blending](https://www.kaggle.com/abhishek/blending-blending-blending) (3K views, 128 Votes)*\n* [Lessons from Toxic : Blending is the new sexy](https://www.kaggle.com/jagangupta/lessons-from-toxic-blending-is-the-new-sexy) (10K views, 128 Votes)*\n* [non-blending lightGBM model LB: 0.977](https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977) (16K views, 127 Votes)*\n* [Bayesian Optimization Seed Blending](https://www.kaggle.com/hengzheng/bayesian-optimization-seed-blending) (9K views, 118 Votes)*\n* [elo\\_world\\_high\\_score\\_without\\_blending](https://www.kaggle.com/gpreda/elo-world-high-score-without-blending) (8K views, 112 Votes)*\n* [competition part-5: blending 101](https://www.kaggle.com/abhishek/competition-part-5-blending-101) (3K views, 105 Votes)*\n* [Beginner EDA with Feature Eng. and Blending Models](https://www.kaggle.com/datafan07/beginner-eda-with-feature-eng-and-blending-models) (3K views, 99 Votes)*\n* [Cross-validation, weighted linear blending, errors](https://www.kaggle.com/tilii7/cross-validation-weighted-linear-blending-errors) (7K views, 84 Votes)*\n* [Blending TensorFlow + PyTorch (th=0.4914)](https://www.kaggle.com/sagarjiyani/blending-tensorflow-pytorch-th-0-4914) (5K views, 82 Votes)*\n* [Blending NN and LGBM/RF](https://www.kaggle.com/gogo827jz/blending-nn-and-lgbm-rf) (6K views, 81 Votes)*\n\n### Stacking\n \n* [Stacked Regressions : Top 4% on LeaderBoard](https://www.kaggle.com/serigne/stacked-regressions-top-4-on-leaderboard) (493K views, 6236 Votes)*\n* [Introduction to Ensembling/Stacking in Python](https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python) (620K views, 5436 Votes)*\n* [Model stacking, feature engineering and EDA](https://www.kaggle.com/dimitreoliveira/model-stacking-feature-engineering-and-eda) (43K views, 409 Votes)*\n* [Stacking House Prices - Walkthrough to Top 5%](https://www.kaggle.com/agodwinp/stacking-house-prices-walkthrough-to-top-5) (32K views, 227 Votes)*\n* [Stacking Starter](https://www.kaggle.com/mmueller/stacking-starter) (41K views, 221 Votes)*\n* [Ensemble Model: Stacked Model Example](https://www.kaggle.com/jimthompson/ensemble-model-stacked-model-example) (68K views, 207 Votes)*\n* [Titanic: Voting, Pipeline, Stack, and Guide](https://www.kaggle.com/nicapotato/titanic-voting-pipeline-stack-and-guide) (21K views, 197 Votes)*\n* [Explore Stacking (LB 0.1463)](https://www.kaggle.com/dongxu027/explore-stacking-lb-0-1463) (18K views, 180 Votes)*\n* [Stacking Test-Sklearn, XGBoost, CatBoost, LightGBM](https://www.kaggle.com/eliotbarr/stacking-test-sklearn-xgboost-catboost-lightgbm) (26K views, 176 Votes)*\n* [Top 1% Approach: EDA, New Models and Stacking](https://www.kaggle.com/datafan07/top-1-approach-eda-new-models-and-stacking) (9K views, 172 Votes)*\n* [Simple Stacker LB 0.284](https://www.kaggle.com/yekenot/simple-stacker-lb-0-284) (20K views, 172 Votes)*\n* [Stack&Blend LRs XGB LGB {House Prices K} v17](https://www.kaggle.com/itslek/stack-blend-lrs-xgb-lgb-house-prices-k-v17) (6K views, 164 Votes)*\n* [Top 10 (0.10943): stacking, MICE and brutal force](https://www.kaggle.com/agehsbarg/top-10-0-10943-stacking-mice-and-brutal-force) (21K views, 156 Votes)*\n* [OOF stacking regime](https://www.kaggle.com/hhstrand/oof-stacking-regime) (14K views, 145 Votes)*\n* [EloDA with Feature Engineering and Stacking](https://www.kaggle.com/tunguz/eloda-with-feature-engineering-and-stacking) (10K views, 141 Votes)*\n\n### Bagging\n \n* [Riiid LGBM bagging2 + SAKT =0.781](https://www.kaggle.com/ammarnassanalhajali/riiid-lgbm-bagging2-sakt-0-781) (11K views, 188 Votes)*\n* [SAKT + Riiid LGBM bagging2](https://www.kaggle.com/leadbest/sakt-riiid-lgbm-bagging2) (8K views, 139 Votes)*\n* [Topic 5. Ensembles. Part 1. Bagging](https://www.kaggle.com/kashnitsky/topic-5-ensembles-part-1-bagging) (16K views, 137 Votes)*\n* [Bagging vs Boosting](https://www.kaggle.com/prashant111/bagging-vs-boosting) (22K views, 117 Votes)*\n* [Fork of Riiid LGBM bagging2.1 471152](https://www.kaggle.com/julianguo/fork-of-riiid-lgbm-bagging2-1-471152) (8K views, 117 Votes)*\n* [Riiid! LGBM bagging2](https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2) (6K views, 113 Votes)*\n* [Predicting House Prices [XGB/RF/Bagging-Reg Pipe]](https://www.kaggle.com/kabure/predicting-house-prices-xgb-rf-bagging-reg-pipe) (10K views, 89 Votes)*\n* [Keras starter with bagging (LB: 1120.596)](https://www.kaggle.com/danijelk/keras-starter-with-bagging-lb-1120-596) (18K views, 77 Votes)*\n* [Keras starter with bagging 1111.84364](https://www.kaggle.com/mtinti/keras-starter-with-bagging-1111-84364) (19K views, 77 Votes)*\n* [Riiid LGBM bagging2.1](https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2-1) (4K views, 62 Votes)*\n* [BaggingRegressor + RAPIDS Ensemble](https://www.kaggle.com/andypenrose/baggingregressor-rapids-ensemble) (4K views, 59 Votes)*\n* [Credit analysis with KNN/DTree/RF/Bagging/ANN🏦](https://www.kaggle.com/fahadmehfoooz/credit-analysis-with-knn-dtree-rf-bagging-ann) (<1K views, 55 Votes)*\n* [Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM](https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm) (1K views, 51 Votes)*\n* [Ensemble ML Algorithms : Bagging, Boosting, Voting](https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting) (8K views, 48 Votes)*\n* [Simple XGBoost CV Bagging](https://www.kaggle.com/cttsai/simple-xgboost-cv-bagging) (3K views, 42 Votes)*\n\n### Boosting\n \n* [Boosting creativity towards feature engineering](https://www.kaggle.com/felipemello/boosting-creativity-towards-feature-engineering) (9K views, 229 Votes)*\n* [Vehicle Insurance EDA and boosting models](https://www.kaggle.com/yashvi/vehicle-insurance-eda-and-boosting-models) (19K views, 200 Votes)*\n* [Topic 10. Gradient Boosting](https://www.kaggle.com/kashnitsky/topic-10-gradient-boosting) (21K views, 188 Votes)*\n* [Gradient boosting simplified](https://www.kaggle.com/grroverpr/gradient-boosting-simplified) (48K views, 147 Votes)*\n* [Bagging vs Boosting](https://www.kaggle.com/prashant111/bagging-vs-boosting) (22K views, 117 Votes)*\n* [Diabetes Prediction (Stacking + Boosting)](https://www.kaggle.com/niteshyadav3103/diabetes-prediction-stacking-boosting) (1K views, 85 Votes)*\n* [Elo Stack With Goss Boosting](https://www.kaggle.com/roydatascience/elo-stack-with-goss-boosting) (6K views, 78 Votes)*\n* [Assignment 10. Gradient boosting and flight delays](https://www.kaggle.com/kashnitsky/assignment-10-gradient-boosting-and-flight-delays) (9K views, 71 Votes)*\n* [Battle of the Boosting Algos: LGB, XGB, Catboost](https://www.kaggle.com/lavanyashukla01/battle-of-the-boosting-algos-lgb-xgb-catboost) (9K views, 55 Votes)*\n* [Introduction to Boosting using LGBM(LB: 0.68357)](https://www.kaggle.com/vinnsvinay/introduction-to-boosting-using-lgbm-lb-0-68357) (14K views, 52 Votes)*\n* [FE-Pipeline with HistGradientBoostingRegressor](https://www.kaggle.com/tunguz/fe-pipeline-with-histgradientboostingregressor) (3K views, 51 Votes)*\n* [Sentiment Analysis:SVM/NB/Bagging/Boosting/RF/LSTM](https://www.kaggle.com/fahadmehfoooz/sentiment-analysis-svm-nb-bagging-boosting-rf-lstm) (1K views, 51 Votes)*\n* [Explainable Boosting machines for Tabular data](https://www.kaggle.com/parulpandey/explainable-boosting-machines-for-tabular-data) (2K views, 49 Votes)*\n* [Ensemble ML Algorithms : Bagging, Boosting, Voting](https://www.kaggle.com/faressayah/ensemble-ml-algorithms-bagging-boosting-voting) (8K views, 48 Votes)*\n* [House Price Prediction : GBoosting,AdaBoost etc.](https://www.kaggle.com/sid321axn/house-price-prediction-gboosting-adaboost-etc) (8K views, 42 Votes)*\n\n\n### Hyperparameters Tuning\n\n* LGBM [hyperparameter tuning](https://www.kaggle.com/mlisovyi/lightgbm-hyperparameter-optimisation-lb-0-761) methods.\n* Automated [model tuning](https://www.kaggle.com/willkoehrsen/automated-model-tuning) methods.\n* Parameter tuning with [hyper plot](https://www.kaggle.com/bigironsphere/parameter-tuning-in-one-function-with-hyperopt).\n* [Bayesian optimization](http://krasserm.github.io/2018/03/21/bayesian-optimization/) for hyperparameter tuning.\n* [Gpyopt Hyperparameter Optimisation](https://www.kaggle.com/nicapotato/gpyopt-hyperparameter-optimisation-gpu-lgbm).",
    "2325446": "This is really very helpful.",
    "2192622": "This is awesome resource , Keep up the great work.👍👍",
    "2064823": "Fantastic resource for beginner, extremely grateful to you! @thedevastator ",
    "2050768": "A fantastic resource all in one spot. Thank you",
    "1911840": "Thank you for taking the time to put this together. Very helpful!",
    "1897431": "This is very useful ",
    "1879277": "very helpful.",
    "1879239": "Hello @thedevastator, so much knowledge, thanks for sharing ✨",
    "1864977": "tttttthhhhhhhhhhhhaaaaaaaaaaaannnnnnnnnnnnnnnnnk",
    "1861816": "Thank you so much for sharing these tips",
    "1860925": "Thank you so much for this! I'm barely starting with Kaggle competitions but I'm saving this one for later:)",
    "1858544": "Thanks so much for putting this together!",
    "1857191": "Hello, @thedevastator, thanks for sharing a massive amount of information to keep improving",
    "1857146": "Wow !!\nThank you very much for the amount of time you have contributed to make this...",
    "1855775": "wow great stuff, thank you for sharing .",
    "1855526": "Nice post, thanks for sharing @thedevastator",
    "1852313": "This is remarkable summary!\nGreat!",
    "1849096": "very helpful.",
    "1848827": "Nice post, thanks for sharing @thedevastator ",
    "1848622": "You have done a aggregate of aggregates. Thanks somethings I was just looking for found here . I think you should put this in a github  repo.. ",
    "1897691": "",
    "1883154": "",
    "1848917": "",
    "1848302": "",
    "1848077": "Wow! Very useful! Thank you!!!",
    "2012013": "Super useful! Thank you! ",
    "2006418": "Thanks for the insight!",
    "1992926": "Thanks for the insight!",
    "1913373": "Thanks for the insight!",
    "1910876": "This is excellent. Thanks for sharing",
    "1888314": "Thank you. This is very useful 👍",
    "1885318": "Thank you. This is very useful 👍",
    "1881383": "Very helpful, Thank you!",
    "1880525": "Thanks for sharing @thedevastator ",
    "1864373": "very good,thanks for sharing",
    "1862027": "Nice post, thanks for sharing",
    "1858651": "Thanks for the tips",
    "1856808": "Wow! Very useful! Thank you!!!",
    "1855356": "Great summary, thanks for sharing!!",
    "1851762": "Thanks for the information",
    "1851541": "Thank a lost. marked."
  }
}