{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Lesson: how to overfit on public LB and get an epic fail on private test data  \n\nOur great fall:  \nPublic LB 2nd place -> private LB 173, 171 position down\n\nMain mistake:  \nwe had no correct CV scheme and / or feature selection to predict unseen day\n\nSo, our model was good to predict unseen donor but not so good to predict unseen day\n\n\nHow did I get good public score on CITE part:  \n1) Data\n- Use several datasets (original data; raw counts data with log1p transformation; CITE feature shop & denoised data by Alexander C. & company)\n- PCA for the whole dataset. N_components determined by loop with RidgeCV model\n- Unchanged features selection by premutation feature importance with Ridge/RidgeCV model\n\n2) Several models ensebling:\n- Full-connected NN (or MLP), idea is taken from https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\n- Convolutional NN, idea is taken from https://www.kaggle.com/competitions/lish-moa/discussion/202256\n- CatBoost. First attempts were made in Multiregression mode, but CPU calculation was too slow. Later switched to regular 140 separate target regressions with GPU calculation.\n- LightGBM, 140 separate target regressions with CPU.\n- RidgeCV  \nAlso were tried but results were bad: KernelRidge, SVM, TabnetRegressor, MLP from sklearn\n\n3) Wide usage of OPTUNA:\n- selection of NN parameters  \n  1st step: number of layers and units in it  \n  2nd step: dropout and regularization values  \n  3rd step (optional): learning rate and batch size  \n- selection of LGB parameters (for each target)\n\n4) Blending with OPTUNA  \nBLEND = K1 * NN_oof + K2 * CNN_oof + K3 * CB_oof + K4 * LGB_oof + K5 * RIDGE_oof  \nscore = correlation_metric(targets, BLEND)  \nscore -> maximize (K1...K5 are selected by OPTUNA)","metadata":{}}]}