{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"A good solution to this competition will most certainly require ensembling. But what are some good ways of ensembling predictions?\n\nIn this notebook, we will look at two approaches:\n* voting ensemble\n* voting ensemble with weights (this allows you to put more weight on predictions that got a better validation/LB score)\n\nAs always, the challenge will be the resources that we have available. With each submissions file at over 5 million rows, each row containing 20 predictions, the proble of available RAM is non-trivial!\n\nTo combat this, we will use the very memory efficient `polars` 🙂\n\nAs a basis for our work, let us use the following three submissions:\n\n* [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) [0.575] by Chris Deotte\n* [Test Dataset Is All We Need?](https://www.kaggle.com/code/tomooinubushi/test-dataset-is-all-we-need/notebook) [0.522] by Tomoo Inubushi\n* [💡Matrix Factorization [PyTorch+Merlin Dataloader]](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader/notebook) [0.493] by yours truly\n\nLet's get started!\n\n**Please upvote if you like his notebook 🙏 It would be of great help to me if you do. Thank you!**\n\n*Please note: In this notebook, we are ensembling 1 good solution with 2 that are not that great, hence we can't expect great results with equal weights. Even when setting the weights to something that is reasonable given the performance of each solution, we still cannot expect a very good result.*\n\n*However, when I used this method locally on my own submissions, I was able to combine several solutions generated with the same ranking model (by varying the seed) to improve my LB score from 0.576 to 0.577. This effect can be even stronger when ensembling more varied solutions.*","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Loading the data","metadata":{}},{"cell_type":"code","source":"!pip install polars # why are we using polars? it has much smaller memory footprint than pandas!","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:35:03.529097Z","iopub.execute_input":"2022-11-27T13:35:03.530619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import polars as pl","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are the submissions that we will use. We order the file paths from best performing to the worst.","metadata":{}},{"cell_type":"code","source":"paths = ['../input/otto-submissions-for-ensembling/submission_rerank_0575.csv', '../input/otto-submissions-for-ensembling/submission_test_all_0522.csv', '../input/otto-submissions-for-ensembling/submission_matrix_factorization_0493.csv']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can load all the submissions at once, but we have to be very careful about what operations we run on the data as it is very simple to run out of RAM.","metadata":{}},{"cell_type":"code","source":"def read_sub(path, weight=1): # by default let us assing the weight of 1 to predictions from each submission, this will be akin to a standard vote ensemble\n    '''a helper function for loading and preprocessing submissions'''\n    return (\n        pl.read_csv(path)\n            .with_column(pl.col('labels').str.split(by=' '))\n            .with_column(pl.lit(weight).alias('vote'))\n            .explode('labels')\n            .rename({'labels': 'aid'})\n            .with_column(pl.col('aid').cast(pl.UInt32)) # we are casting the `aids` to `Int32`! memory management is super important to ensure we don't run out of resources\n            .with_column(pl.col('vote').cast(pl.UInt8))\n    )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Loading all the data at once.","metadata":{}},{"cell_type":"code","source":"subs = [read_sub(path) for path in paths]\nsubs[0].head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Concatenating and grouping won't work due to memory requirements. Our only option are the very efficient joins.","metadata":{}},{"cell_type":"code","source":"subs = subs[0].join(subs[1], how='outer', on=['session_type', 'aid']).join(subs[2], how='outer', on=['session_type', 'aid'], suffix='_right2')\nsubs.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us fill in the `nulls`, sum the votes, and order the predictions so that predictions with more votes appear first.","metadata":{}},{"cell_type":"code","source":"subs = (subs\n    .fill_null(0)\n    .with_column((pl.col('vote') + pl.col('vote_right') + pl.col('vote_right2')).alias('vote_sum'))\n    .drop(['vote', 'vote_right', 'vote_right2'])\n    .sort(by='vote_sum')\n    .reverse()\n)\n\nsubs.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All we have to do now is take the first 20 predictions per `session_type` and turn them into a space seperated string.","metadata":{}},{"cell_type":"code","source":"preds = subs.groupby('session_type').agg([\n    pl.col('aid').head(20).alias('labels')\n])\n\npreds = preds.with_column(pl.col('labels').apply(lambda lst: ' '.join([str(aid) for aid in lst])))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have created a standard voting ensemble and are now ready to output the submission file!","metadata":{}},{"cell_type":"code","source":"%%time\n\npreds.write_csv('submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Voting ensemble is often a great way to go. However, sometimes we might want to weight our submissions. Say, we want to give more weight to the submission that performs better.\n\nHow would we do it?\n\nWe already have all the pieces 🙂\n\nWhen reading the submissions, all you have to do is specify the weight associated with each one using the `read_sub` function, for instance we could do something like this:\n\n`subs = [read_sub(path, weight) for path, weight in zip(paths, [1, 0.55, 0.55])]`\n\nAnd that's it!","metadata":{}},{"cell_type":"markdown","source":"## Summary\n\nWe now have a way to perfom voting ensemble (including using custom weights) even within the limits of a Kaggle VM! Ensembling will certainly be a major component of strong submissions.\n\n**If you enjoyed this notebook, please upvote! 🙏 Thank you!**\n\nThank you for reading, happy Kaggling! 🙂","metadata":{}}]}