{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# MSCI22 keep in mind: 2in1\n\nMental note to myself: keep in mind that this is a 2 in 1 competition. I mean, we are looking for two models (CITE model and MULTIOME model) and these models must be evaluated in isolation.\n\n## How?\n\nAccording the competition evaluation:\n\n\n\\begin{equation}\nLB = \\frac{\\sum \\rho}{T}\n\\end{equation}\n\n\\begin{equation}\nLB = \\frac{\\sum \\rho_C + \\sum \\rho_M}{T}\n\\end{equation}\n\n\\begin{equation}\nLB = \\frac{C}{T} LB_C+ \\frac{M}{T} LB_M\n\\end{equation}\n\n\n* T = Total cells\n* C = Total CITE cells\n* M = Total MULTIOME cells\n\nAs a result of this we can evaluate a MULTIOME model setting to 0 the CITE model predictions because then the CITE model LB must be -1:\n\n\n\\begin{equation}\nLB_M = \\frac{LB + C/T}{M/T} \n\\end{equation}\n\n\n\n\n## Practical example\n\nThe \"[Blending of five submission by @swimmy](https://www.kaggle.com/code/swimmy/blending-of-five-submission)\" notebook submission score 0.810 (top 90 in LB today). On the other hand, \"[Normalized Ensembles for Pearson's r by @vslaykovsky](https://www.kaggle.com/code/vslaykovsky/lb-0-811-normalized-ensembles-for-pearson-s-r/notebook?scriptVersionId=105258966)\" score 0.811 (top 58 in LB today). Then the second ensemble model is better than the first blending model. \n\nBut, evaluating the MULTIOME model in isolation, the first model score LB=-0.438 and the second model LB=-0.439: the firt one is better. If we mix the two submission selecting the first MULTIOME model and the second CITE model then score 0.811 and top 48 in LB.\n","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-09-17T10:56:21.500593Z","iopub.execute_input":"2022-09-17T10:56:21.501301Z","iopub.status.idle":"2022-09-17T10:56:21.529362Z","shell.execute_reply.started":"2022-09-17T10:56:21.501179Z","shell.execute_reply":"2022-09-17T10:56:21.528428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EDGE=48663*140","metadata":{"execution":{"iopub.status.busy":"2022-09-17T10:56:21.531378Z","iopub.execute_input":"2022-09-17T10:56:21.532117Z","iopub.status.idle":"2022-09-17T10:56:21.537378Z","shell.execute_reply.started":"2022-09-17T10:56:21.532071Z","shell.execute_reply":"2022-09-17T10:56:21.536210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CITE = '../input/open-problems-public-submissions/submission.csv' # https://www.kaggle.com/code/vslaykovsky/lb-0-811-normalized-ensembles-for-pearson-s-r/notebook?scriptVersionId=105258966\nMULTIOME = '../input/open-problems-public-submissions/5in1_ensemble_105868499.csv' # https://www.kaggle.com/code/swimmy/blending-of-five-submission","metadata":{"execution":{"iopub.status.busy":"2022-09-17T10:56:21.538664Z","iopub.execute_input":"2022-09-17T10:56:21.539569Z","iopub.status.idle":"2022-09-17T10:56:21.548431Z","shell.execute_reply.started":"2022-09-17T10:56:21.539537Z","shell.execute_reply":"2022-09-17T10:56:21.547263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv(CITE)\nmultiome = pd.read_csv(MULTIOME)\nsubmission.iloc[EDGE:,1] = multiome.iloc[EDGE:,1]\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-17T10:56:21.551261Z","iopub.execute_input":"2022-09-17T10:56:21.551831Z"},"trusted":true},"execution_count":null,"outputs":[]}]}