{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"},{"sourceId":7104571,"sourceType":"datasetVersion","datasetId":4095742},{"sourceId":153760549,"sourceType":"kernelVersion"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Update: \nWe switched to full reproducibility on Kaggle getting the input from our previous ensembling notebook, the performance is slightly reduced from our real final submission as we could not reconstruct every step to 100% but this way we have no blends with models not trained by our team and everything is in our public pipeline.\n\nWe have 3 types of models:\n* [Neural Network](https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net)\n* [LGBM](https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation)\n* [NLP](https://www.kaggle.com/code/kishanvavdara/nlp-regression)\n\nThat are ensembled in:\n* [Ensembling](https://www.kaggle.com/raki21/4th-place-ensembling)\n\nand finally passed to this:\n* [Postprocessing](https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing)","metadata":{}},{"cell_type":"markdown","source":"# Idea","metadata":{}},{"cell_type":"markdown","source":"    We will show that many compound gene combinations are close to 0-mean random distributions.\n    \n    Then we will apply postprocessing that reduces the magnitude of those predictions alot.\n\n    That enables us to increase magnitude of the large variance compounds,\n    we ensembled and tuned the magnitude of predictions before, but having models that overfitted on noise in the large amount of low variance compounds lead to our optimum magnitude being much too low for the higher variance compounds, after increasing the magnitude our score improved alot! \n\n    Because we found this trick only on the last day we could only make a single prediction with applying a MULT=1.1, leading to \n    0.558/0.733 and 4th place. \n    \n    Having just increased the MULT by a bit we get an optimal multiplier for public test lead to optimal MULT=1.2, giving us\n    0.557/0.724 \n    \n    or MULT=1.3 with \n    0.558/0.718.\n    \n    Technically increasing MULT further to 1.5 would have given private test of \n    0.565/0.712, but we would never have selected that in the first place because of the public test performance decreasing already by alot.\n    \n    I also made a single modified submission with the publicly shared 3rd place solution, just adding the 10 lines of post processing decrease public score by 4 points and private by 8 (try it yourself! https://www.kaggle.com/code/jankowalski2000/3rd-place-solution I'll share what to add in the comments here).","metadata":{}},{"cell_type":"code","source":"MULT = 1.2","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:00:46.835304Z","iopub.execute_input":"2023-12-05T18:00:46.835940Z","iopub.status.idle":"2023-12-05T18:00:46.892637Z","shell.execute_reply.started":"2023-12-05T18:00:46.835881Z","shell.execute_reply":"2023-12-05T18:00:46.891011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Imports of Libraries and Previous Results","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:00:46.895311Z","iopub.execute_input":"2023-12-05T18:00:46.896040Z","iopub.status.idle":"2023-12-05T18:00:47.365224Z","shell.execute_reply.started":"2023-12-05T18:00:46.895990Z","shell.execute_reply":"2023-12-05T18:00:47.364142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import_ensemble = pd.read_csv('/kaggle/input/single-cell-4th-before-postprocess/560-743-ensemble.csv', index_col='id')\nimport_ensemble = pd.read_csv('/kaggle/input/4th-place-ensembling/submission.csv', index_col='id')","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:00:47.366978Z","iopub.execute_input":"2023-12-05T18:00:47.367861Z","iopub.status.idle":"2023-12-05T18:00:54.020979Z","shell.execute_reply.started":"2023-12-05T18:00:47.367821Z","shell.execute_reply":"2023-12-05T18:00:54.019225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv', index_col='id')\ncol = list(submission.columns)\nsubmission[col] = import_ensemble[col]","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:00:54.024249Z","iopub.execute_input":"2023-12-05T18:00:54.024653Z","iopub.status.idle":"2023-12-05T18:01:09.562270Z","shell.execute_reply.started":"2023-12-05T18:00:54.024616Z","shell.execute_reply":"2023-12-05T18:01:09.559835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualization of Randomness in Low Variance Compounds","metadata":{}},{"cell_type":"code","source":"def mrrmse_np(y_pred, y_true):\n    return np.sqrt(np.square(y_true - y_pred).mean())\n\ndef cell_type_from_df(cell_type, compounds):\n    cells_from_cell_type = df[df.cell_type == cell_type]\n    cells_compounds_train = cells_from_cell_type[cells_from_cell_type.sm_name.isin(compounds)]\n    cells_num = cells_compounds_train.drop(columns=['sm_name', 'cell_type', 'sm_lincs_id', 'SMILES']).to_numpy()\n    return cells_num\n\ndef plot_de_distr(cells, compound):\n    # Create a 3x5current green card lottery eligible countries subplot\n    fig, axes = plt.subplots(3, 5, figsize=(15, 9))\n    axes = axes.ravel() \n\n    for i, ax in enumerate(axes):\n        mean = cells[i].mean()\n        mrrmse = mrrmse_np(cells[i], cells[i]*0)\n        lower_bound = np.percentile(cells[i], 0.5)\n        upper_bound = np.percentile(cells[i], 99.5)\n        ax.hist(cells[i], bins=200, range=(lower_bound, upper_bound))\n        ax.set_title(f\"{list(compounds)[i]} \\n Mean: {mean:.2f}, MRRMSE: {mrrmse:.2f} \")\n\n    plt.tight_layout()\n    plt.show()\n    print('\\n')\n\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-12-05T18:01:09.564728Z","iopub.execute_input":"2023-12-05T18:01:09.565100Z","iopub.status.idle":"2023-12-05T18:01:09.576897Z","shell.execute_reply.started":"2023-12-05T18:01:09.565067Z","shell.execute_reply":"2023-12-05T18:01:09.575541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize B cell spread for compounds\ndf = pd.read_parquet('/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet')\ndf = df[df.control == 0]\n\ncompounds = df[df.cell_type == 'B cells'].sm_name\n\nfor cell_type in ['B cells', 'Myeloid cells', 'NK cells', 'T cells CD4+', 'T regulatory cells']:\n    print(cell_type)\n    cells_num = cell_type_from_df(cell_type, compounds)\n    plot_de_distr(cells_num, compounds)","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:01:09.578931Z","iopub.execute_input":"2023-12-05T18:01:09.580136Z","iopub.status.idle":"2023-12-05T18:02:01.117288Z","shell.execute_reply.started":"2023-12-05T18:01:09.580097Z","shell.execute_reply":"2023-12-05T18:02:01.116376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualization of Prediction Magnitudes","metadata":{}},{"cell_type":"code","source":"#hide\ncontrol_ids = [\n#     'LSM-36361'  # DMSO\n    'LSM-43181',  # Belinostat\n    'LSM-6303',  # Dabrafenib\n]\n\nprivte_ids = [\n    'LSM-45710', 'LSM-4062', \n    'LSM-2193',  #  'forskolin' -> 'Colforsin'\n    'LSM-4105', 'LSM-4031', 'LSM-1099', 'LSM-45153', 'LSM-3822', 'LSM-4933', \n    'LSM-45630',  # 'KD-025' -> 'SLx-2119'\n    'LSM-6258', 'LSM-1023', 'LSM-2655', 'LSM-47602', 'LSM-3349', 'LSM-1020', 'LSM-1143',\n    'LSM-3828', 'LSM-1051', 'LSM-1120', 'LSM-5467', 'LSM-2292', 'LSM-43293', 'LSM-45437',\n    'LSM-2703', 'LSM-45831', 'LSM-1179', 'LSM-1199', 'LSM-1190', 'LSM-36374', 'LSM-5215',\n    'LSM-1195', 'LSM-45468', 'LSM-45410', 'LSM-47459', 'LSM-45663', 'LSM-45518', 'LSM-1062',\n    'LSM-3667',  # 'BRD-K74305673' -> 'IMD-0354',\n    'LSM-1032', 'LSM-5855', 'LSM-45988',\n    'LSM-24954',  # 'BRD-K98039984' -> 'Prednisolone'\n    'LSM-6286', 'LSM-45984', 'LSM-1124', 'LSM-1165', 'LSM-42802', 'LSM-1121', 'LSM-6308',\n    'LSM-1136', 'LSM-1186', 'LSM-45915', 'LSM-2621', 'LSM-5341', 'LSM-45724', 'LSM-2219',\n    'LSM-2936', 'LSM-3171', 'LSM-46889', 'LSM-2379', 'LSM-47132', 'LSM-47120', 'LSM-47437',\n    'LSM-1139', 'LSM-1144', 'LSM-4353', 'LSM-1210', 'LSM-5887', 'LSM-1025', 'LSM-5771', 'LSM-1132',\n    'LSM-1263',  # 'BRD-A04553218' -> 'Chlorpheniramine'\n    'LSM-1167',\n    'LSM-1194',  # 'BRD-A92800748' -> 'TIE2 Kinase Inhibitor'\n    'LSM-45948', 'LSM-45514', 'LSM-5430', 'LSM-2309', \n]\n\npublic_ids = [\n    'LSM-43216', 'LSM-1050', 'LSM-45849', 'LSM-42800', 'LSM-1131', 'LSM-6335', 'LSM-1211',\n    'LSM-45239', 'LSM-1130', 'LSM-45786', 'LSM-5199', 'LSM-45281',\n    'LSM-6324', # 'ACY-1215' -> 'Ricolinostat'\n    'LSM-3309', 'LSM-1056', 'LSM-45591', 'LSM-46203', 'LSM-5662',\n    'LSM-47134',  # 'SB-2342' -> '5-(9-Isopropyl-8-methyl-2-morpholino-9H-purin-6-yl)pyrimidin-2-amine\t'\n    'LSM-45637', 'LSM-1127', 'LSM-46971', 'LSM-1172', 'LSM-46042', 'LSM-1101', 'LSM-45758',\n    'LSM-5218', 'LSM-2287', 'LSM-1014',\n    'LSM-1040', #  'fostamatinib' -> 'Tamatinib'\n    'LSM-1476;LSM-5290',\n    'LSM-45680',  # 'basimglurant' -> 'RG7090'\n    'LSM-4349',  # '5-iodotubercidin' -> 'IN1451'\n    'LSM-3425', 'LSM-45806',\n    'LSM-45616',  # 'SB-683698' -> 'TR-14035'\n    'LSM-1055',\n    'LSM-43281',  # 'C-646' -> 'STK219801'\n    'LSM-5690', 'LSM-1155', 'LSM-2499',\n    'LSM-2382',  # 'JTC-801' -> 'UNII-BXU45ZH6LI'\n    'LSM-45220', 'LSM-1037', 'LSM-1005', 'LSM-1180', 'LSM-36812',\n    'LSM-45924',  # 'filgotinib' -> 'GLPG0634'\n    'LSM-2013',  # 'TL-HRAS-61' -> TL_HRAS26'\n    'LSM-4738',\n]\n\n\ntrain_ids = [\n    'LSM-1027', 'LSM-1071', 'LSM-45916',\n    'LSM-4944',  # 'ixazomib' -> 'MLN 2238'\n    'LSM-47425',  # 'IWP-L6' -> 'Porcn Inhibitor III'\n    'LSM-1115',\n    'LSM-6237',  # 'CD-437' -> 'O-Demethylated Adapalene'\n    'LSM-1205', 'LSM-45574',\n    'LSM-4255',  # 'NVP-BEZ235' -> 'Dactolisib'\n    'LSM-1181', 'LSM-1158', 'LSM-2334', 'LSM-45496', 'LSM-1011',\n]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-12-05T18:02:01.118835Z","iopub.execute_input":"2023-12-05T18:02:01.119351Z","iopub.status.idle":"2023-12-05T18:02:01.132320Z","shell.execute_reply.started":"2023-12-05T18:02:01.119320Z","shell.execute_reply":"2023-12-05T18:02:01.131344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def classify_id(row):\n    if row['sm_lincs_id'] in privte_ids:\n        return 'Private'\n    elif row['sm_lincs_id'] in public_ids:\n        return 'Public'\n    elif row['sm_lincs_id'] in train_ids:\n        return 'Train'\n    elif row['sm_lincs_id'] in control_ids:\n        return 'Control'\n    else:\n        return 'Other'","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-12-05T18:02:01.133674Z","iopub.execute_input":"2023-12-05T18:02:01.134354Z","iopub.status.idle":"2023-12-05T18:02:01.155771Z","shell.execute_reply.started":"2023-12-05T18:02:01.134311Z","shell.execute_reply":"2023-12-05T18:02:01.154845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_distribution(submission):\n    de_train = pd.read_parquet('/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet')\n    data = de_train.copy()       \n    # Apply the function to create a new column 'id_class' in the 'data' DataFrame\n    lb_split = data.apply(classify_id, axis=1)\n    # Display the updated DataFrame with the new 'id_class' column\n    data.insert(2, 'lb_split', lb_split)\n\n    id_map = pd.read_csv('/kaggle/input/open-problems-single-cell-perturbations/id_map.csv')\n    lb_split = id_map['sm_name'].map(data.set_index('sm_name')['lb_split'].to_dict())\n\n    # Insert the 'per_compound_mean_abs_gene' column\n    pred = submission.copy()\n    # pred = prediction.drop('id', axis=1)\n    pred.insert(0, 'sm_name', id_map['sm_name'])  \n    pred.insert(1, 'lb_split', lb_split)\n    per_compound_mean_abs_gene = pred.iloc[:, 2:].apply(lambda row: np.abs(row).mean(), axis=1)\n\n    pred.insert(0, 'per_compound_mean_abs_gene', per_compound_mean_abs_gene)\n\n    # Filter 'Public' and 'Private' test data\n    public_test = pred[pred['lb_split'] == 'Public']\n    private_test = pred[pred['lb_split'] == 'Private']\n\n    # Plotting\n    plt.figure(figsize=(10, 6))\n\n    # Scatter plot for 'Public' test data\n    plt.scatter(range(len(public_test)), public_test['per_compound_mean_abs_gene'], label='Public', color='blue', alpha=0.7)\n\n    # Scatter plot for 'Private' test data\n    plt.scatter(range(len(private_test)), private_test['per_compound_mean_abs_gene'], label='Private', color='red', alpha=0.7)\n\n    plt.xlabel('Index', fontsize=12)\n    plt.ylabel('Per Compound: abs(predicted gene_expression).mean', fontsize=12)\n    plt.title('Predicted Variance Metric for Public and Private Test Predictions', fontsize=15)\n    plt.legend()\n    plt.grid(True)\n    plt.tight_layout()\n\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-12-05T18:02:01.157136Z","iopub.execute_input":"2023-12-05T18:02:01.158169Z","iopub.status.idle":"2023-12-05T18:02:01.172323Z","shell.execute_reply.started":"2023-12-05T18:02:01.158133Z","shell.execute_reply":"2023-12-05T18:02:01.170872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_distribution(submission)","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:02:01.176571Z","iopub.execute_input":"2023-12-05T18:02:01.177234Z","iopub.status.idle":"2023-12-05T18:02:05.194131Z","shell.execute_reply.started":"2023-12-05T18:02:01.177186Z","shell.execute_reply":"2023-12-05T18:02:05.192827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Postprocess","metadata":{}},{"cell_type":"code","source":"for index, compound_gene_pred in submission.iterrows():\n    if index % 25 == 0:\n        print(index)\n    abs_compound_mean = abs(compound_gene_pred).mean()\n    \n    compound_gene_pred *= min(abs_compound_mean**0.6, 1)\n    submission.loc[index] = compound_gene_pred  \n    \n    abs_compound_mean = abs(compound_gene_pred).mean()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:02:05.195673Z","iopub.execute_input":"2023-12-05T18:02:05.196120Z","iopub.status.idle":"2023-12-05T18:05:56.576023Z","shell.execute_reply.started":"2023-12-05T18:02:05.196081Z","shell.execute_reply":"2023-12-05T18:05:56.574680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_distribution(submission)","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:05:56.577911Z","iopub.execute_input":"2023-12-05T18:05:56.578494Z","iopub.status.idle":"2023-12-05T18:06:00.476843Z","shell.execute_reply.started":"2023-12-05T18:05:56.578457Z","shell.execute_reply":"2023-12-05T18:06:00.475348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission *= MULT","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:06:00.478840Z","iopub.execute_input":"2023-12-05T18:06:00.479732Z","iopub.status.idle":"2023-12-05T18:06:01.404962Z","shell.execute_reply.started":"2023-12-05T18:06:00.479681Z","shell.execute_reply":"2023-12-05T18:06:01.403630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_distribution(submission)","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:06:01.408821Z","iopub.execute_input":"2023-12-05T18:06:01.409227Z","iopub.status.idle":"2023-12-05T18:06:05.133712Z","shell.execute_reply.started":"2023-12-05T18:06:01.409195Z","shell.execute_reply":"2023-12-05T18:06:05.132820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv')\n!ls","metadata":{"execution":{"iopub.status.busy":"2023-12-05T18:06:05.135564Z","iopub.execute_input":"2023-12-05T18:06:05.136352Z","iopub.status.idle":"2023-12-05T18:06:57.987216Z","shell.execute_reply.started":"2023-12-05T18:06:05.136303Z","shell.execute_reply":"2023-12-05T18:06:57.985620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}