{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkcyan;\">🧬Open Problems – Single-Cell Perturbations</h1> \n    <h3 align=\"center\" style=\"color:gray;\">Predict how small molecules change gene expression in different cell types</h3> \n    <h3 align=\"center\" style=\"color:gray;\">By: Somayyeh Gholami & Mehran Kazeminia</h3> \n</div>\n\n# <div style=\"color:white;background-color:darkcyan;padding:2%;border-radius:15px 15px;font-size:1em;text-align:center\">EDA & LinearSVR & RegressorChain</div>\n\nAdvances in single-cell experimental protocols allow for massively multiplexed experiments to measure hundreds of thousands of cells under thousands of unique conditions. These are commonly termed “perturbations”, which are temporary or permanent changes caused by an external influence [Srivatsan et al., 2020].\n\n![](https://cdn-images-1.medium.com/max/1000/1*jCLUHuVRj4ZiBvLFmd8u5g.png)\n\n**Single-Cell Perturbations** datasets measure how gene expression changes in individual cells when they are exposed to different stimuli such as drugs/chemicals/environmental changes. This information can be used to understand how cells work and how they can be manipulated to treat diseases.\n\n![](https://cdn-images-1.medium.com/max/1000/1*El2QMOoqtoUChAZED3w6Kw.png)","metadata":{}},{"cell_type":"code","source":"from IPython.display import HTML\nimport time\n\nhandle = display(HTML(\"\"\"<marquee>👌</marquee>\"\"\"), display_id='html_marquee1')\ntime.sleep(2)\nhandle = display(HTML(\"\"\"<marquee>In this competition setup, participants are tasked with modelling differential expression (DE), which enables us to estimate the impact of an experimental perturbation on the expression level of every gene in the transcription (18211 genes in this dataset). We estimate the impact of each compound by first averaging the raw gene expression counts in each cell of a specific type in each sample, which is called pseudobulking in the single-cell literature. We then fit a linear model to the pseudobulked counts data using Limma and include the library (row), plate, and donor as technical covariates and compound as the experimental covariate.</marquee>\"\"\"), display_id='html_marquee1', update=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-09T15:17:30.882258Z","iopub.execute_input":"2023-10-09T15:17:30.883385Z","iopub.status.idle":"2023-10-09T15:17:32.926665Z","shell.execute_reply.started":"2023-10-09T15:17:30.883327Z","shell.execute_reply":"2023-10-09T15:17:32.925872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![](https://cdn-images-1.medium.com/max/1000/1*KVjhv1iXxYVwmpvG_I8buw.png)","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')\n#:::::::::::::::::::::::::::::::::::\nimport os\nimport gc\nimport glob\nimport random\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\nfrom itertools import groupby\n#:::::::::::::::::::::::::::::::::::\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline\n!ls ../input/*","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-09T15:17:32.928266Z","iopub.execute_input":"2023-10-09T15:17:32.928858Z","iopub.status.idle":"2023-10-09T15:17:38.248855Z","shell.execute_reply.started":"2023-10-09T15:17:32.928828Z","shell.execute_reply":"2023-10-09T15:17:38.247368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install rdkit","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-10-09T15:17:38.251115Z","iopub.execute_input":"2023-10-09T15:17:38.252548Z","iopub.status.idle":"2023-10-09T15:17:52.630183Z","shell.execute_reply.started":"2023-10-09T15:17:38.25246Z","shell.execute_reply":"2023-10-09T15:17:52.628582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from rdkit import Chem\nfrom rdkit.Chem import Draw","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-10-09T15:17:52.634411Z","iopub.execute_input":"2023-10-09T15:17:52.634814Z","iopub.status.idle":"2023-10-09T15:17:52.744344Z","shell.execute_reply.started":"2023-10-09T15:17:52.634781Z","shell.execute_reply":"2023-10-09T15:17:52.743052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div>\n    <h1 align=\"center\" style=\"color:darkred;\">Competition Data (Eight files)</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"The **de_train.parquet** file comprises the main competition data. It contains values for a number of **cell_type / sm_name** pairs. Your goal is to predict corresponding values for the **cell_type / sm_name** pairs given in **id_map.csv**.","metadata":{}},{"cell_type":"code","source":"from IPython.display import HTML\nimport time\n\nhandle = display(HTML(\"\"\"<marquee>👌</marquee>\"\"\"), display_id='html_marquee1')\ntime.sleep(2)\nhandle = display(HTML(\"\"\"<marquee>Note: there is no DE data for the DMSO sample, because it is the negative control. All DE output is calculated in reference to the DMSO, i.e. the DE analysis asks \"how confident am I that each gene increased or decreased relative to DMSO due to the compound treatment\".</marquee>\"\"\"), display_id='html_marquee1', update=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-09T15:17:52.745532Z","iopub.execute_input":"2023-10-09T15:17:52.745847Z","iopub.status.idle":"2023-10-09T15:17:54.760258Z","shell.execute_reply.started":"2023-10-09T15:17:52.745821Z","shell.execute_reply":"2023-10-09T15:17:54.759424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### \"Dataset Description\" by host:\n\nNote that there is no additional test data beyond the indicated cell_type / sm_name pairs. The input to your model will be a tuple of cell_type and sm_name and the output of your model will be predicted signed -log10(p-values) for all 18211 genes.\n\nWe also provide the raw data for the training split, along with 10x Multiome data for the donors at baseline. This raw data is not necessary to compete for the leaderboard prize.\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(1) de_train.parquet</p></div>\n\n#### Aggregated differential expression data in dense array format.\n\n- genes A1BG, A1BG-AS1, …, ZZEF1 (numbering 18,211 in all) - Differential expression value (-log10(p-value) * sign(LFC)) for each gene.\n\n- cell_type - The annotated cell type of each cell based on RNA expression.\n\n- sm_name - The primary name for the (parent) compound (in a standardized representation) as chosen by LINCS. This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n\n- sm_lincs_id - The global LINCS ID (parent) compound (in a standardized representation). This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n\n- SMILES - Simplified molecular-input line-entry system (SMILES) representations of the compounds used in the experiment. This is a 1D representation of molecular structure. These SMILES are provided by Cellarity based on the specific compounds ordered for this experiment.\n\n- control - Boolean indicating whether this instance was used as a control.","metadata":{}},{"cell_type":"code","source":"de_train = pd.read_parquet('../input/open-problems-single-cell-perturbations/de_train.parquet')\nde_train","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:54.761199Z","iopub.execute_input":"2023-10-09T15:17:54.761515Z","iopub.status.idle":"2023-10-09T15:17:57.357568Z","shell.execute_reply.started":"2023-10-09T15:17:54.761469Z","shell.execute_reply":"2023-10-09T15:17:57.356586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"MV = de_train.isnull().sum()\nprint('Missing Value in de_train:', MV[MV > 0])\nprint('Duplicates in cell_type:', de_train.cell_type.duplicated().sum())\nprint('Duplicates in sm_name:', de_train.sm_name.duplicated().sum())","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:57.359015Z","iopub.execute_input":"2023-10-09T15:17:57.360228Z","iopub.status.idle":"2023-10-09T15:17:57.402794Z","shell.execute_reply.started":"2023-10-09T15:17:57.36019Z","shell.execute_reply":"2023-10-09T15:17:57.401548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">cell_type / sm_name</span>","metadata":{}},{"cell_type":"code","source":"de_train.pivot(columns=['sm_lincs_id', 'sm_name', 'cell_type'], values='control')","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:57.404436Z","iopub.execute_input":"2023-10-09T15:17:57.405115Z","iopub.status.idle":"2023-10-09T15:17:57.451827Z","shell.execute_reply.started":"2023-10-09T15:17:57.405078Z","shell.execute_reply":"2023-10-09T15:17:57.450559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"::::::::::::::::::::::::::::::::::::::::::::::\n\ncell_type / sm_name\n\nUnique numbers: 614\n\nTotal numbers:  614\n\n::::::::::::::::::::::::::::::::::::::::::::::","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">cell_type</span>","metadata":{}},{"cell_type":"code","source":"cell_type_list = list(de_train['cell_type'].unique())\nprint(cell_type_list)\nlen(cell_type_list)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:57.453663Z","iopub.execute_input":"2023-10-09T15:17:57.45489Z","iopub.status.idle":"2023-10-09T15:17:57.464675Z","shell.execute_reply.started":"2023-10-09T15:17:57.454835Z","shell.execute_reply":"2023-10-09T15:17:57.462809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.gca().set_facecolor('lightyellow')\nde_train['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,2), color=['darkcyan'])\n\npd.DataFrame(data= {'Number': de_train['cell_type'].value_counts(), \n                    'Percent': de_train['cell_type'].value_counts(normalize=True)})","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:57.468667Z","iopub.execute_input":"2023-10-09T15:17:57.469016Z","iopub.status.idle":"2023-10-09T15:17:57.732639Z","shell.execute_reply.started":"2023-10-09T15:17:57.468988Z","shell.execute_reply":"2023-10-09T15:17:57.731554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.gca().set_facecolor('lightyellow')\nde_train.groupby('cell_type')['control'].value_counts(normalize=True).plot(kind='barh', figsize=(15,3), color=['darkcyan','red'])\n\npd.DataFrame(de_train.groupby('cell_type')['control'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:57.734076Z","iopub.execute_input":"2023-10-09T15:17:57.735225Z","iopub.status.idle":"2023-10-09T15:17:58.034379Z","shell.execute_reply.started":"2023-10-09T15:17:57.735183Z","shell.execute_reply":"2023-10-09T15:17:58.032954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">sm_name</span>","metadata":{}},{"cell_type":"code","source":"sm_name_list = list(de_train['sm_name'].unique())\nlen(sm_name_list)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:58.03628Z","iopub.execute_input":"2023-10-09T15:17:58.036801Z","iopub.status.idle":"2023-10-09T15:17:58.046359Z","shell.execute_reply.started":"2023-10-09T15:17:58.036755Z","shell.execute_reply":"2023-10-09T15:17:58.045049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.gca().set_facecolor('lightyellow')\nde_train['sm_name'].value_counts(normalize=True).plot(kind='barh', figsize=(15,40), color=['lightblue'])\n\npd.DataFrame(data= {'Number': de_train['sm_name'].value_counts(), \n                    'Percent': de_train['sm_name'].value_counts(normalize=True)})","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:58.048391Z","iopub.execute_input":"2023-10-09T15:17:58.049164Z","iopub.status.idle":"2023-10-09T15:17:59.845905Z","shell.execute_reply.started":"2023-10-09T15:17:58.049116Z","shell.execute_reply":"2023-10-09T15:17:59.845104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">SMILES</span>","metadata":{}},{"cell_type":"code","source":"import re\n# Thanks to: https://github.com/DocMinus\n\ndef element_count(data, N):\n    smiles = data['SMILES'].iloc[N]\n\n    pattern = \"Si|Ti|Al|Zn|Pd|Pt|Br?|Cl?|N|O|S|P|F|I|B|b|c|n|o|s|p\" \n    regex = re.compile(pattern)\n    elements = [token for token in regex.findall(smiles)]\n    ele_length = len(elements)\n    lowercase_pattern = \"b|c|n|o|s|p\"\n    regex_low = re.compile(lowercase_pattern)\n    for i in range(ele_length):\n        if regex_low.findall(elements[i]):\n            elements[i] = elements[i].upper()\n    element_count = [elements.count(ele) for ele in elements]\n    formula = dict(zip(elements, element_count)) \n\n    print('==> Element_Count: ', str(formula))","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:59.847217Z","iopub.execute_input":"2023-10-09T15:17:59.848322Z","iopub.status.idle":"2023-10-09T15:17:59.854749Z","shell.execute_reply.started":"2023-10-09T15:17:59.848289Z","shell.execute_reply":"2023-10-09T15:17:59.854068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N = random.randrange(len(de_train))\ncols_sel = ['cell_type','sm_name','sm_lincs_id','SMILES']\n\nsmiles = de_train['SMILES'][N]\nmol = Chem.MolFromSmiles(smiles)\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)    \ndisplay(pd.DataFrame(de_train[cols_sel].iloc[N]))\n\nelement_count(de_train, N)\nimg","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:59.855956Z","iopub.execute_input":"2023-10-09T15:17:59.857175Z","iopub.status.idle":"2023-10-09T15:17:59.91783Z","shell.execute_reply.started":"2023-10-09T15:17:59.857143Z","shell.execute_reply":"2023-10-09T15:17:59.917083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkred;\">Principal Component Analysis (PCA)</span>\n\nPrincipal component analysis, or PCA, is a statistical procedure that allows you to summarize the information content in large data tables by means of a smaller set of “summary indices” that can be more easily visualized and analyzed.\n\nSince our dataset has many features, we reduce these features by using PCA to visualize the distribution of samples in a two-dimensional graph.","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\n\npca = PCA(n_components=2)\ntarget_2d = pca.fit_transform(de_train.iloc[:,5:])\ntarget_2d.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:17:59.918889Z","iopub.execute_input":"2023-10-09T15:17:59.919646Z","iopub.status.idle":"2023-10-09T15:18:01.444026Z","shell.execute_reply.started":"2023-10-09T15:17:59.919617Z","shell.execute_reply":"2023-10-09T15:18:01.44284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_task(target, feature, headline):   \n    colors  = ['darkblue','red','darkcyan','darkviolet','violet','yellow']\n    classes = list(feature.unique())\n    \n    plt.style.use('seaborn-whitegrid') \n    plt.figure(figsize=(10, 4), facecolor='lightyellow')\n    for u, c in zip(classes, colors):\n        plt.scatter(target[:,0][feature==u], target[:,1][feature==u], c=c, s=10)\n        \n    plt.gca().set_facecolor('lightgray')\n    plt.title(headline, fontsize=12)\n    plt.legend(classes, loc=1)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:18:01.446036Z","iopub.execute_input":"2023-10-09T15:18:01.446836Z","iopub.status.idle":"2023-10-09T15:18:01.457141Z","shell.execute_reply.started":"2023-10-09T15:18:01.44679Z","shell.execute_reply":"2023-10-09T15:18:01.455475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_task(target_2d, de_train['cell_type'], 'Scatter graph for cell_type_feature')","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:18:01.459918Z","iopub.execute_input":"2023-10-09T15:18:01.460931Z","iopub.status.idle":"2023-10-09T15:18:01.949179Z","shell.execute_reply.started":"2023-10-09T15:18:01.460877Z","shell.execute_reply":"2023-10-09T15:18:01.94797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_task(target_2d, de_train['control'], 'Scatter graph for control_feature')","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:18:01.9509Z","iopub.execute_input":"2023-10-09T15:18:01.951312Z","iopub.status.idle":"2023-10-09T15:18:02.232697Z","shell.execute_reply.started":"2023-10-09T15:18:01.951273Z","shell.execute_reply":"2023-10-09T15:18:02.231313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkred;\">Shapiro–Wilk test</span>","metadata":{}},{"cell_type":"code","source":"cols_total = list(de_train.columns)\nlen(cols_total)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:18:02.234436Z","iopub.execute_input":"2023-10-09T15:18:02.234988Z","iopub.status.idle":"2023-10-09T15:18:02.246028Z","shell.execute_reply.started":"2023-10-09T15:18:02.23494Z","shell.execute_reply":"2023-10-09T15:18:02.244198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"de_train_drop = de_train.drop(columns=['cell_type','sm_name','sm_lincs_id','SMILES','control'])\ncols = list(de_train_drop.columns)\nlen(cols)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:18:02.248098Z","iopub.execute_input":"2023-10-09T15:18:02.248797Z","iopub.status.idle":"2023-10-09T15:18:02.302349Z","shell.execute_reply.started":"2023-10-09T15:18:02.248765Z","shell.execute_reply":"2023-10-09T15:18:02.300988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_select = []\nalpha = 0.05\n\nfor col in cols:\n    _, p_value = stats.shapiro(de_train_drop[col])\n    \n    if (p_value <= alpha): \n        cols_select.append(col)     \n         \nlen(cols_select)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:18:02.303898Z","iopub.execute_input":"2023-10-09T15:18:02.304493Z","iopub.status.idle":"2023-10-09T15:18:05.55Z","shell.execute_reply.started":"2023-10-09T15:18:02.304444Z","shell.execute_reply":"2023-10-09T15:18:05.548708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(2) adata_train.parquet</p></div>\n\n#### Unaggregated count and normalized data in COO sparse-array format. A supplement to de_train. In addition to the fields in de_train, this data also has:\n\n- obs_id - This is a unique identifier assigned to each cell in the raw dataset.\n\n- gene - Corresponds to the columns of de_train.\n\n- count - The raw molecular counts for the gene expression data measured in the experiment as output by 10x CellRanger.\n\n- normalized_count - These counts have been library size normalized and log(X+1) transformed.","metadata":{}},{"cell_type":"code","source":"adata_train = pd.read_parquet('../input/open-problems-single-cell-perturbations/adata_train.parquet')\nadata_train","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:18:05.551721Z","iopub.execute_input":"2023-10-09T15:18:05.552067Z","iopub.status.idle":"2023-10-09T15:19:10.249385Z","shell.execute_reply.started":"2023-10-09T15:18:05.552039Z","shell.execute_reply":"2023-10-09T15:19:10.248641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#gene_list = list(adata_train['gene'].unique())\n#len(gene_list)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:10.250902Z","iopub.execute_input":"2023-10-09T15:19:10.251472Z","iopub.status.idle":"2023-10-09T15:19:10.255778Z","shell.execute_reply.started":"2023-10-09T15:19:10.251444Z","shell.execute_reply":"2023-10-09T15:19:10.254759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unique number of genes: 21255","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(3) adata_obs_meta.csv</p></div>\n\n#### Observation metadata for adata_train.\n\n- library_id - A unique identifier for each library, which is a measurement made on pooled samples from each row of the plate. All cells from wells on the same row of the same plate will share a library_id.\n\n- plate_name - A unique ID for all samples from the same plate.\n\n- well - The well location of the sample on each plate (this is standard across 96 well plate experiments). It is a concatenation of row and col.\n\n- row - Which row on the plate the sample came from.\n\n- col - Which column on the plate the sample came from.\n\n- donor_id - Identifies the donor source of the sample, one of three.\n\n- cell_type - The annotated cell type of each cell based on RNA expression. This matches the cell_type in the de_train.parquet.\n\n- cell_id - This is included for consistency with LINCS Connectivity Map metadata, which denotes a cell_id for each cell line.\n\n- sm_name - The primary name for the (parent) compound (in a standardized representation) as chosen by LINCS. This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n\n- sm_lincs_id - The global LINCS ID (parent) compound (in a standardized representation). This is provided to map the data in this experiment to the LINCS Connectivity Map data.\n\n- SMILES - Simplified molecular-input line-entry system (SMILES) representations of the compounds used in the experiment. This is a 1D representation of molecular structure. These SMILES are provided by Cellarity based on the specific compounds ordered for this experiment.\n\n- dose_uM - Dose of the compound in on a micro-molar scale. This maps to the pert_idose field in LINCS.\n\n- timepoint_hr - Duration of treatment in hours. This maps to the pert_itime field in LINCS.\n\n- control - Whether this observation was used as a control, True or False.","metadata":{}},{"cell_type":"code","source":"adata_obs_meta = pd.read_csv('../input/open-problems-single-cell-perturbations/adata_obs_meta.csv')\nadata_obs_meta","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:10.257172Z","iopub.execute_input":"2023-10-09T15:19:10.257539Z","iopub.status.idle":"2023-10-09T15:19:12.244497Z","shell.execute_reply.started":"2023-10-09T15:19:10.257512Z","shell.execute_reply":"2023-10-09T15:19:12.243576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">cell_type / sm_name</span>","metadata":{}},{"cell_type":"code","source":"adata_obs_meta.pivot(columns=['sm_lincs_id', 'sm_name', 'cell_type'], values='control')","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:12.246166Z","iopub.execute_input":"2023-10-09T15:19:12.246913Z","iopub.status.idle":"2023-10-09T15:19:16.494773Z","shell.execute_reply.started":"2023-10-09T15:19:12.246877Z","shell.execute_reply":"2023-10-09T15:19:16.493285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"::::::::::::::::::::::::::::::::::::::::::::::\n\ncell_type / sm_name\n\nUnique numbers: 620\n\nTotal numbers: 240090 \n\n::::::::::::::::::::::::::::::::::::::::::::::","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">cell_type</span>","metadata":{}},{"cell_type":"code","source":"cell_type_list = list(adata_obs_meta['cell_type'].unique())\nprint(cell_type_list)\nlen(cell_type_list)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:16.496614Z","iopub.execute_input":"2023-10-09T15:19:16.49708Z","iopub.status.idle":"2023-10-09T15:19:16.519222Z","shell.execute_reply.started":"2023-10-09T15:19:16.497041Z","shell.execute_reply":"2023-10-09T15:19:16.517804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">sm_name</span>","metadata":{}},{"cell_type":"code","source":"sm_name_list = list(adata_obs_meta['sm_name'].unique())\nlen(sm_name_list)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:16.526963Z","iopub.execute_input":"2023-10-09T15:19:16.527662Z","iopub.status.idle":"2023-10-09T15:19:16.550953Z","shell.execute_reply.started":"2023-10-09T15:19:16.52762Z","shell.execute_reply":"2023-10-09T15:19:16.549935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">SMILES</span>","metadata":{}},{"cell_type":"code","source":"N = random.randrange(len(adata_obs_meta))\ncols_sel = ['cell_type','sm_name','sm_lincs_id','SMILES']\n\nsmiles = adata_obs_meta['SMILES'][N]\nmol = Chem.MolFromSmiles(smiles)\nimg = Draw.MolToImage(mol, size=(700, 300), fitImage=True)    \ndisplay(pd.DataFrame(adata_obs_meta[cols_sel].iloc[N]))\n\nelement_count(adata_obs_meta, N)\nimg","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:16.552175Z","iopub.execute_input":"2023-10-09T15:19:16.552475Z","iopub.status.idle":"2023-10-09T15:19:16.60925Z","shell.execute_reply.started":"2023-10-09T15:19:16.552442Z","shell.execute_reply":"2023-10-09T15:19:16.608109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_n = random.sample(range(len(adata_obs_meta)), 30)\n\nimg_list = []\nfor n in sample_n:\n    smiles = adata_obs_meta['SMILES'][n]\n    mol = Chem.MolFromSmiles(smiles)\n    img = Draw.MolToImage(mol, size=(700, 300), fitImage=False)\n    img_list.append(img)\n\nimg_list[0]","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:16.610704Z","iopub.execute_input":"2023-10-09T15:19:16.610975Z","iopub.status.idle":"2023-10-09T15:19:17.345339Z","shell.execute_reply.started":"2023-10-09T15:19:16.610952Z","shell.execute_reply":"2023-10-09T15:19:17.343883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(4) multiome_train.parquet</p></div>\n\n#### This is optional additional 10x Multiome data for each sample at baseline.\n\n- obs_id - Unique identifier for each observation. (Distinct from identifiers used in adata.)\n\n- location - This is a feature ID. If the feature_type in multiome_var_meta.csv is Gene Expression then this is a gene symbol. If feature_type is Peaks, then this is the genomic interval of the peak.\n\n- count - This is the raw molecular counts of the transcript for accessible DNA measurement as output by Cellranger-Arc.\n\n- normalized_count - If the feature_type in multiome_var_meta.csv is Gene Expression then this is library size normalized and log(X+1) transformed counts. If feature_type is Peaks, then this is ATAC-seq peak counts transformed with TF-IDF using the default log(TF) * log(IDF).","metadata":{}},{"cell_type":"code","source":"#multiome_train = pd.read_parquet('../input/open-problems-single-cell-perturbations/multiome_train.parquet')\n#multiome_train","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:17.346873Z","iopub.execute_input":"2023-10-09T15:19:17.347326Z","iopub.status.idle":"2023-10-09T15:19:17.352205Z","shell.execute_reply.started":"2023-10-09T15:19:17.347294Z","shell.execute_reply":"2023-10-09T15:19:17.350828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(5) multiome_obs_meta.csv</p></div>\n\n- obs_id - Identifier corresponding to that in multiome_train.parquet.\n\n- cell_type - The annotated cell type of each cell based on RNA expression.\n\n- donor_id - Identifies the donor source of the sample, one of three.","metadata":{}},{"cell_type":"code","source":"#multiome_obs_meta = pd.read_csv('../input/open-problems-single-cell-perturbations/multiome_obs_meta.csv')\n#multiome_obs_meta","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:17.353991Z","iopub.execute_input":"2023-10-09T15:19:17.354371Z","iopub.status.idle":"2023-10-09T15:19:17.36707Z","shell.execute_reply.started":"2023-10-09T15:19:17.354343Z","shell.execute_reply":"2023-10-09T15:19:17.365763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(6) multiome_var_meta.csv</p></div>\n\n- location - This is a feature ID. If the feature_type is Gene Expression then this is a gene symbol. If feature_type is Peaks, then this is the genomic interval of the peak.\n\n- gene_id - This is an alternative unique feature ID. If the feature_type is Gene Expression then this is an Ensembl Stable Gene ID. If feature_type is Peaks, then this is the genomic interval of the peak.\n\n- feature_type - Denotes whether the feature is an RNA expression measurement or a Chromatin Accessibility measurement.\n\n- genome - The genome version used when running CellRanger-Arc\n\n- interval - The genomic coordinates of each feature on reference genome GRCh38. Genomic coordinates are directly related to the reference genome and include the chromosome name, start position, and end position in the following format: chr1:1234570-1234870","metadata":{}},{"cell_type":"code","source":"#multiome_var_meta = pd.read_csv('../input/open-problems-single-cell-perturbations/multiome_var_meta.csv')\n#multiome_var_meta","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:17.368897Z","iopub.execute_input":"2023-10-09T15:19:17.369377Z","iopub.status.idle":"2023-10-09T15:19:17.382848Z","shell.execute_reply.started":"2023-10-09T15:19:17.369333Z","shell.execute_reply":"2023-10-09T15:19:17.381577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(7) id_map.csv</p></div>\n\n#### Identifies the cell_type / sm_name pair to be predicted for the given id.","metadata":{}},{"cell_type":"code","source":"id_map = pd.read_csv('../input/open-problems-single-cell-perturbations/id_map.csv')\nid_map","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:17.384391Z","iopub.execute_input":"2023-10-09T15:19:17.385238Z","iopub.status.idle":"2023-10-09T15:19:17.411134Z","shell.execute_reply.started":"2023-10-09T15:19:17.385204Z","shell.execute_reply":"2023-10-09T15:19:17.410061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">cell_type / sm_name</span>","metadata":{}},{"cell_type":"code","source":"id_map.pivot(columns=['cell_type', 'sm_name'], values='id')","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:17.412428Z","iopub.execute_input":"2023-10-09T15:19:17.412918Z","iopub.status.idle":"2023-10-09T15:19:17.457792Z","shell.execute_reply.started":"2023-10-09T15:19:17.412891Z","shell.execute_reply":"2023-10-09T15:19:17.45637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"::::::::::::::::::::::::::::::::::::::::::::::\n\ncell_type / sm_name\n\nUnique numbers: 255\n\nTotal numbers:  255\n\n::::::::::::::::::::::::::::::::::::::::::::::","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">cell_type</span>","metadata":{}},{"cell_type":"code","source":"plt.gca().set_facecolor('lightyellow')\nid_map['cell_type'].value_counts(normalize=True).plot(kind='barh', figsize=(15,1), color=['darkcyan'])\n\npd.DataFrame(data= {'Number': id_map['cell_type'].value_counts(), \n                    'Percent': id_map['cell_type'].value_counts(normalize=True)})","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:17.459232Z","iopub.execute_input":"2023-10-09T15:19:17.459728Z","iopub.status.idle":"2023-10-09T15:19:17.667861Z","shell.execute_reply.started":"2023-10-09T15:19:17.459696Z","shell.execute_reply":"2023-10-09T15:19:17.667045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">sm_name</span>","metadata":{}},{"cell_type":"code","source":"plt.gca().set_facecolor('lightyellow')\nid_map['sm_name'].value_counts(normalize=True).plot(kind='barh', figsize=(15,40), color=['lightblue'])\n\npd.DataFrame(data= {'Number': id_map['sm_name'].value_counts(), \n                    'Percent': id_map['sm_name'].value_counts(normalize=True)})","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:17.668987Z","iopub.execute_input":"2023-10-09T15:19:17.669999Z","iopub.status.idle":"2023-10-09T15:19:19.329195Z","shell.execute_reply.started":"2023-10-09T15:19:17.669968Z","shell.execute_reply":"2023-10-09T15:19:19.328244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>(8) sample_submission.csv</p></div>\n\n#### A sample submission file in the correct format.","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('../input/open-problems-single-cell-perturbations/sample_submission.csv', index_col='id')\nsample_submission","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:19.330463Z","iopub.execute_input":"2023-10-09T15:19:19.331248Z","iopub.status.idle":"2023-10-09T15:19:22.247101Z","shell.execute_reply.started":"2023-10-09T15:19:19.331215Z","shell.execute_reply":"2023-10-09T15:19:22.246227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">B cells || common_member > check</span>","metadata":{}},{"cell_type":"code","source":"de_train_b = de_train[de_train['cell_type']=='B cells']\n\nnamelist_b = list(de_train_b['sm_name'])\nnamelist_b , len(namelist_b)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.248347Z","iopub.execute_input":"2023-10-09T15:19:22.248816Z","iopub.status.idle":"2023-10-09T15:19:22.265217Z","shell.execute_reply.started":"2023-10-09T15:19:22.248787Z","shell.execute_reply":"2023-10-09T15:19:22.264186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id_map_b = id_map[id_map['cell_type']=='B cells']\n\nnamelist_b_sub = list(id_map_b['sm_name'])\nlen(namelist_b_sub)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.266221Z","iopub.execute_input":"2023-10-09T15:19:22.266503Z","iopub.status.idle":"2023-10-09T15:19:22.294956Z","shell.execute_reply.started":"2023-10-09T15:19:22.266457Z","shell.execute_reply":"2023-10-09T15:19:22.293739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"common_member_b = [f for f in namelist_b_sub if f in namelist_b]\ncommon_member_b","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.296431Z","iopub.execute_input":"2023-10-09T15:19:22.297104Z","iopub.status.idle":"2023-10-09T15:19:22.30894Z","shell.execute_reply.started":"2023-10-09T15:19:22.297074Z","shell.execute_reply":"2023-10-09T15:19:22.307546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">Myeloid cells || common_member > check</span>","metadata":{}},{"cell_type":"code","source":"de_train_m = de_train[de_train['cell_type']=='Myeloid cells']\n\nnamelist_m = list(de_train_m['sm_name'])\nnamelist_m , len(namelist_m)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.310536Z","iopub.execute_input":"2023-10-09T15:19:22.310945Z","iopub.status.idle":"2023-10-09T15:19:22.329562Z","shell.execute_reply.started":"2023-10-09T15:19:22.310914Z","shell.execute_reply":"2023-10-09T15:19:22.328303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id_map_m = id_map[id_map['cell_type']=='Myeloid cells']\n\nnamelist_m_sub = list(id_map_m['sm_name'])\nlen(namelist_m_sub)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.331112Z","iopub.execute_input":"2023-10-09T15:19:22.331754Z","iopub.status.idle":"2023-10-09T15:19:22.34594Z","shell.execute_reply.started":"2023-10-09T15:19:22.331722Z","shell.execute_reply":"2023-10-09T15:19:22.344692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"common_member_m = [f for f in namelist_m_sub if f in namelist_m]\ncommon_member_m","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.34738Z","iopub.execute_input":"2023-10-09T15:19:22.347879Z","iopub.status.idle":"2023-10-09T15:19:22.355832Z","shell.execute_reply.started":"2023-10-09T15:19:22.347839Z","shell.execute_reply":"2023-10-09T15:19:22.354997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n# <span style=\"color:darkred;\">train | test | target</span>","metadata":{}},{"cell_type":"code","source":"xlist  = ['cell_type','sm_name']\n_ylist = ['cell_type','sm_name','sm_lincs_id','SMILES','control']\n\ny = de_train.drop(columns=_ylist)\ny.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.356854Z","iopub.execute_input":"2023-10-09T15:19:22.357614Z","iopub.status.idle":"2023-10-09T15:19:22.421579Z","shell.execute_reply.started":"2023-10-09T15:19:22.357583Z","shell.execute_reply":"2023-10-09T15:19:22.420723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:darkcyan;\">get_dummies</span>","metadata":{}},{"cell_type":"code","source":"train  = pd.get_dummies(de_train[xlist], columns=xlist)\ntrain.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.422658Z","iopub.execute_input":"2023-10-09T15:19:22.423127Z","iopub.status.idle":"2023-10-09T15:19:22.435192Z","shell.execute_reply.started":"2023-10-09T15:19:22.4231Z","shell.execute_reply":"2023-10-09T15:19:22.434429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.get_dummies(id_map[xlist], columns=xlist)\ntest.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.436197Z","iopub.execute_input":"2023-10-09T15:19:22.437009Z","iopub.status.idle":"2023-10-09T15:19:22.448962Z","shell.execute_reply.started":"2023-10-09T15:19:22.43698Z","shell.execute_reply":"2023-10-09T15:19:22.448173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"uncommon = [f for f in train if f not in test]\n\nX = train.drop(columns=uncommon)\nX.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.449997Z","iopub.execute_input":"2023-10-09T15:19:22.450812Z","iopub.status.idle":"2023-10-09T15:19:22.465389Z","shell.execute_reply.started":"2023-10-09T15:19:22.450779Z","shell.execute_reply":"2023-10-09T15:19:22.463885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n## <span style=\"color:darkred;\">Evaluation</span>\n\n#### Mean Rowwise Root Mean Squared Error","metadata":{}},{"cell_type":"code","source":"def mrrmse_pd(y_pred: pd.DataFrame, y_true: pd.DataFrame):\n    \n    return ((y_pred - y_true)**2).mean(axis=1).apply(np.sqrt).mean()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.466914Z","iopub.execute_input":"2023-10-09T15:19:22.467273Z","iopub.status.idle":"2023-10-09T15:19:22.477Z","shell.execute_reply.started":"2023-10-09T15:19:22.467244Z","shell.execute_reply":"2023-10-09T15:19:22.475817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def mrrmse_np(y_pred, y_true):\n    \n    return np.sqrt(np.square(y_true - y_pred).mean(axis=1)).mean()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.478438Z","iopub.execute_input":"2023-10-09T15:19:22.47882Z","iopub.status.idle":"2023-10-09T15:19:22.490869Z","shell.execute_reply.started":"2023-10-09T15:19:22.47879Z","shell.execute_reply":"2023-10-09T15:19:22.489422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n# <span style=\"color:darkred;\">LinearSVR & MultiOutputRegressor</span>","metadata":{}},{"cell_type":"code","source":"from sklearn.svm import LinearSVR\nfrom sklearn.multioutput import MultiOutputRegressor\nfrom sklearn.model_selection import train_test_split\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=421)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.492429Z","iopub.execute_input":"2023-10-09T15:19:22.493083Z","iopub.status.idle":"2023-10-09T15:19:22.570459Z","shell.execute_reply.started":"2023-10-09T15:19:22.493047Z","shell.execute_reply":"2023-10-09T15:19:22.569072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = LinearSVR(max_iter= 1000, epsilon= 0.1)\nwrapper = MultiOutputRegressor(model)\nwrapper.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:19:22.571512Z","iopub.execute_input":"2023-10-09T15:19:22.571786Z","iopub.status.idle":"2023-10-09T15:20:43.406768Z","shell.execute_reply.started":"2023-10-09T15:19:22.571762Z","shell.execute_reply":"2023-10-09T15:20:43.405566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict = wrapper.predict(X_test) \n#mrrmse_pd(pd.DataFrame(predict), pd.DataFrame(y_test.values))","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:20:43.408515Z","iopub.execute_input":"2023-10-09T15:20:43.409595Z","iopub.status.idle":"2023-10-09T15:21:29.594648Z","shell.execute_reply.started":"2023-10-09T15:20:43.409551Z","shell.execute_reply":"2023-10-09T15:21:29.593404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N = random.randrange(predict.shape[1])\n\nplt.style.use('seaborn-whitegrid') \nplt.figure(figsize=(8, 4), facecolor='lightyellow')\nplt.title(f'Column:  #{N}', fontsize=12)\nplt.gca().set_facecolor('lightgray')\n\nsns.distplot(y_test.values[:,N]-predict[:,N], bins=100, color='red')\nplt.legend(['y_true','y_pred'], loc=1)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:21:29.596313Z","iopub.execute_input":"2023-10-09T15:21:29.596956Z","iopub.status.idle":"2023-10-09T15:21:30.03508Z","shell.execute_reply.started":"2023-10-09T15:21:29.596918Z","shell.execute_reply":"2023-10-09T15:21:30.034009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>\n\n# <span style=\"color:darkred;\">Chained Multioutput Regression</span>\n\n- **Regression:** Predict a single numeric output given an input.\n\n- **Multioutput Regression:** Predict two or more numeric outputs given an input.\n\nIn multioutput regression, typically the outputs are dependent upon the input and upon each other. This means that often the outputs are not independent of each other and may require a model that predicts both outputs together or each output contingent upon the other outputs.\n\n### RegressorChain\n\nThe first model in the sequence uses the input and predicts one output; the second model uses the input and the output from the first model to make a prediction; the third model uses the input and output from the first two models to make a prediction, and so on.\n\nThis work is usually done for outputs that depend on both the input and the other outputs, and the RegressorChain library in sklearn can be used. But in this challenge with about eighteen thousand outputs, it is not possible to use this type of library. But to demonstrate this process, we will only repeat this chain three times and for each time, we will predict only 1000 outputs.\n\nDue to the long running time, we did this on [our second notebook](https://www.kaggle.com/code/mehrankazeminia/2-op2-regressor-chain/notebook) and import the results here. Go to the following address to see the details:\n\nhttps://www.kaggle.com/code/mehrankazeminia/2-op2-regressor-chain/notebook\n","metadata":{}},{"cell_type":"code","source":"sub_chain = pd.read_csv('../input/2-op2-regressor-chain/submission.csv', index_col='id')\nsub_chain.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:21:30.036699Z","iopub.execute_input":"2023-10-09T15:21:30.037358Z","iopub.status.idle":"2023-10-09T15:21:34.42764Z","shell.execute_reply.started":"2023-10-09T15:21:30.03732Z","shell.execute_reply":"2023-10-09T15:21:34.426547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Please note** that due to the large number of columns and the long execution time, we only performed an example of the RegressorChain method and certainly did not expect much impact on the results. By the way, in this method, usually the prioritization for the columns is very important. That is, we must know which column values should be predicted first.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <span style=\"color:darkred;\">Ensembling</span>","metadata":{}},{"cell_type":"markdown","source":"The following files are actually the result of various guesses, optimization and ensembling of the results presented in several public notebooks of this challenge, sometimes their mathematical logic is not clear and in any case there is a risk of overfitting.","metadata":{}},{"cell_type":"code","source":"sub_import1 = pd.read_csv('../input/op2-603/op2_603.csv', index_col='id')\nsub_import1.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:21:34.428998Z","iopub.execute_input":"2023-10-09T15:21:34.429334Z","iopub.status.idle":"2023-10-09T15:21:39.117359Z","shell.execute_reply.started":"2023-10-09T15:21:34.429305Z","shell.execute_reply":"2023-10-09T15:21:39.116299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: **@vendekagonlabs**  https://www.kaggle.com/code/vendekagonlabs/jax-autoencoder-quickstart","metadata":{}},{"cell_type":"code","source":"sub_import2 = pd.read_csv('../input/op2-720/op2_720.csv', index_col='id')\nsub_import2.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:21:39.118787Z","iopub.execute_input":"2023-10-09T15:21:39.120168Z","iopub.status.idle":"2023-10-09T15:21:43.487445Z","shell.execute_reply.started":"2023-10-09T15:21:39.12013Z","shell.execute_reply":"2023-10-09T15:21:43.486163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: **@kishanvavdara**  https://www.kaggle.com/code/kishanvavdara/neural-network-regression","metadata":{}},{"cell_type":"code","source":"sub_import3 = pd.read_csv('../input/op2-604/submission_df.csv', index_col='id')\nsub_import3.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:21:43.489296Z","iopub.execute_input":"2023-10-09T15:21:43.489691Z","iopub.status.idle":"2023-10-09T15:21:47.496674Z","shell.execute_reply.started":"2023-10-09T15:21:43.489662Z","shell.execute_reply":"2023-10-09T15:21:47.495389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_import4 = pd.read_csv('/kaggle/input/op-v3-neural-network-regression/submission_df.csv', index_col='id')\nsub_import4.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:21:47.49802Z","iopub.execute_input":"2023-10-09T15:21:47.499004Z","iopub.status.idle":"2023-10-09T15:21:51.14488Z","shell.execute_reply.started":"2023-10-09T15:21:47.498944Z","shell.execute_reply":"2023-10-09T15:21:51.143579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col = list(de_train.columns[5:])\nsubmission = sample_submission.copy()\n\nsubmission[col] = (sub_import4[col] *0.201) +(sub_import1[col] *0.240) + (sub_import2[col] *0.120) + (sub_import3[col] *0.440) - (sub_chain[col] *0.099)\n\nsubmission.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:21:51.146848Z","iopub.execute_input":"2023-10-09T15:21:51.147198Z","iopub.status.idle":"2023-10-09T15:22:00.688328Z","shell.execute_reply.started":"2023-10-09T15:21:51.147171Z","shell.execute_reply":"2023-10-09T15:22:00.687029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv')\n!ls","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:22:00.689704Z","iopub.execute_input":"2023-10-09T15:22:00.690017Z","iopub.status.idle":"2023-10-09T15:22:41.772421Z","shell.execute_reply.started":"2023-10-09T15:22:00.689991Z","shell.execute_reply":"2023-10-09T15:22:41.770898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}}]}