{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59094,"databundleVersionId":7010844,"sourceType":"competition"}],"dockerImageVersionId":30579,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Credit\nFork from [Single Cells Perturbations -- 1 | Kaggle](https://www.kaggle.com/code/amanmukati/single-cells-perturbations-1)","metadata":{"_uuid":"09360334-a18b-4238-bb8f-3acf8bbded42","_cell_guid":"23d56b2e-217f-455a-beb9-e4fbd648e64e","trusted":true}},{"cell_type":"markdown","source":"## Data Loading","metadata":{"_uuid":"b9944023-633f-4cc3-91e7-cfb8b3e79098","_cell_guid":"6308e395-60f3-4226-970e-5b857958281d","trusted":true}},{"cell_type":"code","source":"# Standard libraries\nimport pandas as pd\nimport numpy as np\nimport time\n\n# Visualization libraries\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport networkx as nx\n#from pyvis.network import Network\n\n# Machine learning libraries\nfrom sklearn.cluster import DBSCAN\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.linear_model import Ridge\nfrom sklearn.manifold import TSNE\nfrom sklearn.decomposition import PCA, FastICA, TruncatedSVD\nfrom sklearn.model_selection import train_test_split, KFold\nfrom sklearn.metrics import r2_score, mean_squared_error\nfrom sklearn.preprocessing import LabelEncoder, OrdinalEncoder, OneHotEncoder\n\n# Advanced analysis libraries\nimport umap\nimport shap\nimport scipy.cluster.hierarchy as sch\nimport scipy.stats as stats\nfrom scipy.spatial.distance import squareform\n\n# CatBoost\nimport catboost\nfrom catboost import CatBoostRegressor, Pool\nfrom sklearn.multioutput import MultiOutputRegressor","metadata":{"_uuid":"9fa2ebdb-95ef-4edc-87e2-ef0b7cb124d5","_cell_guid":"d9e2ad50-fd36-4c9d-99b3-06ceb4aac464","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:24:33.789138Z","iopub.execute_input":"2023-11-15T09:24:33.790276Z","iopub.status.idle":"2023-11-15T09:24:48.549754Z","shell.execute_reply.started":"2023-11-15T09:24:33.790223Z","shell.execute_reply":"2023-11-15T09:24:48.548268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading additional data required for mapping\nBASE = '/kaggle/input/open-problems-single-cell-perturbations/'","metadata":{"_uuid":"87f7a5a6-2d0a-43bc-9cc7-c997e1f82ea9","_cell_guid":"da4b12be-defb-446e-9663-d595ddfa1bfc","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:24:48.557015Z","iopub.execute_input":"2023-11-15T09:24:48.557847Z","iopub.status.idle":"2023-11-15T09:24:48.563477Z","shell.execute_reply.started":"2023-11-15T09:24:48.557811Z","shell.execute_reply":"2023-11-15T09:24:48.562239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Displaying the first few rows of the ID map for inspection\ntrain_data = pd.read_parquet(f'{BASE}de_train.parquet')","metadata":{"_uuid":"b7d63bfd-f163-4790-85f0-67bd50429e33","_cell_guid":"917043ea-3bc2-4674-9cbf-44c01dabe18c","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:24:48.564710Z","iopub.execute_input":"2023-11-15T09:24:48.565005Z","iopub.status.idle":"2023-11-15T09:24:50.250997Z","shell.execute_reply.started":"2023-11-15T09:24:48.564977Z","shell.execute_reply":"2023-11-15T09:24:50.249703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"----\n# EDA","metadata":{"_uuid":"8b9ab052-1e20-4433-bca1-fa93c4cc8c95","_cell_guid":"1d6e3ed4-468f-4540-a10d-6ecb1432a315","trusted":true}},{"cell_type":"code","source":"# Display general information about the dataset\ntrain_data.info()","metadata":{"_uuid":"cf31ec9d-2f6a-408f-aa21-5584a7f5e4ee","_cell_guid":"e8058336-a7eb-47ee-9e9c-900b110e90fe","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:24:50.255120Z","iopub.execute_input":"2023-11-15T09:24:50.255718Z","iopub.status.idle":"2023-11-15T09:24:51.608331Z","shell.execute_reply.started":"2023-11-15T09:24:50.255637Z","shell.execute_reply":"2023-11-15T09:24:51.606895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.pairplot(train_data.sample(613)[['A1BG', 'A1BG-AS1', 'A2M', 'cell_type']], hue='cell_type', plot_kws={'alpha':0.5})\nplt.suptitle('Pairwise Relationships among Selected Genes Across Cell Types')\nplt.show()","metadata":{"_uuid":"425af249-ad7a-4efc-8058-c9d09feb9178","_cell_guid":"acdfc6dd-93e5-4917-a16b-ed2c01635bee","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:24:51.609954Z","iopub.execute_input":"2023-11-15T09:24:51.610945Z","iopub.status.idle":"2023-11-15T09:24:57.072046Z","shell.execute_reply.started":"2023-11-15T09:24:51.610891Z","shell.execute_reply":"2023-11-15T09:24:57.071020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x='sm_name', data=train_data)\nplt.title('Distribution of Small Molecules')\nplt.xticks(rotation=90)\nplt.show()","metadata":{"_uuid":"84024f9f-9fc1-4dee-af68-f988f1c2d48c","_cell_guid":"f1b5de04-2949-4144-afb8-f596d2e99f98","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:24:57.073281Z","iopub.execute_input":"2023-11-15T09:24:57.073626Z","iopub.status.idle":"2023-11-15T09:24:58.656863Z","shell.execute_reply.started":"2023-11-15T09:24:57.073592Z","shell.execute_reply":"2023-11-15T09:24:58.655107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Adjust the sample size to the number of rows in your DataFrame if needed\nsample_size = min(1000, len(train_data))  # Ensuring sample size is not greater than the dataset size\n\n# Perform t-SNE\ntsne = TSNE(n_components=2, perplexity=30, random_state=42)\ntsne_results = tsne.fit_transform(train_data.drop(['cell_type', 'sm_name', 'sm_lincs_id', 'SMILES', 'control'], axis=1).sample(sample_size))\n\n# Visualization\ntrain_data_subset = train_data.sample(sample_size)  # Ensure you're sampling the same rows for coloring by cell_type\n\nplt.figure(figsize=(16,10))\nsns.scatterplot(\n    x=tsne_results[:,0], y=tsne_results[:,1],\n    hue=train_data_subset['cell_type'],\n    palette=sns.color_palette(\"hsv\", len(train_data_subset['cell_type'].unique())),\n    legend=\"full\",\n    alpha=0.7\n)\nplt.title('t-SNE Visualization of Gene Expressions')\nplt.show()","metadata":{"_uuid":"862b360f-b551-4b44-95ae-ba1893cf9943","_cell_guid":"5274d56d-fb46-4b0d-a138-e0591f9c4f0a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:24:58.658923Z","iopub.execute_input":"2023-11-15T09:24:58.659365Z","iopub.status.idle":"2023-11-15T09:25:04.479741Z","shell.execute_reply.started":"2023-11-15T09:24:58.659325Z","shell.execute_reply":"2023-11-15T09:25:04.478252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate correlations on a smaller sample if needed\n#sample_size = min(500, len(train_data))\n#sampled_data = train_data.sample(sample_size).drop(['cell_type', 'sm_name', 'sm_lincs_id', 'SMILES', 'control'], axis=1)\n#corr = sampled_data.corr()\n#\n# Filter out the lower triangle of the correlation matrix to avoid duplicate edges in the graph\n#tri_upper = np.triu(np.ones(corr.shape), k=1).astype(bool)\n#\n# Apply the threshold to the upper triangle\n#threshold = 0.8\n#filtered_corr = corr.where((corr.abs() > threshold) & tri_upper)\n\n# Create a list of edges and corresponding weights\n#edges = []\n#for src, dst in zip(*np.where(filtered_corr.notnull())):\n#    gene1, gene2 = sampled_data.columns[src], sampled_data.columns[dst]\n#    weight = filtered_corr.iloc[src, dst]\n#    edges.append((gene1, gene2, weight))\n#\n# Create the network graph\n#G = nx.Graph()\n#G.add_weighted_edges_from(edges)\n#\n# NOW create the network graph using PyVis\n#net = Network(notebook=True, height=\"750px\", width=\"100%\")\n#net.from_nx(G)\n#net.show(\"gene_correlation.html\")","metadata":{"_uuid":"bee35af2-9251-4b87-a35b-a7234eee54ea","_cell_guid":"f2f62f66-8fc5-40ab-8f15-b853f41099bd","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:04.481912Z","iopub.execute_input":"2023-11-15T09:25:04.482459Z","iopub.status.idle":"2023-11-15T09:25:04.490796Z","shell.execute_reply.started":"2023-11-15T09:25:04.482421Z","shell.execute_reply":"2023-11-15T09:25:04.489371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# UMAP for dimensionality reduction\nreducer = umap.UMAP(random_state=42)\nembedding = reducer.fit_transform(train_data.drop(['cell_type', 'sm_name', 'sm_lincs_id', 'SMILES', 'control'], axis=1))\n\n# Using DBSCAN for clustering UMAP output\nclustering = DBSCAN(eps=0.5, min_samples=10).fit(embedding)\ntrain_data['umap_cluster'] = clustering.labels_\n\n# Visualizing UMAP output with clusters\nplt.figure(figsize=(12, 10))\nsns.scatterplot(\n    x=embedding[:, 0], \n    y=embedding[:, 1], \n    hue=train_data['umap_cluster'], \n    palette=sns.color_palette(\"hsv\", len(set(clustering.labels_))),\n    legend='full'\n)\nplt.title('UMAP projection of the Dataset', fontsize=20)\nplt.xlabel('UMAP Dimension 1')\nplt.ylabel('UMAP Dimension 2')\nplt.show()\n\n\n# Prepare data for SHAP analysis\nX = train_data.drop(['cell_type', 'sm_name', 'sm_lincs_id', 'SMILES', 'control', 'umap_cluster'], axis=1)\ny = train_data['umap_cluster']\n\n# Train a model for SHAP analysis\nmodel = RandomForestClassifier(random_state=0, n_estimators=100)\nmodel.fit(X, y)\n\n# Explain the model's predictions using SHAP\nexplainer = shap.TreeExplainer(model)\nshap_values = explainer.shap_values(X.sample(100))  # Using a sample for speed\n\n# Visualize the first prediction's explanation\nshap.initjs()\nshap.force_plot(explainer.expected_value[0], shap_values[0][0], X.iloc[0])","metadata":{"_uuid":"8cd4fd73-9b95-47e0-874a-8d4ea4004489","_cell_guid":"c0814c8a-6263-4859-b9bb-a215017dcbc6","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:04.492658Z","iopub.execute_input":"2023-11-15T09:25:04.493235Z","iopub.status.idle":"2023-11-15T09:25:26.860897Z","shell.execute_reply.started":"2023-11-15T09:25:04.493172Z","shell.execute_reply":"2023-11-15T09:25:26.859403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for gene in ['A1BG', 'A1BG-AS1', 'A2M']:  # Replace with genes of interest\n    result = stats.kruskal(*[group[gene].values for name, group in train_data.groupby('cell_type')])\n    print(f\"Kruskal-Wallis test for {gene}: H-statistic={result.statistic}, p-value={result.pvalue}\")","metadata":{"_uuid":"a1f63f67-02b8-434b-bea1-420464de0109","_cell_guid":"825d5ea0-6956-404a-b316-68f67001a42d","collapsed":false,"_kg_hide-input":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:26.862921Z","iopub.execute_input":"2023-11-15T09:25:26.863516Z","iopub.status.idle":"2023-11-15T09:25:27.009210Z","shell.execute_reply.started":"2023-11-15T09:25:26.863462Z","shell.execute_reply":"2023-11-15T09:25:27.007533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le = LabelEncoder()\ntrain_data['cell_type_encoded'] = le.fit_transform(train_data['cell_type'])\n\n# Define features and target\nX = train_data.drop(['cell_type', 'cell_type_encoded', 'sm_name', 'sm_lincs_id', 'SMILES', 'control'], axis=1)\ny = train_data['cell_type_encoded']\n\n# Split the data into training and test sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n\n# Fit the RandomForest model on the training data\nrf = RandomForestClassifier(n_estimators=100, random_state=42)\nrf.fit(X_train, y_train)\n\n# Extract feature importances from the model\nimportances = rf.feature_importances_\n\n# Plot the top N feature importances\ntop_n = 20  # Specify how many top features you'd like to visualize\nindices = np.argsort(importances)[-top_n:]\nplt.figure(figsize=(15, 7))\nplt.title('Top Feature Importances')\nplt.barh(range(len(indices)), importances[indices], color='b', align='center')\nplt.yticks(range(len(indices)), [X_train.columns[i] for i in indices])\nplt.xlabel('Relative Importance')\nplt.show()","metadata":{"_uuid":"3d5215fb-50b7-4f26-b954-24c069bf0c69","_cell_guid":"fd9fe387-29fe-4096-80c6-615131765a71","collapsed":false,"_kg_hide-input":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:27.010890Z","iopub.execute_input":"2023-11-15T09:25:27.011373Z","iopub.status.idle":"2023-11-15T09:25:30.773336Z","shell.execute_reply.started":"2023-11-15T09:25:27.011339Z","shell.execute_reply":"2023-11-15T09:25:30.772200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{"_uuid":"8b38df8a-0d49-4356-912d-cb41b8a28a05","_cell_guid":"170afc69-5b4c-45a4-a575-fc9de1311814","trusted":true}},{"cell_type":"code","source":"# Displaying the shape of the training and ID map datasets\nid_map = pd.read_csv(f'{BASE}id_map.csv')","metadata":{"_uuid":"387b874e-9795-409b-b3d8-74d37fa12489","_cell_guid":"ea0f271d-329c-4df0-b834-09ea8d306c4a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:30.774845Z","iopub.execute_input":"2023-11-15T09:25:30.776422Z","iopub.status.idle":"2023-11-15T09:25:30.786122Z","shell.execute_reply.started":"2023-11-15T09:25:30.776362Z","shell.execute_reply":"2023-11-15T09:25:30.785142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id_map.head()","metadata":{"_uuid":"a6efe78e-3e08-4a59-a0f8-c82cb97c83ab","_cell_guid":"6c58b296-ff73-4970-93a8-d696399c8623","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:30.792396Z","iopub.execute_input":"2023-11-15T09:25:30.793761Z","iopub.status.idle":"2023-11-15T09:25:30.808679Z","shell.execute_reply.started":"2023-11-15T09:25:30.793700Z","shell.execute_reply":"2023-11-15T09:25:30.807255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_data.head())","metadata":{"_uuid":"b7d840b5-8a11-4128-a711-83686bee92ad","_cell_guid":"b808bbdc-3e72-413f-807a-70ddff6f20dc","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:30.810298Z","iopub.execute_input":"2023-11-15T09:25:30.811073Z","iopub.status.idle":"2023-11-15T09:25:30.840401Z","shell.execute_reply.started":"2023-11-15T09:25:30.811039Z","shell.execute_reply":"2023-11-15T09:25:30.839377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_data.describe())","metadata":{"_uuid":"607adea3-c676-4403-af1a-2d6cb4b6d75a","_cell_guid":"e45b46de-3db5-43e5-a896-5a94f43f2d61","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:25:30.842234Z","iopub.execute_input":"2023-11-15T09:25:30.842570Z","iopub.status.idle":"2023-11-15T09:26:12.341104Z","shell.execute_reply.started":"2023-11-15T09:25:30.842540Z","shell.execute_reply":"2023-11-15T09:26:12.339868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Training and Evaluation","metadata":{"_uuid":"66ca301d-bcf8-40cc-a1ad-acd2798152de","_cell_guid":"682a046a-ca74-4cf4-ba8d-05335e7a84e1","trusted":true}},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\nplt.hist(train_data['A1BG'], bins=50, color='blue', alpha=0.7)\nplt.title(\"Distribution of Gene Expression (A1BG)\")\nplt.xlabel(\"Gene Expression Value\")\nplt.ylabel(\"Frequency\")\nplt.show()","metadata":{"_uuid":"739b7f92-b859-41af-b71e-a4ccd7a498f6","_cell_guid":"8c646986-81b9-4730-b7a2-30dbc785cbfa","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:12.342916Z","iopub.execute_input":"2023-11-15T09:26:12.343312Z","iopub.status.idle":"2023-11-15T09:26:12.718784Z","shell.execute_reply.started":"2023-11-15T09:26:12.343277Z","shell.execute_reply":"2023-11-15T09:26:12.717610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Training Data Shape:\", train_data.shape)\nprint(\"Id Map Data Shape:\", id_map.shape)","metadata":{"_uuid":"831fe51a-dfef-4e1e-9b99-ec0b57850ad5","_cell_guid":"c9be5a80-1d5c-4da1-986f-2cad088c21c1","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:12.720613Z","iopub.execute_input":"2023-11-15T09:26:12.721013Z","iopub.status.idle":"2023-11-15T09:26:12.726739Z","shell.execute_reply.started":"2023-11-15T09:26:12.720979Z","shell.execute_reply":"2023-11-15T09:26:12.725927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t0start = time.time()","metadata":{"_uuid":"96010ed1-9153-4d03-8f6c-b8d7197d22d3","_cell_guid":"b03eb6ee-9901-433f-aaf6-0b0f59031beb","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:12.728161Z","iopub.execute_input":"2023-11-15T09:26:12.728763Z","iopub.status.idle":"2023-11-15T09:26:12.743394Z","shell.execute_reply.started":"2023-11-15T09:26:12.728729Z","shell.execute_reply":"2023-11-15T09:26:12.742451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\ndf_de_train = pd.read_parquet(fn)# , index_col = 0)\nprint(df_de_train.shape)\ndf_de_train","metadata":{"_uuid":"03abd835-c72c-4c37-b1e7-115c09ef1554","_cell_guid":"306ba247-79e8-462b-8216-8af0a9382838","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:12.745006Z","iopub.execute_input":"2023-11-15T09:26:12.745651Z","iopub.status.idle":"2023-11-15T09:26:14.507550Z","shell.execute_reply.started":"2023-11-15T09:26:12.745615Z","shell.execute_reply":"2023-11-15T09:26:14.506621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/open-problems-single-cell-perturbations/id_map.csv'\ndf_id_map = pd.read_csv(fn)\nprint(df_id_map.shape)\ndisplay(df_id_map)\nfn = '/kaggle/input/open-problems-single-cell-perturbations/sample_submission.csv'\ndf = pd.read_csv(fn, index_col = 0)\nprint(df.shape)\ndf","metadata":{"_uuid":"bbf930dd-3576-442a-9345-80e8bb7be1b3","_cell_guid":"a49f8819-b518-4ac9-b96b-f0da6b5b2019","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:14.508989Z","iopub.execute_input":"2023-11-15T09:26:14.510302Z","iopub.status.idle":"2023-11-15T09:26:19.392354Z","shell.execute_reply.started":"2023-11-15T09:26:14.510264Z","shell.execute_reply":"2023-11-15T09:26:19.390806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \nn_components = 35\n\npredict_method = 'train_aggregation_by_compounds_with_denoising_TSVD'\n\nif '_pca' in predict_method:\n    str_inf_target_dimred = 'PCA' \n    reducer = PCA(n_components=n_components )\nelif '_ICA' in predict_method:\n    str_inf_target_dimred = 'ICA' \n    reducer = FastICA(n_components=n_components, random_state=0, whiten='unit-variance')\nelif '_TSVD' in predict_method:\n    str_inf_target_dimred = 'TSVD' \n    reducer = TruncatedSVD(n_components=n_components, n_iter=7, random_state=42)\nelse:\n    str_inf_target_dimred = ''\n    \nprint(str_inf_target_dimred, reducer)","metadata":{"_uuid":"5d1fb9cd-d24a-4b6b-8f47-56b07f448b3c","_cell_guid":"75b58c9d-e04f-4d31-ad27-ee367d0944a2","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:19.394079Z","iopub.execute_input":"2023-11-15T09:26:19.394504Z","iopub.status.idle":"2023-11-15T09:26:19.407151Z","shell.execute_reply.started":"2023-11-15T09:26:19.394467Z","shell.execute_reply":"2023-11-15T09:26:19.405621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Y = df_de_train.iloc[:,5:].values\nYr = reducer.fit_transform(Y)\npd.DataFrame(Yr).corr(method = 'spearman').round(2)","metadata":{"_uuid":"2e32b93b-908f-4749-a957-869bfe10654a","_cell_guid":"0b788811-db75-4fca-bc75-e91db901a2d4","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:19.408975Z","iopub.execute_input":"2023-11-15T09:26:19.409505Z","iopub.status.idle":"2023-11-15T09:26:21.858886Z","shell.execute_reply.started":"2023-11-15T09:26:19.409461Z","shell.execute_reply":"2023-11-15T09:26:21.857362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nn_components_for_compound_encoding = 3\nquantile = 0.54\ndf_tmp = pd.DataFrame(Yr[:, :n_components_for_compound_encoding  ], index = df_de_train.index  )\ndf_tmp['column for aggregation'] = df_de_train['sm_name']\ndf_compound_encoded = df_tmp.groupby('column for aggregation').quantile( quantile )\nprint('df_compound_encoded.shape', df_compound_encoded.shape )\ndisplay( df_compound_encoded )\n\nX = pd.DataFrame(index = df_de_train['sm_name']) # \nX['IX'] = np.arange(len(df_de_train))\nX = X.join( df_compound_encoded , how = 'left').sort_values('IX')\nX = X.iloc[:,1:]\nX = X.values \nX.shape","metadata":{"_uuid":"e4922094-a16c-47ad-8d50-9cec8d4620b2","_cell_guid":"13c39e1e-5a0e-4b06-9b93-12a08504055b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:21.860393Z","iopub.execute_input":"2023-11-15T09:26:21.860773Z","iopub.status.idle":"2023-11-15T09:26:21.896539Z","shell.execute_reply.started":"2023-11-15T09:26:21.860711Z","shell.execute_reply":"2023-11-15T09:26:21.895734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"enc = OrdinalEncoder()\nX = enc.fit_transform(df_de_train[ ['cell_type', 'sm_name']])\nX.shape","metadata":{"_uuid":"bb9e27a6-9b69-413c-a968-ab0f043b7ed7","_cell_guid":"61976963-0f2c-41ca-999e-10311a8d2c3b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:21.898142Z","iopub.execute_input":"2023-11-15T09:26:21.898776Z","iopub.status.idle":"2023-11-15T09:26:21.911375Z","shell.execute_reply.started":"2023-11-15T09:26:21.898746Z","shell.execute_reply":"2023-11-15T09:26:21.910121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nn_components_for_compound_encoding = 20\nquantile = 0.54\ndf_tmp = pd.DataFrame(Yr[:, :n_components_for_compound_encoding  ], index = df_de_train.index  )\ndf_tmp['column for aggregation'] = df_de_train['sm_name']\ndf_compound_encoded = df_tmp.groupby('column for aggregation').mean() # quantile( quantile )\nprint('df_compound_encoded.shape', df_compound_encoded.shape )\n\nX = pd.DataFrame(index = df_de_train['sm_name']) # \nX['IX'] = np.arange(len(df_de_train))\nX = X.join( df_compound_encoded , how = 'left').sort_values('IX')\nX = X.iloc[:,1:]\nX = X.values \nprint( X.shape )\n\nX_submit = pd.DataFrame(index = df_id_map['sm_name']) # \nX_submit['IX'] = np.arange(len(X_submit))\nX_submit = X_submit.join( df_compound_encoded , how = 'left').sort_values('IX')\nX_submit = X_submit.iloc[:,1:]\nX_submit = X_submit.values \nprint( X_submit.shape )","metadata":{"_uuid":"fe48258c-3317-4b14-af29-56ac581a1713","_cell_guid":"97c0e760-d485-4c9a-88b6-c7cca7b94d29","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:21.913567Z","iopub.execute_input":"2023-11-15T09:26:21.914575Z","iopub.status.idle":"2023-11-15T09:26:21.941345Z","shell.execute_reply.started":"2023-11-15T09:26:21.914519Z","shell.execute_reply":"2023-11-15T09:26:21.939770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \nn_splits = 3\nkf = KFold(n_splits=n_splits, random_state = 42, shuffle = True )\nIX = pd.DataFrame(); IX['IX'] = range(len(df_de_train))\n\nalpha_regularization_for_linear_models = 100000\nmodel = Ridge(alpha=alpha_regularization_for_linear_models)\n\nY = df_de_train.iloc[:,5:].values\n\nY_reduced_submit = np.zeros( (len(df_id_map) , Yr.shape[1] )   ); cnt_blend_submit = 0\nY_submit = np.zeros( (len(df_id_map) , 18211 )   ); cnt_blend_submit = 0\nY_oof = np.zeros( (len(df_de_train) , 18211 )   ); cnt_blend_submit = 0","metadata":{"_uuid":"bf46dc8a-fdda-48a8-8429-2fdc5894d7f3","_cell_guid":"6fd7b11a-c40f-446a-8b2f-1a3eb08a7819","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:21.942965Z","iopub.execute_input":"2023-11-15T09:26:21.943327Z","iopub.status.idle":"2023-11-15T09:26:22.008201Z","shell.execute_reply.started":"2023-11-15T09:26:21.943297Z","shell.execute_reply":"2023-11-15T09:26:22.006865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X.shape)\nfor i, (train_index, test_index) in enumerate(kf.split(IX)):\n    print(f\"Fold {i}:\", len(test_index) )\n    \n    Yr_train = reducer.fit_transform(Y[train_index,:])\n    Yr_test = reducer.transform(Y[test_index,:])\n    X_train = X[train_index,:]\n    X_test = X[test_index,:]\n    model.fit(X_train, Yr_train)\n    Yr_pred = model.predict(X_test) # , Yr_train)\n    r2 = r2_score(Yr_test,  Yr_pred )\n    print('r2 test:', r2)\n    r2_test = r2\n    Y_oof[test_index,: ] = reducer.inverse_transform(Yr_pred)\n    \n    Yr_pred = model.predict(X_train) # , Yr_train)\n    r2 = r2_score(Yr_train,  Yr_pred )\n    print('r2 train', r2)\n    Y_reduced_submit =  (Y_reduced_submit * cnt_blend_submit + model.predict(X_submit) ) / ( cnt_blend_submit + 1)\n    Y_submit =  (Y_submit * cnt_blend_submit + reducer.inverse_transform(model.predict(X_submit) ) ) / ( cnt_blend_submit + 1)\n    cnt_blend_submit += 1\n    \ns = mean_squared_error(Y,Y_oof, squared = False)\nprint(s)","metadata":{"_uuid":"67657a47-8497-4db8-821c-191feb6f568a","_cell_guid":"16afdc61-40dd-4850-a7d9-f5edeb614ea2","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:22.010096Z","iopub.execute_input":"2023-11-15T09:26:22.010452Z","iopub.status.idle":"2023-11-15T09:26:28.611463Z","shell.execute_reply.started":"2023-11-15T09:26:22.010422Z","shell.execute_reply":"2023-11-15T09:26:28.610253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \nn_splits = 3\nkf = KFold(n_splits=n_splits, random_state = 42, shuffle = True )\nIX = pd.DataFrame(); IX['IX'] = range(len(df_de_train))\n\nalpha_regularization_for_linear_models = 100000\nmodel = Ridge(alpha=alpha_regularization_for_linear_models)\n\ncategorical_features = ['cell_type','sm_name']\nmodel = MultiOutputRegressor( CatBoostRegressor(cat_features=categorical_features, verbose = 0,  # Categorical features ) ) \n                        iterations=300,  # Number of boosting iterations\n                          depth=5,        # Depth of the trees\n                          learning_rate=0.015,  # Learning rate\n                          loss_function='RMSE'))  # Specify your loss function (e.g., RMSE for regression)\n#                           verbose=0) ) # Set verbose to 0 to suppress output\n\n\nY = df_de_train.iloc[:,5:].values","metadata":{"_uuid":"9726065c-5d50-49c4-a31f-69d90379bdb0","_cell_guid":"d3421dee-606e-4513-b28f-9c3df2214f40","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:28.612951Z","iopub.execute_input":"2023-11-15T09:26:28.613329Z","iopub.status.idle":"2023-11-15T09:26:28.663288Z","shell.execute_reply.started":"2023-11-15T09:26:28.613295Z","shell.execute_reply":"2023-11-15T09:26:28.661708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Y_reduced_submit = np.zeros( (len(df_id_map) , Yr.shape[1] )   ); cnt_blend_submit = 0\nY_submit = np.zeros( (len(df_id_map) , 18211 )   ); cnt_blend_submit = 0\nY_oof = np.zeros( (len(df_de_train) , 18211 )   ); cnt_blend_submit = 0\n\ndf_train = df_de_train[['cell_type','sm_name']]\n\nX_submit = df_id_map[ ['cell_type','sm_name'] ]","metadata":{"_uuid":"dab72540-f315-4208-a90f-2ed74c4dc50d","_cell_guid":"8c92ffdd-db1f-4524-8440-e686b63f92a5","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:28.665898Z","iopub.execute_input":"2023-11-15T09:26:28.666468Z","iopub.status.idle":"2023-11-15T09:26:28.687026Z","shell.execute_reply.started":"2023-11-15T09:26:28.666415Z","shell.execute_reply":"2023-11-15T09:26:28.685799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X.shape)\nfor i, (train_index, test_index) in enumerate(kf.split(IX)):\n    print(f\"Fold {i}:\", len(test_index) )\n    \n    Yr_train = reducer.fit_transform(Y[train_index,:])\n    Yr_test = reducer.transform(Y[test_index,:])\n    X_train = df_de_train[categorical_features].iloc[train_index,:]\n    X_test = df_de_train[categorical_features].iloc[test_index,:]\n    \n    model.fit(X_train, Yr_train)\n    Yr_pred = model.predict(X_test) # , Yr_train)\n    r2 = r2_score(Yr_test,  Yr_pred )\n    print('r2 test:', r2)\n    r2_test = r2\n    Y_oof[test_index,: ] = reducer.inverse_transform(Yr_pred)\n    \n    Yr_pred = model.predict(X_train) # , Yr_train)\n    r2 = r2_score(Yr_train,  Yr_pred )\n    print('r2 train', r2)\n    Y_reduced_submit =  (Y_reduced_submit * cnt_blend_submit + model.predict(X_submit) ) / ( cnt_blend_submit + 1)\n    Y_submit =  (Y_submit * cnt_blend_submit + reducer.inverse_transform(model.predict(X_submit) ) ) / ( cnt_blend_submit + 1)\n    cnt_blend_submit += 1\n    \n\ns = mean_squared_error(Y,Y_oof, squared = False)\nprint(s)","metadata":{"_uuid":"3a580f57-dca1-490f-82d7-999e519f0070","_cell_guid":"25e04fe1-fc7f-40fa-9357-dc47e4d2d8f6","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:26:28.688688Z","iopub.execute_input":"2023-11-15T09:26:28.689092Z","iopub.status.idle":"2023-11-15T09:27:01.339690Z","shell.execute_reply.started":"2023-11-15T09:26:28.689057Z","shell.execute_reply":"2023-11-15T09:27:01.338058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission Preparation","metadata":{"_uuid":"5927397c-9e46-4948-9184-f41eb076fbb6","_cell_guid":"7c744a7a-ea08-4d0b-950c-4216dfe47369","trusted":true}},{"cell_type":"code","source":"df_submit = pd.DataFrame(Y_submit, columns = df_de_train.columns[5:])\ndf_submit.index.name = 'id'\nprint( df_submit.shape )\ndisplay(df_submit)\ndf_submit.to_csv('submission.csv')","metadata":{"_uuid":"47765817-1607-4f2b-89cd-cbef1d271826","_cell_guid":"c0d5d008-34ab-41dc-9869-8b54289cf165","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-11-15T09:27:01.341938Z","iopub.execute_input":"2023-11-15T09:27:01.342503Z","iopub.status.idle":"2023-11-15T09:27:14.886596Z","shell.execute_reply.started":"2023-11-15T09:27:01.342452Z","shell.execute_reply":"2023-11-15T09:27:14.885505Z"},"trusted":true},"execution_count":null,"outputs":[]}]}