{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Data Loading**:\n- Begin by importing the necessary libraries, including Pandas for data manipulation.\n- Load the provided datasets using Pandas `read_parquet` and `read_csv` functions. These datasets include differential expression data, observation metadata, and an ID mapping file.","metadata":{}},{"cell_type":"code","source":"# Import necessary libraries\nimport pandas as pd\n\n# Load the provided datasets\nde_train = pd.read_parquet('/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet')\nadata_obs = pd.read_csv('/kaggle/input/open-problems-single-cell-perturbations/adata_obs_meta.csv')\nid_map = pd.read_csv('/kaggle/input/open-problems-single-cell-perturbations/id_map.csv')\n","metadata":{"execution":{"iopub.status.busy":"2023-09-24T10:06:49.852525Z","iopub.execute_input":"2023-09-24T10:06:49.853047Z","iopub.status.idle":"2023-09-24T10:06:52.355848Z","shell.execute_reply.started":"2023-09-24T10:06:49.853006Z","shell.execute_reply":"2023-09-24T10:06:52.354189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Preprocessing**:\n- Preprocessing is essential to prepare the data for modeling. It may involve tasks such as data cleaning and normalization.\n- In this example, we demonstrate how to handle missing values by dropping rows with missing values in the `de_train` dataset.","metadata":{}},{"cell_type":"code","source":"# Preprocess the data (e.g., data cleaning, normalization)\n# Example: Drop rows with missing values in de_train\nde_train.dropna(inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-09-24T10:15:37.439941Z","iopub.execute_input":"2023-09-24T10:15:37.440432Z","iopub.status.idle":"2023-09-24T10:15:37.500698Z","shell.execute_reply.started":"2023-09-24T10:15:37.440398Z","shell.execute_reply":"2023-09-24T10:15:37.499134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Data Splitting**:\n- For modeling purposes, data is typically split into training and testing sets. Here, we split the `de_train` dataset into training and testing sets based on the cell type.\n- We create a `train` dataset containing rows with cell types 'T' and 'NK' and a `test` dataset containing cell types 'B' and 'myeloid.'","metadata":{}},{"cell_type":"code","source":"# Split data into train and test sets based on cell type\ntrain = de_train[de_train['cell_type'].isin(['T', 'NK'])]\ntest = de_train[de_train['cell_type'].isin(['B', 'myeloid'])]","metadata":{"execution":{"iopub.status.busy":"2023-09-24T10:16:54.708929Z","iopub.execute_input":"2023-09-24T10:16:54.709484Z","iopub.status.idle":"2023-09-24T10:16:54.724863Z","shell.execute_reply.started":"2023-09-24T10:16:54.709447Z","shell.execute_reply":"2023-09-24T10:16:54.723033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Step 3: Model Building\n\n**Model Splitting**:\n- Split the training data into training and validation sets. This allows us to train the model on a portion of the data and evaluate its performance on a separate dataset.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}