{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center>\n</center>\n<div style=\"padding:20px; \n            color:#150d0a;\n            margin:10px;\n            font-size:220%;\n            text-align:center;\n            display:fill;\n            border-radius:20px;\n            border-width: 5px;\n            border-style: solid;\n            border-color: #150d0a;\n            background-color:#a98ebf;\n            overflow:hidden;\n            font-weight:500\">Open Problems – Single-Cell Perturbations</div>","metadata":{}},{"cell_type":"markdown","source":"### <font color='289C4E'>Table of contents<font><a class='anchor' id='top'></a>\n- [Introduction to the Competition](#chapter1)\n- [Single-Cell Analysis: Understanding Human Biology at the Cellular Level](#chapter2)\n- [The Challenge of Predicting Chemical Perturbations](#chapter3)\n- [Existing Methods and Challenges](#chapter4)\n- [The Dataset](#chapter5)\n- [Technical Details of the Experiment](#chapter6)\n- [Cell Multiplexing and Differential Expression Analysis](#chapter7)\n- [Data Splits](#chapter8)\n- [Data](#chapter9)","metadata":{}},{"cell_type":"markdown","source":"If you're like me and got a bit puzzled 😕 when reading all the competition stuff, don't worry! I'm here to help you understand what this competition is all about. I'll explain the experiment, break down the tricky words, talk about the data they've given us, and tell you what they're trying to achieve in simple terms.","metadata":{}},{"cell_type":"markdown","source":"## 1. Introduction to the Competition <a class=\"anchor\"  id=\"chapter1\"></a>[↑](#top)\n\nThe overarching goal of the competition is to use machine learning in the development of new medicines. The challenge is to predict how chemical substances affect the behaviour of human cells at a molecular level. We can think of this like, when a person takes an ibuprofen tablet it affects our cells in some way to relieve us from pain, fever, and inflammation. Similarly in this competition, we are trying to predict the effects of compounds on human cells at the molecular level. This is a very crucial step in drug discovery as understanding this molecular effect can lead to the discovery of new treatments for various diseases.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## Single-Cell Analysis: Understanding Human Biology at the Cellular Level <a class=\"anchor\"  id=\"chapter2\"></a>[↑](#top)\n\nHuman biology is very complex, our body is made up of around 37 trillion individual cells. These cells work together to from tissues, organs and systems, and their behaviour is controlled by the information encoded in their DNA and RNA, (Think of DNA as the \"instruction manual\" for living organisms, like a recipe book for a cake. RNA is like a \"messenger\" that carries out instructions from DNA to make proteins). Recent technological advances have allowed scientists to study individual cells at a molecular level, looking at their DNA, RNA, and protein profiles. This field is known as \"single-cell analysis.\"\n","metadata":{}},{"cell_type":"markdown","source":"## Chemical Perturbations <a class=\"anchor\"  id=\"chapter3\"></a>[↑](#top)\n\nWhen a single cell is treated with a chemical/compound this treatment is referred to as \"chemical perturbations\". These treatments can have a significant impact on cell behaviour, being able to predict the response of cells to different treatments would be crucial for advancing the field of drug development.\n","metadata":{}},{"cell_type":"markdown","source":"## The Challenge of Predicting Chemical Perturbations <a class=\"anchor\"  id=\"chapter4\"></a>[↑](#top)\n\nThe task of predicting the impact of chemical perturbation on a single cell is very challenging, it involves understanding the affects on gene expression of the cells. Genes are like the instructions or recipes in a cookbook. In the cell's \"cookbook,\" which is the DNA, genes are the specific recipes that provide the step-by-step instructions for making important molecules like proteins, gene expression is the cell's way of reading and using those recipes to make the proteins required for its activities. By treating the cell with a compound scientists are trying to understand the effect on gene expression. Conducting experiments to gather this gene expression data is quite expensive and labour-intensive. Also, not all types of cells are suitable for high-throughput transcriptomic screening, which is a sophisticated and automated laboratory technique used to study the expression of genes on a large scale.\n","metadata":{}},{"cell_type":"markdown","source":"## Existing Methods and Challenges <a class=\"anchor\"  id=\"chapter5\"></a>[↑](#top)\n\nThere has been much work done in the field of prediction drug perturbations, but a significant challenge is the lack of diverse benchmarking datasets that represent various cells in the human body. The largest available dataset, called Connectivity Map (CMap) primarily consists of data from cancer cells, which are abnormal human cells.","metadata":{}},{"cell_type":"markdown","source":"## The Dataset <a class=\"anchor\"  id=\"chapter6\"></a>[↑](#top)\n\nA dataset was created using human peripheral blood mononuclear cells (PBMCs) from three healthy human donors. PBMCs are a special type of cells that can be found in human blood. They include white blood cells like lymphocytes and monocytes, and they play an important role in your immune system, helping your body fight off infections and diseases. Scientists often study PBMCs to learn more about how the immune system works and how it responds to different conditions or treatments.\nThese cells were treated with 144 different compounds, and their gene expression profiles were measured after 24 hours of treatment.","metadata":{}},{"cell_type":"markdown","source":"# Technical Details of the Experiment <a class=\"anchor\"  id=\"chapter7\"></a>[↑](#top)\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4308072%2Fdbce2b5e0a2b8b9691502b31c5feb5b6%2FScreenshot%202023-08-25%20at%206.20.53%20PM.png?generation=1693002071834598&alt=media\"/>\n\n\nFirstly let us break down some technical terms to understand this experiment.\n\n**96-Well Plate**: A 96-well plate is a tray with 96 tiny wells or compartments, kind of like an ice cube tray but smaller. Each well can hold a small amount of liquid, like a drop of water.\n\n**Positive Controls**: Positive controls are known examples or standards. Scientists use positive controls to make sure their experiments are working correctly. In this experiment, two compounds, Dabrafenib and Belinostat, are used as positive controls. These are well-understood substances with known effects on gene expression and cellular processes\n\n**Negative Controls**: Negative controls are like blank tests. Imagine if you wanted to see if a chemical changed the colour of a water. You'd compare it to a glass of water with nothing added. Negative controls in the experiment serve a similar purpose. They help scientists see what happens when nothing special is added.\n\n\nThe experimental design involved plating PBMCs (a collection of different cell types) on 96-well plates, with specific columns dedicated to positive and negative controls. The remaining wells were allocated to different compounds for treatment.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"\n## Cell Multiplexing and Differential Expression Analysis <a class=\"anchor\"  id=\"chapter8\"></a>[↑](#top)\n\nCel Multiplexing is a molecular-level labelling technique that helps organize and track cells for better organization and increased sample throughput in a single experiment. Differential Expression is a technique that estimates how gene expression changes in response to different compounds. There are 18211 genes in this dataset, this analysis helps researchers understand how compounds affect the cells at the genetic level.\n\nScientists use statistical techniques to analyze gene expression data. They look for patterns and changes in gene activity, determining which genes are more or less active in response to different treatments. The output of this differential expression analysis is a list of genes that a differentially expressed, along with information that they are upregulated (more active) or downregulated (less active) in response to the treatments.","metadata":{}},{"cell_type":"markdown","source":"\n### Data Splits <a class=\"anchor\"  id=\"chapter9\"></a>[↑](#top)\n\nIn this experiment, the data is divided into different sets for training and testing purposes.\n\n#### Train Data:\n\n**All compounds in T, NK cells:** This part of the data includes all the information about treatments (compounds) applied to T cells and natural killer (NK) cells. These are specific types of immune cells.\n\n**15 compounds + positive and negative controls in B and myeloid cells:** This part of the data includes information about 15 specific compounds applied to B cells and myeloid cells. B cells are another type of immune cell, and myeloid cells include cells like macrophages and monocytes. The \"positive\" and \"negative\" controls are substances with known effects that are included for reference.\n\n#### Public Test Data:\n\n**50 randomly selected compounds in B and myeloid cells:** This part of the data is used for testing the model's predictions. It includes information about 50 compounds applied to B cells and myeloid cells. The model will make predictions for this data, and those predictions will be compared to the actual results to evaluate how well the model performs.\n\n#### Private Test Data:\n\n**79 randomly selected compounds in B and myeloid cells:** It includes information about 79 compounds applied to B cells and myeloid cells. The model's predictions for this data will be used to assess its performance in a private evaluation.\n\n### Data Representation:\n\nEach row in the data represents a specific combination of a cell type and a treatment compound (SM_NAME). So, for example, a row could represent a B cell treated with a particular compound and the output of your model will be predicted  $signed -log10(p-values) $ for all 18211 genes.\n\nThe task of this competition is to predict the gene expression value of all 18,211 genes for given cell type and treatment pair.\n","metadata":{}},{"cell_type":"markdown","source":"## Data <a class=\"anchor\"  id=\"chapter8\"></a>[↑](#top)","metadata":{}},{"cell_type":"markdown","source":"## de_train","metadata":{}},{"cell_type":"code","source":"de_train_path = '/kaggle/input/open-problems-single-cell-perturbations/de_train.parquet'\nde_train = pd.read_parquet(de_train_path)\nprint(de_train.shape)\nde_train","metadata":{"execution":{"iopub.status.busy":"2023-09-20T13:19:56.835798Z","iopub.execute_input":"2023-09-20T13:19:56.836222Z","iopub.status.idle":"2023-09-20T13:19:58.270540Z","shell.execute_reply.started":"2023-09-20T13:19:56.836192Z","shell.execute_reply":"2023-09-20T13:19:58.269343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* We have 18216 columns, out of which 18211 are differential expression (DE) values for each\n* sm_name is the treatment/compound applied to the gene\n* cell_type is the type of cell\n* sm_lincs_id and SMILES is standard name representation of the compound\n* control refers to if this instance was a control, only positive control values are present, and the DE value is calculated by reference to negative control, as in what change happened compared to when nothing was done.\n","metadata":{}},{"cell_type":"markdown","source":"## adata_train","metadata":{}},{"cell_type":"code","source":"adata_train_path = '/kaggle/input/open-problems-single-cell-perturbations/adata_train.parquet'\nadata_train = pd.read_parquet(adata_train_path)\nprint(adata_train.shape)\nadata_train","metadata":{"execution":{"iopub.status.busy":"2023-09-20T13:23:12.017636Z","iopub.execute_input":"2023-09-20T13:23:12.018045Z","iopub.status.idle":"2023-09-20T13:24:27.190534Z","shell.execute_reply.started":"2023-09-20T13:23:12.018016Z","shell.execute_reply":"2023-09-20T13:24:27.189355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The unaggregated count and normalized data in COO sparse-array format is a supplemental dataset provided in the competition. It contains information about gene expression in individual cells.\n\nHow this data can be used:\n\n**Single-Cell Analysis:**  This data can be used for single-cell RNA sequencing analysis. It can be used to investigate the gene expression profiles of individual cells, which is crucial for understanding cellular diversity and heterogeneity within a tissue or sample.\n\n**Cell Clustering:**  We can cluster cells based on similar gene expression patterns to analyze gene expression profiles across cells. It can help identify distinct cell types or subpopulations within a larger cell population.\n\n**Differential Expression Analysis:** Differential expression analysis can be done to identify genes that are significantly upregulated or downregulated in specific cell types or under different experimental conditions. This is useful for understanding the molecular mechanisms underlying cellular responses.\n\n**Biological Insights:** The data can provide insights into the biology of cells. For example, it can help identify marker genes that are specific to certain cell types or conditions.\n\n**Integration with Other Data:** This data can be integrated with other datasets, such as the aggregated differential expression data, to gain a comprehensive understanding of cellular responses to perturbations.","metadata":{}},{"cell_type":"markdown","source":"## adata_obs_meta","metadata":{}},{"cell_type":"code","source":"\nadata_obs_meta_path = '/kaggle/input/open-problems-single-cell-perturbations/adata_obs_meta.csv'\nadata_obs_meta = pd.read_csv(adata_obs_meta_path)\nprint(adata_obs_meta.shape)\nadata_obs_meta","metadata":{"execution":{"iopub.status.busy":"2023-09-20T13:30:25.220440Z","iopub.execute_input":"2023-09-20T13:30:25.220836Z","iopub.status.idle":"2023-09-20T13:30:26.470540Z","shell.execute_reply.started":"2023-09-20T13:30:25.220807Z","shell.execute_reply":"2023-09-20T13:30:26.469423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The \"adata_obs_meta.csv\" file contains observation metadata that provides additional information about the samples used in the experiment. \n\nSome Ideas on how this data could be used.\n\n**Experimental Design Analysis:** Researchers can use this metadata to analyze how samples were organized on plates, the donor source of samples, the cell types involved, and the specific compounds and their concentrations used in the experiment. This information is crucial for understanding the experimental setup.\n\n**Treatment Duration Analysis:** The \"dose_uM\" and \"timepoint_hr\" columns provide information about the treatment conditions. Researchers can investigate how different doses and treatment durations affect gene expression.\n\n**Quality Control:** Metadata can be used for quality control checks to ensure that samples were processed correctly and that controls were appropriately included in the experiment.\n\n**Comparative Analysis:** Researchers may use this metadata to compare the effects of different compounds on specific cell types or under different conditions.\n\nOverall, this metadata complements the gene expression data and provides essential context for interpreting the results of the experiment and designing analytical approaches.","metadata":{}},{"cell_type":"markdown","source":"## multiome_train","metadata":{}},{"cell_type":"code","source":"\nmultiome_train_path = '/kaggle/input/open-problems-single-cell-perturbations/multiome_train.parquet'\nmultiome_train = pd.read_parquet(multiome_train_path)\nprint(multiome_train.shape)\nmultiome_train","metadata":{"execution":{"iopub.status.busy":"2023-09-21T07:08:26.048458Z","iopub.execute_input":"2023-09-21T07:08:26.048982Z","iopub.status.idle":"2023-09-21T07:09:40.914171Z","shell.execute_reply.started":"2023-09-21T07:08:26.048951Z","shell.execute_reply":"2023-09-21T07:09:40.913053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This dataset contains additional information about individual cells at their baseline state, using a technology known as \"10x Multiome.\" This data can be used to understand the cellular activity, to identify regulatory regions in a cell and to compare cells.","metadata":{}},{"cell_type":"markdown","source":"Other data files contains mappings for these datasets.","metadata":{}}]}