{
  "id": 448668,
  "title": "Overview and Data Explained",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/448668",
  "author_name": "Raki",
  "post_date": "2023-10-20T16:33:41.898000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'm not sure how useful this is, but I thought I just post my process of reading into a competition I don't know much about and how I make sense of it / how long it takes. Please note that I miss clicked in the beggining on data instead of overview making understanding harder than it should have been.</p>\n<h1>Open problems single cell perturbations</h1>\n<p><strong>Description:</strong> “For this competition, we designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). We selected 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells (data not released). We performed this experiment in human PBMCs because the cells are commercially available with pre-obtained consent for public release and PBMCs are a primary, disease-relevant tissue that contains multiple mature cell types (including T-cells, B-cells, myeloid cells, and NK cells) with established markers for annotation of cell types. To supplement this dataset, we also measured cells from each donor at baseline with joint scRNA and single-cell chromatin accessibility measurements using the 10x Multiome assay. We hope that the addition of rich multi-omic data for each donor and cell type at baseline will help establish biological priors that explain the susceptibility of particular genes to exhibit perturbation responses in difference biological contexts.”</p>\n<p>I understand almost nothing here and consulted Wikipedia and GPT4: </p>\n<p>Perturbational means disrupting usual state. PBMCs are from blood that cirulates outside heart and lungs, mononuclear is cell with single round nucleus. PBMCs are a mixture of cell types important for the immune system like lymphocytes and monocytes.</p>\n<p>In this context compounds are probably chemical substances used to treat cells to observe responses. LINCS is a project that aims at creating a network-based understanding of biology cataloguing changes in gene expression and other cellular processes that occur when cells are exposed to perturbations like the compounds. The PMID is a paper link. Gene expression is the process where information from a gene’s DNA is translated into a substance like proteins that cells use to perform functions. Single-cell stands in contrast to measuring gene expression in many cells at once. </p>\n<p>=&gt; Seems like this is a competition trying to map precise changes in single cells exposed to a range of substances. </p>\n<p>Next part is about why to measure this stuff on human PBMCs.</p>\n<p><em>Break after 15 min.</em></p>\n<p>Compounds based on diverse transcriptional signatures, meaning unique pattern or gene set for copying of DNA into RNA, which then is a template for protein proudction. CD34+ is a cell with a specific protein CD34 on the surface, which is a marker for identifying stem cells. Hematopoietic Stem Cells (HSCs) are found in bone marrow and can give rise to all other blood cells. Crucial for maintaining blood and particularly immune system. PBMCs are commercially easy to get and relevant for medical research, they also contain different mature cell types like T, B, myeloid,NK (immune response, producing antibodies, immune responeses, vs tumors and viral infections), they have established markers for annotation of cell type, which means they have molecules on cell surface with which you can categorize them. </p>\n<p>Additional measurements were taken: baseline as reference point for each individual, scRNA = single cell RNA sequencing for measuring single cell gene expression, the chromatin accesibility measures how tightly DNA is packed in the cell, affecting gene expression. The 10x multiome assay is a technology that allows both measurements from same cell (multi-omic data).</p>\n<p>They aim to use this data to help establish or confirm biological assumptions that can explain why certain genes in certain cells or individuals are more or less responsive to the treatments used in their experiment, across different biological contexts.</p>\n<p><strong>In summary:</strong></p>\n<ol>\n<li>Goal: The main objective was to understand how individual cells react when exposed to different substances (compounds).</li>\n<li>Selection of Compounds: They picked certain substances based on how these substances previously affected a specific type of stem cell. They wanted to see if similar or diverse reactions would happen in other types of blood cells.</li>\n<li>Using Human PBMCs: They chose to use a certain type of blood cell known as PBMCs because these cells are easy to obtain for research and are relevant to studying diseases. These cells also represent a mix of different mature blood cells which are central to our immune system.</li>\n<li>24-Hour Treatment: They exposed these blood cells to the chosen substances for 24 hours to observe any changes, specifically in the activity of genes within the cells.</li>\n<li>Baseline Measurements: Before exposing the cells to the substances, they measured the natural state of these cells using advanced techniques that look at gene activity and how accessible the DNA is within each cell. This gives them a reference point to understand the changes caused by the substances.</li>\n</ol>\n<h1>Starting from overview:</h1>\n<p>The goal of this competition is to predict how small molecules change gene expression in different cell types. This helps in drug discovery and basic biology.</p>\n<p><strong>Metric:</strong> <br>\nMean Rowwise Root Mean Squared Error = <br>\nmean_over_rows(root(mean_over_columns(error ** 2)))<br>\nFor each id (row) predict value for each of 18211 genes.</p>\n<p><strong>Prices:</strong> around 10k for top 5 based on score and 10k for top 5 based on judges</p>\n<p><em>Break after 25 min</em></p>\n<p>Judges: 6 categories, each 1-5, Integration of bio knowledge,exploration of problem, model design, robustness, documentation &amp; code style, reproducibility.<br>\nSet up ‘looking for team’ in the competition, I want to go for team gold to progress on master/grandmaster track this time, maybe also win money again.</p>\n<p><strong>Back to data:</strong><br>\nPBMCs plated on 96 well plates. Two columns for ‘positive controls’,  one for ‘negative controls’, meaning testing if substances known to affect gene activity and known to not affect gene activity had the expected effect. Wells contained mixed cell types. 72 compounds tried. 3 donors x 2 plates = 144 compounds x 3 donors. Compared measurements with control group/baseline.<br>\nIt may be that not all cell types are in every well (350 cells/well and some toxic interactions).</p>\n<p><em>Break after 15 min</em></p>\n<p>Task is prediction of differential expression (DE), for each of the 18211 genes. <br>\nData is ‘pseudobulked’ across row, plate, donor as technical covariate and compound as experimental covariate. Pseudobulked = summing of raw cell counts for each type and well.</p>\n<p>Row shouldn’t make an obvious difference here, the reason it is included is that there is some technical bias created from the 10x multiome assay ‘cell multiplexing’ across the rows. </p>\n<p>A limma function is created with f(g) = x0 + x1 * compound + x2 * row + x3 * donor + x4 * plate.<br>\nThis is done so that the influcence of donors having different base levels for example does not influence the outcome (at least the linear part of donor difference). </p>\n<p>CARE! This limma DE model is fitted to the train set for the training samples, the public and private test DE model is fit to all data. This is needed for privacy of donors.</p>\n<p>=&gt; So we have to predict x1? Also not directly but -log10(p – value)?</p>\n<p>Gets even more complicated:<br>\n‘The output of this model is an estimated fold-change in gene expression and a multiple-testing corrected p-value that a given gene's expression is dependent on the compound experimental variable. There is a long rabbit hole to go down in the world of differential expression testing. We don't have a complete mechanistic model of the data generative process of collecting scRNA data, and many groups disagree on the best way to account for nuisance variables or technical noise. We picked limma because it performs well in our testing.’.</p>\n<p>Fold change is simple multiple (change in gene expression).<br>\nMultiple Testing corrected p-value, is a measure of strength of evidence against the hypothesis that there is no perturbation from adding the compound to that gene expression. Multiple-testing makes it more robust. </p>\n<p><strong>Dataset:</strong> <br>\nAll 144 compounds for T and NK cells, 15 compounts + positive/negative controls in B and myeloid cells.</p>\n<p>Public test:  50 randomly selected compounds in B and myeloid cells</p>\n<p>Private test: 79 randomly selected compounds in B and meyeloid cells</p>\n<p>=====&gt; We have a full set of compounds for some cell types and only a few entries for other cell types and our task is to predict the likelihood in -log10(p-values) that a gene expression is influcenced by the addition of a compound for all 18211 genes.</p>\n<p><em>Break after 25 min</em></p>\n<p>Files: de_train: <br>\ncell type, compound name and id (sm), SMILES simplefied molecular representation, genes, control<br>\nAdditional metadata (raw counts for example): adata_train, adata_obs_meta, multiome stuff<br>\n id_map and sample submission<br>\n=====&gt; You probably don’t need the metadata, or maybe you need it at some point to squeeze out the last tiny percentile score points, but in the beginning it’s enough to only download de_train + sample submission (only 100 MB instead of 5 GB).</p>\n<p><em>Total reading time ~90 mins, maybe a bit less than 2 hours in total.</em></p>",
  "messages": [
    {
      "id": 2490403,
      "postDate": "2023-10-20T16:33:41.897Z",
      "content": "<p>I'm not sure how useful this is, but I thought I just post my process of reading into a competition I don't know much about and how I make sense of it / how long it takes. Please note that I miss clicked in the beggining on data instead of overview making understanding harder than it should have been.</p>\n<h1>Open problems single cell perturbations</h1>\n<p><strong>Description:</strong> “For this competition, we designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). We selected 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells (data not released). We performed this experiment in human PBMCs because the cells are commercially available with pre-obtained consent for public release and PBMCs are a primary, disease-relevant tissue that contains multiple mature cell types (including T-cells, B-cells, myeloid cells, and NK cells) with established markers for annotation of cell types. To supplement this dataset, we also measured cells from each donor at baseline with joint scRNA and single-cell chromatin accessibility measurements using the 10x Multiome assay. We hope that the addition of rich multi-omic data for each donor and cell type at baseline will help establish biological priors that explain the susceptibility of particular genes to exhibit perturbation responses in difference biological contexts.”</p>\n<p>I understand almost nothing here and consulted Wikipedia and GPT4: </p>\n<p>Perturbational means disrupting usual state. PBMCs are from blood that cirulates outside heart and lungs, mononuclear is cell with single round nucleus. PBMCs are a mixture of cell types important for the immune system like lymphocytes and monocytes.</p>\n<p>In this context compounds are probably chemical substances used to treat cells to observe responses. LINCS is a project that aims at creating a network-based understanding of biology cataloguing changes in gene expression and other cellular processes that occur when cells are exposed to perturbations like the compounds. The PMID is a paper link. Gene expression is the process where information from a gene’s DNA is translated into a substance like proteins that cells use to perform functions. Single-cell stands in contrast to measuring gene expression in many cells at once. </p>\n<p>=&gt; Seems like this is a competition trying to map precise changes in single cells exposed to a range of substances. </p>\n<p>Next part is about why to measure this stuff on human PBMCs.</p>\n<p><em>Break after 15 min.</em></p>\n<p>Compounds based on diverse transcriptional signatures, meaning unique pattern or gene set for copying of DNA into RNA, which then is a template for protein proudction. CD34+ is a cell with a specific protein CD34 on the surface, which is a marker for identifying stem cells. Hematopoietic Stem Cells (HSCs) are found in bone marrow and can give rise to all other blood cells. Crucial for maintaining blood and particularly immune system. PBMCs are commercially easy to get and relevant for medical research, they also contain different mature cell types like T, B, myeloid,NK (immune response, producing antibodies, immune responeses, vs tumors and viral infections), they have established markers for annotation of cell type, which means they have molecules on cell surface with which you can categorize them. </p>\n<p>Additional measurements were taken: baseline as reference point for each individual, scRNA = single cell RNA sequencing for measuring single cell gene expression, the chromatin accesibility measures how tightly DNA is packed in the cell, affecting gene expression. The 10x multiome assay is a technology that allows both measurements from same cell (multi-omic data).</p>\n<p>They aim to use this data to help establish or confirm biological assumptions that can explain why certain genes in certain cells or individuals are more or less responsive to the treatments used in their experiment, across different biological contexts.</p>\n<p><strong>In summary:</strong></p>\n<ol>\n<li>Goal: The main objective was to understand how individual cells react when exposed to different substances (compounds).</li>\n<li>Selection of Compounds: They picked certain substances based on how these substances previously affected a specific type of stem cell. They wanted to see if similar or diverse reactions would happen in other types of blood cells.</li>\n<li>Using Human PBMCs: They chose to use a certain type of blood cell known as PBMCs because these cells are easy to obtain for research and are relevant to studying diseases. These cells also represent a mix of different mature blood cells which are central to our immune system.</li>\n<li>24-Hour Treatment: They exposed these blood cells to the chosen substances for 24 hours to observe any changes, specifically in the activity of genes within the cells.</li>\n<li>Baseline Measurements: Before exposing the cells to the substances, they measured the natural state of these cells using advanced techniques that look at gene activity and how accessible the DNA is within each cell. This gives them a reference point to understand the changes caused by the substances.</li>\n</ol>\n<h1>Starting from overview:</h1>\n<p>The goal of this competition is to predict how small molecules change gene expression in different cell types. This helps in drug discovery and basic biology.</p>\n<p><strong>Metric:</strong> <br>\nMean Rowwise Root Mean Squared Error = <br>\nmean_over_rows(root(mean_over_columns(error ** 2)))<br>\nFor each id (row) predict value for each of 18211 genes.</p>\n<p><strong>Prices:</strong> around 10k for top 5 based on score and 10k for top 5 based on judges</p>\n<p><em>Break after 25 min</em></p>\n<p>Judges: 6 categories, each 1-5, Integration of bio knowledge,exploration of problem, model design, robustness, documentation &amp; code style, reproducibility.<br>\nSet up ‘looking for team’ in the competition, I want to go for team gold to progress on master/grandmaster track this time, maybe also win money again.</p>\n<p><strong>Back to data:</strong><br>\nPBMCs plated on 96 well plates. Two columns for ‘positive controls’,  one for ‘negative controls’, meaning testing if substances known to affect gene activity and known to not affect gene activity had the expected effect. Wells contained mixed cell types. 72 compounds tried. 3 donors x 2 plates = 144 compounds x 3 donors. Compared measurements with control group/baseline.<br>\nIt may be that not all cell types are in every well (350 cells/well and some toxic interactions).</p>\n<p><em>Break after 15 min</em></p>\n<p>Task is prediction of differential expression (DE), for each of the 18211 genes. <br>\nData is ‘pseudobulked’ across row, plate, donor as technical covariate and compound as experimental covariate. Pseudobulked = summing of raw cell counts for each type and well.</p>\n<p>Row shouldn’t make an obvious difference here, the reason it is included is that there is some technical bias created from the 10x multiome assay ‘cell multiplexing’ across the rows. </p>\n<p>A limma function is created with f(g) = x0 + x1 * compound + x2 * row + x3 * donor + x4 * plate.<br>\nThis is done so that the influcence of donors having different base levels for example does not influence the outcome (at least the linear part of donor difference). </p>\n<p>CARE! This limma DE model is fitted to the train set for the training samples, the public and private test DE model is fit to all data. This is needed for privacy of donors.</p>\n<p>=&gt; So we have to predict x1? Also not directly but -log10(p – value)?</p>\n<p>Gets even more complicated:<br>\n‘The output of this model is an estimated fold-change in gene expression and a multiple-testing corrected p-value that a given gene's expression is dependent on the compound experimental variable. There is a long rabbit hole to go down in the world of differential expression testing. We don't have a complete mechanistic model of the data generative process of collecting scRNA data, and many groups disagree on the best way to account for nuisance variables or technical noise. We picked limma because it performs well in our testing.’.</p>\n<p>Fold change is simple multiple (change in gene expression).<br>\nMultiple Testing corrected p-value, is a measure of strength of evidence against the hypothesis that there is no perturbation from adding the compound to that gene expression. Multiple-testing makes it more robust. </p>\n<p><strong>Dataset:</strong> <br>\nAll 144 compounds for T and NK cells, 15 compounts + positive/negative controls in B and myeloid cells.</p>\n<p>Public test:  50 randomly selected compounds in B and myeloid cells</p>\n<p>Private test: 79 randomly selected compounds in B and meyeloid cells</p>\n<p>=====&gt; We have a full set of compounds for some cell types and only a few entries for other cell types and our task is to predict the likelihood in -log10(p-values) that a gene expression is influcenced by the addition of a compound for all 18211 genes.</p>\n<p><em>Break after 25 min</em></p>\n<p>Files: de_train: <br>\ncell type, compound name and id (sm), SMILES simplefied molecular representation, genes, control<br>\nAdditional metadata (raw counts for example): adata_train, adata_obs_meta, multiome stuff<br>\n id_map and sample submission<br>\n=====&gt; You probably don’t need the metadata, or maybe you need it at some point to squeeze out the last tiny percentile score points, but in the beginning it’s enough to only download de_train + sample submission (only 100 MB instead of 5 GB).</p>\n<p><em>Total reading time ~90 mins, maybe a bit less than 2 hours in total.</em></p>",
      "rawMarkdown": "I'm not sure how useful this is, but I thought I just post my process of reading into a competition I don't know much about and how I make sense of it / how long it takes. Please note that I miss clicked in the beggining on data instead of overview making understanding harder than it should have been.\n\n# Open problems single cell perturbations\n\n**Description:** “For this competition, we designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). We selected 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells (data not released). We performed this experiment in human PBMCs because the cells are commercially available with pre-obtained consent for public release and PBMCs are a primary, disease-relevant tissue that contains multiple mature cell types (including T-cells, B-cells, myeloid cells, and NK cells) with established markers for annotation of cell types. To supplement this dataset, we also measured cells from each donor at baseline with joint scRNA and single-cell chromatin accessibility measurements using the 10x Multiome assay. We hope that the addition of rich multi-omic data for each donor and cell type at baseline will help establish biological priors that explain the susceptibility of particular genes to exhibit perturbation responses in difference biological contexts.”\n\nI understand almost nothing here and consulted Wikipedia and GPT4: \n\nPerturbational means disrupting usual state. PBMCs are from blood that cirulates outside heart and lungs, mononuclear is cell with single round nucleus. PBMCs are a mixture of cell types important for the immune system like lymphocytes and monocytes.\n\nIn this context compounds are probably chemical substances used to treat cells to observe responses. LINCS is a project that aims at creating a network-based understanding of biology cataloguing changes in gene expression and other cellular processes that occur when cells are exposed to perturbations like the compounds. The PMID is a paper link. Gene expression is the process where information from a gene’s DNA is translated into a substance like proteins that cells use to perform functions. Single-cell stands in contrast to measuring gene expression in many cells at once. \n\n=> Seems like this is a competition trying to map precise changes in single cells exposed to a range of substances. \n\nNext part is about why to measure this stuff on human PBMCs.\n\n*Break after 15 min.*\n\nCompounds based on diverse transcriptional signatures, meaning unique pattern or gene set for copying of DNA into RNA, which then is a template for protein proudction. CD34+ is a cell with a specific protein CD34 on the surface, which is a marker for identifying stem cells. Hematopoietic Stem Cells (HSCs) are found in bone marrow and can give rise to all other blood cells. Crucial for maintaining blood and particularly immune system. PBMCs are commercially easy to get and relevant for medical research, they also contain different mature cell types like T, B, myeloid,NK (immune response, producing antibodies, immune responeses, vs tumors and viral infections), they have established markers for annotation of cell type, which means they have molecules on cell surface with which you can categorize them. \n\nAdditional measurements were taken: baseline as reference point for each individual, scRNA = single cell RNA sequencing for measuring single cell gene expression, the chromatin accesibility measures how tightly DNA is packed in the cell, affecting gene expression. The 10x multiome assay is a technology that allows both measurements from same cell (multi-omic data).\n\nThey aim to use this data to help establish or confirm biological assumptions that can explain why certain genes in certain cells or individuals are more or less responsive to the treatments used in their experiment, across different biological contexts.\n\n**In summary:**\n\n1. Goal: The main objective was to understand how individual cells react when exposed to different substances (compounds).\n2. Selection of Compounds: They picked certain substances based on how these substances previously affected a specific type of stem cell. They wanted to see if similar or diverse reactions would happen in other types of blood cells.\n3. Using Human PBMCs: They chose to use a certain type of blood cell known as PBMCs because these cells are easy to obtain for research and are relevant to studying diseases. These cells also represent a mix of different mature blood cells which are central to our immune system.\n4. 24-Hour Treatment: They exposed these blood cells to the chosen substances for 24 hours to observe any changes, specifically in the activity of genes within the cells.\n5. Baseline Measurements: Before exposing the cells to the substances, they measured the natural state of these cells using advanced techniques that look at gene activity and how accessible the DNA is within each cell. This gives them a reference point to understand the changes caused by the substances.\n          \n#Starting from overview:\nThe goal of this competition is to predict how small molecules change gene expression in different cell types. This helps in drug discovery and basic biology.\n\n**Metric:** \nMean Rowwise Root Mean Squared Error = \nmean_over_rows(root(mean_over_columns(error ** 2)))\nFor each id (row) predict value for each of 18211 genes.\n\n**Prices:** around 10k for top 5 based on score and 10k for top 5 based on judges\n\n*Break after 25 min*\n\nJudges: 6 categories, each 1-5, Integration of bio knowledge,exploration of problem, model design, robustness, documentation & code style, reproducibility.\nSet up ‘looking for team’ in the competition, I want to go for team gold to progress on master/grandmaster track this time, maybe also win money again.\n\n**Back to data:**\nPBMCs plated on 96 well plates. Two columns for ‘positive controls’,  one for ‘negative controls’, meaning testing if substances known to affect gene activity and known to not affect gene activity had the expected effect. Wells contained mixed cell types. 72 compounds tried. 3 donors x 2 plates = 144 compounds x 3 donors. Compared measurements with control group/baseline.\nIt may be that not all cell types are in every well (350 cells/well and some toxic interactions).\n\n*Break after 15 min*\n\nTask is prediction of differential expression (DE), for each of the 18211 genes. \nData is ‘pseudobulked’ across row, plate, donor as technical covariate and compound as experimental covariate. Pseudobulked = summing of raw cell counts for each type and well.\n\nRow shouldn’t make an obvious difference here, the reason it is included is that there is some technical bias created from the 10x multiome assay ‘cell multiplexing’ across the rows. \n\nA limma function is created with f(g) = x0 + x1 * compound + x2 * row + x3 * donor + x4 * plate.\nThis is done so that the influcence of donors having different base levels for example does not influence the outcome (at least the linear part of donor difference). \n\nCARE! This limma DE model is fitted to the train set for the training samples, the public and private test DE model is fit to all data. This is needed for privacy of donors.\n\n=> So we have to predict x1? Also not directly but -log10(p – value)?\n\nGets even more complicated:\n‘The output of this model is an estimated fold-change in gene expression and a multiple-testing corrected p-value that a given gene's expression is dependent on the compound experimental variable. There is a long rabbit hole to go down in the world of differential expression testing. We don't have a complete mechanistic model of the data generative process of collecting scRNA data, and many groups disagree on the best way to account for nuisance variables or technical noise. We picked limma because it performs well in our testing.’.\n\nFold change is simple multiple (change in gene expression).\nMultiple Testing corrected p-value, is a measure of strength of evidence against the hypothesis that there is no perturbation from adding the compound to that gene expression. Multiple-testing makes it more robust. \n\n**Dataset:** \nAll 144 compounds for T and NK cells, 15 compounts + positive/negative controls in B and myeloid cells.\n\nPublic test:  50 randomly selected compounds in B and myeloid cells\n\nPrivate test: 79 randomly selected compounds in B and meyeloid cells\n\n=====> We have a full set of compounds for some cell types and only a few entries for other cell types and our task is to predict the likelihood in -log10(p-values) that a gene expression is influcenced by the addition of a compound for all 18211 genes.\n\n*Break after 25 min*\n\nFiles: de_train: \ncell type, compound name and id (sm), SMILES simplefied molecular representation, genes, control\nAdditional metadata (raw counts for example): adata_train, adata_obs_meta, multiome stuff\n id_map and sample submission\n=====> You probably don’t need the metadata, or maybe you need it at some point to squeeze out the last tiny percentile score points, but in the beginning it’s enough to only download de_train + sample submission (only 100 MB instead of 5 GB).\n\n*Total reading time ~90 mins, maybe a bit less than 2 hours in total.*",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2490403": "I'm not sure how useful this is, but I thought I just post my process of reading into a competition I don't know much about and how I make sense of it / how long it takes. Please note that I miss clicked in the beggining on data instead of overview making understanding harder than it should have been.\n\n# Open problems single cell perturbations\n\n**Description:** “For this competition, we designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). We selected 144 compounds from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells (data not released). We performed this experiment in human PBMCs because the cells are commercially available with pre-obtained consent for public release and PBMCs are a primary, disease-relevant tissue that contains multiple mature cell types (including T-cells, B-cells, myeloid cells, and NK cells) with established markers for annotation of cell types. To supplement this dataset, we also measured cells from each donor at baseline with joint scRNA and single-cell chromatin accessibility measurements using the 10x Multiome assay. We hope that the addition of rich multi-omic data for each donor and cell type at baseline will help establish biological priors that explain the susceptibility of particular genes to exhibit perturbation responses in difference biological contexts.”\n\nI understand almost nothing here and consulted Wikipedia and GPT4: \n\nPerturbational means disrupting usual state. PBMCs are from blood that cirulates outside heart and lungs, mononuclear is cell with single round nucleus. PBMCs are a mixture of cell types important for the immune system like lymphocytes and monocytes.\n\nIn this context compounds are probably chemical substances used to treat cells to observe responses. LINCS is a project that aims at creating a network-based understanding of biology cataloguing changes in gene expression and other cellular processes that occur when cells are exposed to perturbations like the compounds. The PMID is a paper link. Gene expression is the process where information from a gene’s DNA is translated into a substance like proteins that cells use to perform functions. Single-cell stands in contrast to measuring gene expression in many cells at once. \n\n=> Seems like this is a competition trying to map precise changes in single cells exposed to a range of substances. \n\nNext part is about why to measure this stuff on human PBMCs.\n\n*Break after 15 min.*\n\nCompounds based on diverse transcriptional signatures, meaning unique pattern or gene set for copying of DNA into RNA, which then is a template for protein proudction. CD34+ is a cell with a specific protein CD34 on the surface, which is a marker for identifying stem cells. Hematopoietic Stem Cells (HSCs) are found in bone marrow and can give rise to all other blood cells. Crucial for maintaining blood and particularly immune system. PBMCs are commercially easy to get and relevant for medical research, they also contain different mature cell types like T, B, myeloid,NK (immune response, producing antibodies, immune responeses, vs tumors and viral infections), they have established markers for annotation of cell type, which means they have molecules on cell surface with which you can categorize them. \n\nAdditional measurements were taken: baseline as reference point for each individual, scRNA = single cell RNA sequencing for measuring single cell gene expression, the chromatin accesibility measures how tightly DNA is packed in the cell, affecting gene expression. The 10x multiome assay is a technology that allows both measurements from same cell (multi-omic data).\n\nThey aim to use this data to help establish or confirm biological assumptions that can explain why certain genes in certain cells or individuals are more or less responsive to the treatments used in their experiment, across different biological contexts.\n\n**In summary:**\n\n1. Goal: The main objective was to understand how individual cells react when exposed to different substances (compounds).\n2. Selection of Compounds: They picked certain substances based on how these substances previously affected a specific type of stem cell. They wanted to see if similar or diverse reactions would happen in other types of blood cells.\n3. Using Human PBMCs: They chose to use a certain type of blood cell known as PBMCs because these cells are easy to obtain for research and are relevant to studying diseases. These cells also represent a mix of different mature blood cells which are central to our immune system.\n4. 24-Hour Treatment: They exposed these blood cells to the chosen substances for 24 hours to observe any changes, specifically in the activity of genes within the cells.\n5. Baseline Measurements: Before exposing the cells to the substances, they measured the natural state of these cells using advanced techniques that look at gene activity and how accessible the DNA is within each cell. This gives them a reference point to understand the changes caused by the substances.\n          \n#Starting from overview:\nThe goal of this competition is to predict how small molecules change gene expression in different cell types. This helps in drug discovery and basic biology.\n\n**Metric:** \nMean Rowwise Root Mean Squared Error = \nmean_over_rows(root(mean_over_columns(error ** 2)))\nFor each id (row) predict value for each of 18211 genes.\n\n**Prices:** around 10k for top 5 based on score and 10k for top 5 based on judges\n\n*Break after 25 min*\n\nJudges: 6 categories, each 1-5, Integration of bio knowledge,exploration of problem, model design, robustness, documentation & code style, reproducibility.\nSet up ‘looking for team’ in the competition, I want to go for team gold to progress on master/grandmaster track this time, maybe also win money again.\n\n**Back to data:**\nPBMCs plated on 96 well plates. Two columns for ‘positive controls’,  one for ‘negative controls’, meaning testing if substances known to affect gene activity and known to not affect gene activity had the expected effect. Wells contained mixed cell types. 72 compounds tried. 3 donors x 2 plates = 144 compounds x 3 donors. Compared measurements with control group/baseline.\nIt may be that not all cell types are in every well (350 cells/well and some toxic interactions).\n\n*Break after 15 min*\n\nTask is prediction of differential expression (DE), for each of the 18211 genes. \nData is ‘pseudobulked’ across row, plate, donor as technical covariate and compound as experimental covariate. Pseudobulked = summing of raw cell counts for each type and well.\n\nRow shouldn’t make an obvious difference here, the reason it is included is that there is some technical bias created from the 10x multiome assay ‘cell multiplexing’ across the rows. \n\nA limma function is created with f(g) = x0 + x1 * compound + x2 * row + x3 * donor + x4 * plate.\nThis is done so that the influcence of donors having different base levels for example does not influence the outcome (at least the linear part of donor difference). \n\nCARE! This limma DE model is fitted to the train set for the training samples, the public and private test DE model is fit to all data. This is needed for privacy of donors.\n\n=> So we have to predict x1? Also not directly but -log10(p – value)?\n\nGets even more complicated:\n‘The output of this model is an estimated fold-change in gene expression and a multiple-testing corrected p-value that a given gene's expression is dependent on the compound experimental variable. There is a long rabbit hole to go down in the world of differential expression testing. We don't have a complete mechanistic model of the data generative process of collecting scRNA data, and many groups disagree on the best way to account for nuisance variables or technical noise. We picked limma because it performs well in our testing.’.\n\nFold change is simple multiple (change in gene expression).\nMultiple Testing corrected p-value, is a measure of strength of evidence against the hypothesis that there is no perturbation from adding the compound to that gene expression. Multiple-testing makes it more robust. \n\n**Dataset:** \nAll 144 compounds for T and NK cells, 15 compounts + positive/negative controls in B and myeloid cells.\n\nPublic test:  50 randomly selected compounds in B and myeloid cells\n\nPrivate test: 79 randomly selected compounds in B and meyeloid cells\n\n=====> We have a full set of compounds for some cell types and only a few entries for other cell types and our task is to predict the likelihood in -log10(p-values) that a gene expression is influcenced by the addition of a compound for all 18211 genes.\n\n*Break after 25 min*\n\nFiles: de_train: \ncell type, compound name and id (sm), SMILES simplefied molecular representation, genes, control\nAdditional metadata (raw counts for example): adata_train, adata_obs_meta, multiome stuff\n id_map and sample submission\n=====> You probably don’t need the metadata, or maybe you need it at some point to squeeze out the last tiny percentile score points, but in the beginning it’s enough to only download de_train + sample submission (only 100 MB instead of 5 GB).\n\n*Total reading time ~90 mins, maybe a bit less than 2 hours in total.*"
  }
}