{
  "id": 461166,
  "title": "178th Place Solution for the Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/insiya-jafferji-178th-place-solution-for-the-open-",
  "author_name": "",
  "post_date": "2023-12-23T01:33:08.827Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Many thanks to the organizers for hosting the Open Problems – Single-Cell Perturbations competition. Congratulations to the winners and everyone who participated, I really learnt alot from the discussions and great notebooks!</p>\n<p>I thought particularly great contributions and insights were from the following:</p>\n<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a><br>\n<a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a><br>\n<a href=\"https://www.kaggle.com/mehrankazeminia\" target=\"_blank\">@mehrankazeminia</a><br>\n<a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a></p>\n<p>Context:</p>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a><br>\nIntegration of Biological Knowledge:</p>\n<p>Even though I was not at the top of the leader-board I wanted to share my experience in this competition and my approach! </p>\n<p>One of the objectives of this competition is to determine how compounds are involved in changes in gene expression within cells. For this competition, the organizers designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). 144 compounds were selected from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells.</p>\n<p>Before I started to use ML models, I was interested in exploring the dataset within this competition to determine the compounds involved and the distribution of cell types associated with them as outlined in the following notebook:</p>\n<p><a href=\"https://www.kaggle.com/code/insiyajafferji/op-single-cell-dataset-and-cell-type-distribution\" target=\"_blank\">https://www.kaggle.com/code/insiyajafferji/op-single-cell-dataset-and-cell-type-distribution</a></p>\n<p>In this notebook some initial dataset exploration helped me to understand cell type differences within PBMC samples treated with the compounds. I found that NK, CD4+ T cells, CD8+ T cells and Tregs are in all PBMC samples of the compounds being treated and there are a subset of PBMCs where compounds treated also have B and myeloid cells. PBMCs have been reported to contain a multitude of distinct multipotent progenitor cell populations and therefore treatment of certain drug compounds could have the ability to affect cell type distribution aswell as gene expression which can have an impact on the DE/DGE analysis.</p>\n<p>Interestingly some notebooks have suggested that the highest bias (differences between predicted and true values) and variability of gene DE predictions are related to individual drugs rather than cell types. <br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions</a></p>\n<p>Based on this information, the specific embeddings that I thought would help my ML models would be based on cell type, compound (sm_name and SMILES), and DE of genes (gene name).</p>\n<p>Exploration of the problem:</p>\n<p>When I was exploring the ways to approach the challenges associated with this competition I found the following paper useful:<br>\n<a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3</a></p>\n<p>Here, feature selection an essential technique for single-cell data analysis in high-dimensional datasets.<br>\nImportantly, feature selection is an effective strategy to reduce the feature dimension and redundancy and can alleviate issues such as model overfitting in downstream analysis. Different from dimension reduction methods (e.g. principal component analysis) where features in a dataset are combined and/or transformed to derive a lower feature dimension, feature selection methods do not alter the original features in the dataset but only identify and select features that satisfy certain pre-defined criteria or optimise certain computational procedures. Some of the most popular research directions include selecting genes that can discriminate certain cell types. </p>\n<p>In this competition setup, participants were tasked with modelling differential expression (DE), which enables us to estimate the impact of an experimental perturbation on the expression level of every gene in the transcription (18211 genes in this dataset). The Limma model was used to determine ‘differential expression’ (DE) methods for biological data analysis in this dataset.</p>\n<p>The cell type proportion differs in PBMCs treated with different compounds and would therefore have different differential gene expression. For example T cells have a very distinct gene expression profile (i.e express CD3E) compared to B cells (CD79a, MSA41) and Myeloid cells (CD14, FCGR3A(CD16)). Also T cell sub-types also express differences, for example Regulatory T cells are classified by the expression of FOXP3, IL2RA and CTLA4 and Cytotoxic CD8 T cells are enriched in cytotoxic-related genes including GNLY, CCL5, NKG7, GZMH, LYZ, GZMB, and GZMK. CD4 T cells will express the CD4 gene whereas CD8 T cells will express CD8A and CD8B genes. Typically in single cell analysis the result is displayed as TSNE or UMAP plot showing distinct clusters with ideally one cell type and marker plots and DGE analysis can be used to confirm cell types  such as in the following paper (<a href=\"https://pubmed.ncbi.nlm.nih.gov/34911770/\" target=\"_blank\">https://pubmed.ncbi.nlm.nih.gov/34911770/</a>)</p>\n<p>The following paper shows the combined protein and transcript analysis of single-cell RNA sequencing in human peripheral blood mononuclear cells (<a href=\"https://bmcbiol.biomedcentral.com/articles/10.1186/s12915-022-01382-4\" target=\"_blank\">https://bmcbiol.biomedcentral.com/articles/10.1186/s12915-022-01382-4</a>)</p>\n<p>Supervised ensemble classification models are popular among bioinformatics applications and have recently seen their increasing integration with deep learning models. Ensemble feature selection methods, typically, rely on either perturbation to the dataset or hyperparameters of the feature selection algorithms for creating ‘base selectors’ from which the ensemble could be derived. Generally, hybrid methods are motivated by the aim of taking advantage of the strengths of individual methods while alleviating or avoiding their weaknesses.</p>\n<p>Model design:</p>\n<p>I tried a number of approaches for the ML model and found that ensembeling models gave the best approach as described in section 2 of exploration of the problem. I used a combination of blends and ridge models.  Selection of certain cell types such as B cells, T regulatory cells and NK cells did help improve the model and including CD8 T cells gave a worse score.</p>\n<p>I used the features \"cell_type\" and \"sm_name\" and used certain a 'cell_type' with a selection of compounds that will have similar responses on each of these cell divisions. </p>\n<p>RMSE was calculated for each line and then to select the best model for each compound there were two submissions for every line that has RMSE &gt;1</p>\n<p>I experimented with a number of blends and I would like to acknowledge the following notebooks from the following participants:<br>\n-Daphne Anga <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457081\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457081</a><br>\n-Mehran Kazeminia, Somayyeh Gholami <a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a></p>\n<p>For including the Ridge model thanks to AMBROSM and MT for the following:<br>\n<a href=\"https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook</a></p>\n<p>Robustness:<br>\nWhen using the cross-validation it did help in improving the robustness of the model in this competition and could help with over fitting.  I included the cross-validation strategy that makes four folds.</p>\n<p>Documentation, Code and Reproducibility:<br>\nPlease see the Kaggle link to my approach below.<br>\n<a href=\"https://www.kaggle.com/code/insiyajafferji/code-178th-priv-pub-0-759-0-531\" target=\"_blank\">https://www.kaggle.com/code/insiyajafferji/code-178th-priv-pub-0-759-0-531</a></p>",
  "messages": [
    {
      "id": "2559545",
      "postDate": "12/12/2023 23:05:04",
      "content": "<p>Many thanks to the organizers for hosting the Open Problems – Single-Cell Perturbations competition. Congratulations to the winners and everyone who participated, I really learnt alot from the discussions and great notebooks!</p>\n<p>I thought particularly great contributions and insights were from the following:</p>\n<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a><br>\n<a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a><br>\n<a href=\"https://www.kaggle.com/mehrankazeminia\" target=\"_blank\">@mehrankazeminia</a><br>\n<a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a></p>\n<p>Context:</p>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a><br>\nIntegration of Biological Knowledge:</p>\n<p>Even though I was not at the top of the leader-board I wanted to share my experience in this competition and my approach! </p>\n<p>One of the objectives of this competition is to determine how compounds are involved in changes in gene expression within cells. For this competition, the organizers designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). 144 compounds were selected from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells.</p>\n<p>Before I started to use ML models, I was interested in exploring the dataset within this competition to determine the compounds involved and the distribution of cell types associated with them as outlined in the following notebook:</p>\n<p><a href=\"https://www.kaggle.com/code/insiyajafferji/op-single-cell-dataset-and-cell-type-distribution\" target=\"_blank\">https://www.kaggle.com/code/insiyajafferji/op-single-cell-dataset-and-cell-type-distribution</a></p>\n<p>In this notebook some initial dataset exploration helped me to understand cell type differences within PBMC samples treated with the compounds. I found that NK, CD4+ T cells, CD8+ T cells and Tregs are in all PBMC samples of the compounds being treated and there are a subset of PBMCs where compounds treated also have B and myeloid cells. PBMCs have been reported to contain a multitude of distinct multipotent progenitor cell populations and therefore treatment of certain drug compounds could have the ability to affect cell type distribution aswell as gene expression which can have an impact on the DE/DGE analysis.</p>\n<p>Interestingly some notebooks have suggested that the highest bias (differences between predicted and true values) and variability of gene DE predictions are related to individual drugs rather than cell types. <br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions</a></p>\n<p>Based on this information, the specific embeddings that I thought would help my ML models would be based on cell type, compound (sm_name and SMILES), and DE of genes (gene name).</p>\n<p>Exploration of the problem:</p>\n<p>When I was exploring the ways to approach the challenges associated with this competition I found the following paper useful:<br>\n<a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3</a></p>\n<p>Here, feature selection an essential technique for single-cell data analysis in high-dimensional datasets.<br>\nImportantly, feature selection is an effective strategy to reduce the feature dimension and redundancy and can alleviate issues such as model overfitting in downstream analysis. Different from dimension reduction methods (e.g. principal component analysis) where features in a dataset are combined and/or transformed to derive a lower feature dimension, feature selection methods do not alter the original features in the dataset but only identify and select features that satisfy certain pre-defined criteria or optimise certain computational procedures. Some of the most popular research directions include selecting genes that can discriminate certain cell types. </p>\n<p>In this competition setup, participants were tasked with modelling differential expression (DE), which enables us to estimate the impact of an experimental perturbation on the expression level of every gene in the transcription (18211 genes in this dataset). The Limma model was used to determine ‘differential expression’ (DE) methods for biological data analysis in this dataset.</p>\n<p>The cell type proportion differs in PBMCs treated with different compounds and would therefore have different differential gene expression. For example T cells have a very distinct gene expression profile (i.e express CD3E) compared to B cells (CD79a, MSA41) and Myeloid cells (CD14, FCGR3A(CD16)). Also T cell sub-types also express differences, for example Regulatory T cells are classified by the expression of FOXP3, IL2RA and CTLA4 and Cytotoxic CD8 T cells are enriched in cytotoxic-related genes including GNLY, CCL5, NKG7, GZMH, LYZ, GZMB, and GZMK. CD4 T cells will express the CD4 gene whereas CD8 T cells will express CD8A and CD8B genes. Typically in single cell analysis the result is displayed as TSNE or UMAP plot showing distinct clusters with ideally one cell type and marker plots and DGE analysis can be used to confirm cell types  such as in the following paper (<a href=\"https://pubmed.ncbi.nlm.nih.gov/34911770/\" target=\"_blank\">https://pubmed.ncbi.nlm.nih.gov/34911770/</a>)</p>\n<p>The following paper shows the combined protein and transcript analysis of single-cell RNA sequencing in human peripheral blood mononuclear cells (<a href=\"https://bmcbiol.biomedcentral.com/articles/10.1186/s12915-022-01382-4\" target=\"_blank\">https://bmcbiol.biomedcentral.com/articles/10.1186/s12915-022-01382-4</a>)</p>\n<p>Supervised ensemble classification models are popular among bioinformatics applications and have recently seen their increasing integration with deep learning models. Ensemble feature selection methods, typically, rely on either perturbation to the dataset or hyperparameters of the feature selection algorithms for creating ‘base selectors’ from which the ensemble could be derived. Generally, hybrid methods are motivated by the aim of taking advantage of the strengths of individual methods while alleviating or avoiding their weaknesses.</p>\n<p>Model design:</p>\n<p>I tried a number of approaches for the ML model and found that ensembeling models gave the best approach as described in section 2 of exploration of the problem. I used a combination of blends and ridge models.  Selection of certain cell types such as B cells, T regulatory cells and NK cells did help improve the model and including CD8 T cells gave a worse score.</p>\n<p>I used the features \"cell_type\" and \"sm_name\" and used certain a 'cell_type' with a selection of compounds that will have similar responses on each of these cell divisions. </p>\n<p>RMSE was calculated for each line and then to select the best model for each compound there were two submissions for every line that has RMSE &gt;1</p>\n<p>I experimented with a number of blends and I would like to acknowledge the following notebooks from the following participants:<br>\n-Daphne Anga <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457081\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457081</a><br>\n-Mehran Kazeminia, Somayyeh Gholami <a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a></p>\n<p>For including the Ridge model thanks to AMBROSM and MT for the following:<br>\n<a href=\"https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook</a></p>\n<p>Robustness:<br>\nWhen using the cross-validation it did help in improving the robustness of the model in this competition and could help with over fitting.  I included the cross-validation strategy that makes four folds.</p>\n<p>Documentation, Code and Reproducibility:<br>\nPlease see the Kaggle link to my approach below.<br>\n<a href=\"https://www.kaggle.com/code/insiyajafferji/code-178th-priv-pub-0-759-0-531\" target=\"_blank\">https://www.kaggle.com/code/insiyajafferji/code-178th-priv-pub-0-759-0-531</a></p>",
      "rawMarkdown": "Many thanks to the organizers for hosting the Open Problems – Single-Cell Perturbations competition. Congratulations to the winners and everyone who participated, I really learnt alot from the discussions and great notebooks!\n\nI thought particularly great contributions and insights were from the following:\n\n@alexandervc\n@antoninadolgorukova\n@mehrankazeminia\n@somayyehgholami\n\n\nContext:\n\nBusiness context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\nData context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\nIntegration of Biological Knowledge:\n\nEven though I was not at the top of the leader-board I wanted to share my experience in this competition and my approach! \n\nOne of the objectives of this competition is to determine how compounds are involved in changes in gene expression within cells. For this competition, the organizers designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). 144 compounds were selected from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells.\n\nBefore I started to use ML models, I was interested in exploring the dataset within this competition to determine the compounds involved and the distribution of cell types associated with them as outlined in the following notebook:\n\nhttps://www.kaggle.com/code/insiyajafferji/op-single-cell-dataset-and-cell-type-distribution\n\nIn this notebook some initial dataset exploration helped me to understand cell type differences within PBMC samples treated with the compounds. I found that NK, CD4+ T cells, CD8+ T cells and Tregs are in all PBMC samples of the compounds being treated and there are a subset of PBMCs where compounds treated also have B and myeloid cells. PBMCs have been reported to contain a multitude of distinct multipotent progenitor cell populations and therefore treatment of certain drug compounds could have the ability to affect cell type distribution aswell as gene expression which can have an impact on the DE/DGE analysis.\n\nInterestingly some notebooks have suggested that the highest bias (differences between predicted and true values) and variability of gene DE predictions are related to individual drugs rather than cell types. \nhttps://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions\n\nBased on this information, the specific embeddings that I thought would help my ML models would be based on cell type, compound (sm_name and SMILES), and DE of genes (gene name).\n\nExploration of the problem:\n\nWhen I was exploring the ways to approach the challenges associated with this competition I found the following paper useful:\nhttps://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3\n\nHere, feature selection an essential technique for single-cell data analysis in high-dimensional datasets.\nImportantly, feature selection is an effective strategy to reduce the feature dimension and redundancy and can alleviate issues such as model overfitting in downstream analysis. Different from dimension reduction methods (e.g. principal component analysis) where features in a dataset are combined and/or transformed to derive a lower feature dimension, feature selection methods do not alter the original features in the dataset but only identify and select features that satisfy certain pre-defined criteria or optimise certain computational procedures. Some of the most popular research directions include selecting genes that can discriminate certain cell types. \n\nIn this competition setup, participants were tasked with modelling differential expression (DE), which enables us to estimate the impact of an experimental perturbation on the expression level of every gene in the transcription (18211 genes in this dataset). The Limma model was used to determine ‘differential expression’ (DE) methods for biological data analysis in this dataset.\n\nThe cell type proportion differs in PBMCs treated with different compounds and would therefore have different differential gene expression. For example T cells have a very distinct gene expression profile (i.e express CD3E) compared to B cells (CD79a, MSA41) and Myeloid cells (CD14, FCGR3A(CD16)). Also T cell sub-types also express differences, for example Regulatory T cells are classified by the expression of FOXP3, IL2RA and CTLA4 and Cytotoxic CD8 T cells are enriched in cytotoxic-related genes including GNLY, CCL5, NKG7, GZMH, LYZ, GZMB, and GZMK. CD4 T cells will express the CD4 gene whereas CD8 T cells will express CD8A and CD8B genes. Typically in single cell analysis the result is displayed as TSNE or UMAP plot showing distinct clusters with ideally one cell type and marker plots and DGE analysis can be used to confirm cell types  such as in the following paper (https://pubmed.ncbi.nlm.nih.gov/34911770/)\n\nThe following paper shows the combined protein and transcript analysis of single-cell RNA sequencing in human peripheral blood mononuclear cells (https://bmcbiol.biomedcentral.com/articles/10.1186/s12915-022-01382-4)\n\nSupervised ensemble classification models are popular among bioinformatics applications and have recently seen their increasing integration with deep learning models. Ensemble feature selection methods, typically, rely on either perturbation to the dataset or hyperparameters of the feature selection algorithms for creating ‘base selectors’ from which the ensemble could be derived. Generally, hybrid methods are motivated by the aim of taking advantage of the strengths of individual methods while alleviating or avoiding their weaknesses.\n\nModel design:\n\nI tried a number of approaches for the ML model and found that ensembeling models gave the best approach as described in section 2 of exploration of the problem. I used a combination of blends and ridge models.  Selection of certain cell types such as B cells, T regulatory cells and NK cells did help improve the model and including CD8 T cells gave a worse score.\n\nI used the features \"cell_type\" and \"sm_name\" and used certain a 'cell_type' with a selection of compounds that will have similar responses on each of these cell divisions. \n\nRMSE was calculated for each line and then to select the best model for each compound there were two submissions for every line that has RMSE >1\n\nI experimented with a number of blends and I would like to acknowledge the following notebooks from the following participants:\n-Daphne Anga https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457081\n-Mehran Kazeminia, Somayyeh Gholami https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\n\nFor including the Ridge model thanks to AMBROSM and MT for the following:\nhttps://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook\n\nRobustness:\nWhen using the cross-validation it did help in improving the robustness of the model in this competition and could help with over fitting.  I included the cross-validation strategy that makes four folds.\n\nDocumentation, Code and Reproducibility:\nPlease see the Kaggle link to my approach below.\nhttps://www.kaggle.com/code/insiyajafferji/code-178th-priv-pub-0-759-0-531",
      "votes": null
    },
    {
      "id": "2584443",
      "postDate": "01/03/2024 00:02:59",
      "content": "<p>Thanks for sharing your solution and also including the contributions of Alexander, Antonina, Somayeh and Mehran.</p>",
      "rawMarkdown": "Thanks for sharing your solution and also including the contributions of Alexander, Antonina, Somayeh and Mehran.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2584443,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "01/03/2024 00:02:59",
      "content": "<p>Thanks for sharing your solution and also including the contributions of Alexander, Antonina, Somayeh and Mehran.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2559545": "Many thanks to the organizers for hosting the Open Problems – Single-Cell Perturbations competition. Congratulations to the winners and everyone who participated, I really learnt alot from the discussions and great notebooks!\n\nI thought particularly great contributions and insights were from the following:\n\n@alexandervc\n@antoninadolgorukova\n@mehrankazeminia\n@somayyehgholami\n\n\nContext:\n\nBusiness context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\nData context: https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\nIntegration of Biological Knowledge:\n\nEven though I was not at the top of the leader-board I wanted to share my experience in this competition and my approach! \n\nOne of the objectives of this competition is to determine how compounds are involved in changes in gene expression within cells. For this competition, the organizers designed and generated a novel single-cell perturbational dataset in human peripheral blood mononuclear cells (PBMCs). 144 compounds were selected from the Library of Integrated Network-Based Cellular Signatures (LINCS) Connectivity Map dataset (PMID: 29195078) and measured single-cell gene expression profiles after 24 hours of treatment. The experiment was repeated in three healthy human donors, and the compounds were selected based on diverse transcriptional signatures observed in CD34+ hematopoietic stem cells.\n\nBefore I started to use ML models, I was interested in exploring the dataset within this competition to determine the compounds involved and the distribution of cell types associated with them as outlined in the following notebook:\n\nhttps://www.kaggle.com/code/insiyajafferji/op-single-cell-dataset-and-cell-type-distribution\n\nIn this notebook some initial dataset exploration helped me to understand cell type differences within PBMC samples treated with the compounds. I found that NK, CD4+ T cells, CD8+ T cells and Tregs are in all PBMC samples of the compounds being treated and there are a subset of PBMCs where compounds treated also have B and myeloid cells. PBMCs have been reported to contain a multitude of distinct multipotent progenitor cell populations and therefore treatment of certain drug compounds could have the ability to affect cell type distribution aswell as gene expression which can have an impact on the DE/DGE analysis.\n\nInterestingly some notebooks have suggested that the highest bias (differences between predicted and true values) and variability of gene DE predictions are related to individual drugs rather than cell types. \nhttps://www.kaggle.com/code/antoninadolgorukova/op2-analysis-of-different-models-predictions\n\nBased on this information, the specific embeddings that I thought would help my ML models would be based on cell type, compound (sm_name and SMILES), and DE of genes (gene name).\n\nExploration of the problem:\n\nWhen I was exploring the ways to approach the challenges associated with this competition I found the following paper useful:\nhttps://genomebiology.biomedcentral.com/articles/10.1186/s13059-021-02544-3\n\nHere, feature selection an essential technique for single-cell data analysis in high-dimensional datasets.\nImportantly, feature selection is an effective strategy to reduce the feature dimension and redundancy and can alleviate issues such as model overfitting in downstream analysis. Different from dimension reduction methods (e.g. principal component analysis) where features in a dataset are combined and/or transformed to derive a lower feature dimension, feature selection methods do not alter the original features in the dataset but only identify and select features that satisfy certain pre-defined criteria or optimise certain computational procedures. Some of the most popular research directions include selecting genes that can discriminate certain cell types. \n\nIn this competition setup, participants were tasked with modelling differential expression (DE), which enables us to estimate the impact of an experimental perturbation on the expression level of every gene in the transcription (18211 genes in this dataset). The Limma model was used to determine ‘differential expression’ (DE) methods for biological data analysis in this dataset.\n\nThe cell type proportion differs in PBMCs treated with different compounds and would therefore have different differential gene expression. For example T cells have a very distinct gene expression profile (i.e express CD3E) compared to B cells (CD79a, MSA41) and Myeloid cells (CD14, FCGR3A(CD16)). Also T cell sub-types also express differences, for example Regulatory T cells are classified by the expression of FOXP3, IL2RA and CTLA4 and Cytotoxic CD8 T cells are enriched in cytotoxic-related genes including GNLY, CCL5, NKG7, GZMH, LYZ, GZMB, and GZMK. CD4 T cells will express the CD4 gene whereas CD8 T cells will express CD8A and CD8B genes. Typically in single cell analysis the result is displayed as TSNE or UMAP plot showing distinct clusters with ideally one cell type and marker plots and DGE analysis can be used to confirm cell types  such as in the following paper (https://pubmed.ncbi.nlm.nih.gov/34911770/)\n\nThe following paper shows the combined protein and transcript analysis of single-cell RNA sequencing in human peripheral blood mononuclear cells (https://bmcbiol.biomedcentral.com/articles/10.1186/s12915-022-01382-4)\n\nSupervised ensemble classification models are popular among bioinformatics applications and have recently seen their increasing integration with deep learning models. Ensemble feature selection methods, typically, rely on either perturbation to the dataset or hyperparameters of the feature selection algorithms for creating ‘base selectors’ from which the ensemble could be derived. Generally, hybrid methods are motivated by the aim of taking advantage of the strengths of individual methods while alleviating or avoiding their weaknesses.\n\nModel design:\n\nI tried a number of approaches for the ML model and found that ensembeling models gave the best approach as described in section 2 of exploration of the problem. I used a combination of blends and ridge models.  Selection of certain cell types such as B cells, T regulatory cells and NK cells did help improve the model and including CD8 T cells gave a worse score.\n\nI used the features \"cell_type\" and \"sm_name\" and used certain a 'cell_type' with a selection of compounds that will have similar responses on each of these cell divisions. \n\nRMSE was calculated for each line and then to select the best model for each compound there were two submissions for every line that has RMSE >1\n\nI experimented with a number of blends and I would like to acknowledge the following notebooks from the following participants:\n-Daphne Anga https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/457081\n-Mehran Kazeminia, Somayyeh Gholami https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\n\nFor including the Ridge model thanks to AMBROSM and MT for the following:\nhttps://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook\n\nRobustness:\nWhen using the cross-validation it did help in improving the robustness of the model in this competition and could help with over fitting.  I included the cross-validation strategy that makes four folds.\n\nDocumentation, Code and Reproducibility:\nPlease see the Kaggle link to my approach below.\nhttps://www.kaggle.com/code/insiyajafferji/code-178th-priv-pub-0-759-0-531",
    "2584443": "Thanks for sharing your solution and also including the contributions of Alexander, Antonina, Somayeh and Mehran."
  },
  "source": "meta"
}