{
  "id": 460693,
  "title": "The 440th solution of the 'Open Problems – Single-Cell Perturbations' competition",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/bio-stat-uwm-the-440th-solution-of-the-open-proble",
  "author_name": "",
  "post_date": "2023-12-10T17:09:04.184422100Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Business context:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></p>\n<p>Data context: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></p>\n<p>Code link:<br>\n<a href=\"https://github.com/900Step/Open-challenge\" target=\"_blank\">https://github.com/900Step/Open-challenge</a></p>\n<p>Overview<br>\nThis solution considers the impact of drugs on the gene expression of unknown cell types from three aspects: dimensional change, feature construction, and regression algorithms. Due to the high dimensionality of genes, PCA is used to reduce the dimensions of the response variable. Natural language processing, combined with SMILE chemical structures, is employed to encode the drugs. For the regression algorithm, we attempt to use a simple regression model (SVR) for prediction. A tensor model is utilized to conduct stratified analysis on plate-class data. The regression effect of experiments on linear combinations of some public datasets.</p>\n<p>Data Processing</p>\n<p>Due to the project providing both public and private test sets, we do not perform any additional train-test splitting during the modeling process.</p>\n<p>Response variable<br>\nOur objective is to predict highly-dimensional gene expression data using existing cell type and drug information. Initially, this requires altering the data dimensions and structure. Given over 18,000 genes as response variables and the limited sample size, it becomes challenging to draw reliable conclusions. Further-more, this scenario often leads to model overfitting and weak interpretability. Hence, analogous to the independent variables, it is imperative to apply dimensionality reduction to the response variables. We employ the Principal Component Analysis (PCA) method, preserving 10-20 dimensions to maintain a cumulative contribution rate of 95%, thus constructing new response variables.</p>\n<p>Feature construction<br>\nThe column ”SMILE” (Simplified Molecular Input Line Entry System [2]) encompasses the chemical properties of drugs, for instance, Clotrimazole, with its chemical encoding as<br>\n”Clc1ccccc1C(c1ccccc1)(c1ccccc1)n1ccnc1”. We utilize natural language processing techniques<br>\nto encode the chemical expressions of drugs using strings. For example, ”c1ccccc1” represents<br>\na benzene ring, and ”Cl” signifies a chlorine atom. Certain punctuation marks also bear specific meanings, such as ”=” for a double bond and ”#” for a triple bond. Through encoding chemical structures, we extract chemical structural features pertinent to the drugs. The S and R configurations in chiral structures are denoted by ”@” and ”@@”, respectively.</p>\n<p>After these transformations, we convert the boolean values of ”control” to {0,1}, remove the original three columns of drug information from the dataset, and thus obtain a preliminarily processed dataset.</p>\n<p>Modeling</p>\n<p>SVR<br>\nUsing the SVM-based regression algorithm from scikit-learn for prediction, it can be anticipated that due to the low dimensionality of the independent variables and the complex relationships of the response variables, there will be a significant prediction error. This method does not incorporate any specificity processing based on the model's background.</p>\n<p>Tensor<br>\nWe adapted Algorithm 3 from the article \"Tensor Completion with Noisy Side Information\" for tensor-based analysis. This approach involved transforming the original dataset into a 3-dimensional tensor object with the dimensions (146, 18211, 6), corresponding to (drug, gene, cell type). To address missing values in the gene expression data, a \"warm start\" matrix is employed. This matrix is initially populated with zeroes or randomized values for the absent data points. This technique provides a preliminary structure, aiding in more effectively estimating and filling in the missing gene expression values. </p>\n<p>CPA method (didn’t been realized but we think it is a good method)<br>\nI explored the CPA (Component-Based Pathway Analysis) method, which aims to uncover latent variables that influence response prediction. This algorithm converts gene expression data, along with all other features, into the 'adata' data type. By using a more complex data structure, it enables the full utilization of the information. Our goal is to garner more gene-related information, such as gene length, expression characteristics, and whether certain genes exhibit synergistic or antagonistic effects with specific cells or drugs. Enriching the feature information of the existing data could significantly reduce the prediction error rate for differential expression in the forecasting process.</p>\n<p>In Kaggle, there are some public datasets related to specific competitions. Building upon the corresponding files, I attempted to apply a linear combination with weighting to them. The results of this approach were superior to those of each individual file used in the process.<br>\n/kaggle/input/open-problems-2-submits-collection<br>\n/kaggle/input/op2-603<br>\n/kaggle/input/op2-604<br>\n/kaggle/input/op2-607<br>\n/kaggle/input/op2-720</p>",
  "messages": [
    {
      "id": "2556430",
      "postDate": "12/10/2023 17:09:04",
      "content": "<p>Business context:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></p>\n<p>Data context: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></p>\n<p>Code link:<br>\n<a href=\"https://github.com/900Step/Open-challenge\" target=\"_blank\">https://github.com/900Step/Open-challenge</a></p>\n<p>Overview<br>\nThis solution considers the impact of drugs on the gene expression of unknown cell types from three aspects: dimensional change, feature construction, and regression algorithms. Due to the high dimensionality of genes, PCA is used to reduce the dimensions of the response variable. Natural language processing, combined with SMILE chemical structures, is employed to encode the drugs. For the regression algorithm, we attempt to use a simple regression model (SVR) for prediction. A tensor model is utilized to conduct stratified analysis on plate-class data. The regression effect of experiments on linear combinations of some public datasets.</p>\n<p>Data Processing</p>\n<p>Due to the project providing both public and private test sets, we do not perform any additional train-test splitting during the modeling process.</p>\n<p>Response variable<br>\nOur objective is to predict highly-dimensional gene expression data using existing cell type and drug information. Initially, this requires altering the data dimensions and structure. Given over 18,000 genes as response variables and the limited sample size, it becomes challenging to draw reliable conclusions. Further-more, this scenario often leads to model overfitting and weak interpretability. Hence, analogous to the independent variables, it is imperative to apply dimensionality reduction to the response variables. We employ the Principal Component Analysis (PCA) method, preserving 10-20 dimensions to maintain a cumulative contribution rate of 95%, thus constructing new response variables.</p>\n<p>Feature construction<br>\nThe column ”SMILE” (Simplified Molecular Input Line Entry System [2]) encompasses the chemical properties of drugs, for instance, Clotrimazole, with its chemical encoding as<br>\n”Clc1ccccc1C(c1ccccc1)(c1ccccc1)n1ccnc1”. We utilize natural language processing techniques<br>\nto encode the chemical expressions of drugs using strings. For example, ”c1ccccc1” represents<br>\na benzene ring, and ”Cl” signifies a chlorine atom. Certain punctuation marks also bear specific meanings, such as ”=” for a double bond and ”#” for a triple bond. Through encoding chemical structures, we extract chemical structural features pertinent to the drugs. The S and R configurations in chiral structures are denoted by ”@” and ”@@”, respectively.</p>\n<p>After these transformations, we convert the boolean values of ”control” to {0,1}, remove the original three columns of drug information from the dataset, and thus obtain a preliminarily processed dataset.</p>\n<p>Modeling</p>\n<p>SVR<br>\nUsing the SVM-based regression algorithm from scikit-learn for prediction, it can be anticipated that due to the low dimensionality of the independent variables and the complex relationships of the response variables, there will be a significant prediction error. This method does not incorporate any specificity processing based on the model's background.</p>\n<p>Tensor<br>\nWe adapted Algorithm 3 from the article \"Tensor Completion with Noisy Side Information\" for tensor-based analysis. This approach involved transforming the original dataset into a 3-dimensional tensor object with the dimensions (146, 18211, 6), corresponding to (drug, gene, cell type). To address missing values in the gene expression data, a \"warm start\" matrix is employed. This matrix is initially populated with zeroes or randomized values for the absent data points. This technique provides a preliminary structure, aiding in more effectively estimating and filling in the missing gene expression values. </p>\n<p>CPA method (didn’t been realized but we think it is a good method)<br>\nI explored the CPA (Component-Based Pathway Analysis) method, which aims to uncover latent variables that influence response prediction. This algorithm converts gene expression data, along with all other features, into the 'adata' data type. By using a more complex data structure, it enables the full utilization of the information. Our goal is to garner more gene-related information, such as gene length, expression characteristics, and whether certain genes exhibit synergistic or antagonistic effects with specific cells or drugs. Enriching the feature information of the existing data could significantly reduce the prediction error rate for differential expression in the forecasting process.</p>\n<p>In Kaggle, there are some public datasets related to specific competitions. Building upon the corresponding files, I attempted to apply a linear combination with weighting to them. The results of this approach were superior to those of each individual file used in the process.<br>\n/kaggle/input/open-problems-2-submits-collection<br>\n/kaggle/input/op2-603<br>\n/kaggle/input/op2-604<br>\n/kaggle/input/op2-607<br>\n/kaggle/input/op2-720</p>",
      "rawMarkdown": "Business context:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n\nData context: \nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\nCode link:\nhttps://github.com/900Step/Open-challenge\n\nOverview\nThis solution considers the impact of drugs on the gene expression of unknown cell types from three aspects: dimensional change, feature construction, and regression algorithms. Due to the high dimensionality of genes, PCA is used to reduce the dimensions of the response variable. Natural language processing, combined with SMILE chemical structures, is employed to encode the drugs. For the regression algorithm, we attempt to use a simple regression model (SVR) for prediction. A tensor model is utilized to conduct stratified analysis on plate-class data. The regression effect of experiments on linear combinations of some public datasets.\n\nData Processing\n\nDue to the project providing both public and private test sets, we do not perform any additional train-test splitting during the modeling process.\n\nResponse variable\nOur objective is to predict highly-dimensional gene expression data using existing cell type and drug information. Initially, this requires altering the data dimensions and structure. Given over 18,000 genes as response variables and the limited sample size, it becomes challenging to draw reliable conclusions. Further-more, this scenario often leads to model overfitting and weak interpretability. Hence, analogous to the independent variables, it is imperative to apply dimensionality reduction to the response variables. We employ the Principal Component Analysis (PCA) method, preserving 10-20 dimensions to maintain a cumulative contribution rate of 95%, thus constructing new response variables.\n\nFeature construction\nThe column ”SMILE” (Simplified Molecular Input Line Entry System [2]) encompasses the chemical properties of drugs, for instance, Clotrimazole, with its chemical encoding as\n”Clc1ccccc1C(c1ccccc1)(c1ccccc1)n1ccnc1”. We utilize natural language processing techniques\nto encode the chemical expressions of drugs using strings. For example, ”c1ccccc1” represents\na benzene ring, and ”Cl” signifies a chlorine atom. Certain punctuation marks also bear specific meanings, such as ”=” for a double bond and ”#” for a triple bond. Through encoding chemical structures, we extract chemical structural features pertinent to the drugs. The S and R configurations in chiral structures are denoted by ”@” and ”@@”, respectively.\n\nAfter these transformations, we convert the boolean values of ”control” to {0,1}, remove the original three columns of drug information from the dataset, and thus obtain a preliminarily processed dataset.\n\nModeling\n\nSVR\nUsing the SVM-based regression algorithm from scikit-learn for prediction, it can be anticipated that due to the low dimensionality of the independent variables and the complex relationships of the response variables, there will be a significant prediction error. This method does not incorporate any specificity processing based on the model's background.\n\nTensor\nWe adapted Algorithm 3 from the article \"Tensor Completion with Noisy Side Information\" for tensor-based analysis. This approach involved transforming the original dataset into a 3-dimensional tensor object with the dimensions (146, 18211, 6), corresponding to (drug, gene, cell type). To address missing values in the gene expression data, a \"warm start\" matrix is employed. This matrix is initially populated with zeroes or randomized values for the absent data points. This technique provides a preliminary structure, aiding in more effectively estimating and filling in the missing gene expression values. \n\nCPA method (didn’t been realized but we think it is a good method)\nI explored the CPA (Component-Based Pathway Analysis) method, which aims to uncover latent variables that influence response prediction. This algorithm converts gene expression data, along with all other features, into the 'adata' data type. By using a more complex data structure, it enables the full utilization of the information. Our goal is to garner more gene-related information, such as gene length, expression characteristics, and whether certain genes exhibit synergistic or antagonistic effects with specific cells or drugs. Enriching the feature information of the existing data could significantly reduce the prediction error rate for differential expression in the forecasting process.\n\nIn Kaggle, there are some public datasets related to specific competitions. Building upon the corresponding files, I attempted to apply a linear combination with weighting to them. The results of this approach were superior to those of each individual file used in the process.\n/kaggle/input/open-problems-2-submits-collection\n/kaggle/input/op2-603\n/kaggle/input/op2-604\n/kaggle/input/op2-607\n/kaggle/input/op2-720",
      "votes": null
    },
    {
      "id": "2556437",
      "postDate": "12/10/2023 17:13:42",
      "content": "<p>reference:<br>\nDaniel Burkhardt, Andrew Benz, Richard Lieberman, Scott Gigante, Ashley Chow, Ryan Holbrook, Robrecht Cannoodt, Malte Luecken. (2023). Open Problems – Single-Cell Perturbations. Kaggle. <a href=\"https://kaggle.com/competitions/open-problems-single-cell-perturbations\" target=\"_blank\">https://kaggle.com/competitions/open-problems-single-cell-perturbations</a><br>\nDeutsch, Francine M., Dorothy LeBaron, and Maury March Fryer. \"What is in a smile?.\" Psychology of Women Quarterly 11.3 (1987): 341-352.<br>\nBertsimas, Dimitris, and Colin Pawlowski. \"Tensor completion with noisy side information.\" Machine Learning 112.10 (2023): 3945-3976.<br>\nLotfollahi, Mohammad, et al. \"Predicting cellular responses to complex perturbations in high‐throughput screens.\" Molecular Systems Biology (2023): e11517.</p>",
      "rawMarkdown": "reference:\nDaniel Burkhardt, Andrew Benz, Richard Lieberman, Scott Gigante, Ashley Chow, Ryan Holbrook, Robrecht Cannoodt, Malte Luecken. (2023). Open Problems – Single-Cell Perturbations. Kaggle. https://kaggle.com/competitions/open-problems-single-cell-perturbations\nDeutsch, Francine M., Dorothy LeBaron, and Maury March Fryer. \"What is in a smile?.\" Psychology of Women Quarterly 11.3 (1987): 341-352.\nBertsimas, Dimitris, and Colin Pawlowski. \"Tensor completion with noisy side information.\" Machine Learning 112.10 (2023): 3945-3976.\nLotfollahi, Mohammad, et al. \"Predicting cellular responses to complex perturbations in high‐throughput screens.\" Molecular Systems Biology (2023): e11517.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2556437,
      "author_name": "bhengchen",
      "author_url": "",
      "post_date": "12/10/2023 17:13:42",
      "content": "<p>reference:<br>\nDaniel Burkhardt, Andrew Benz, Richard Lieberman, Scott Gigante, Ashley Chow, Ryan Holbrook, Robrecht Cannoodt, Malte Luecken. (2023). Open Problems – Single-Cell Perturbations. Kaggle. <a href=\"https://kaggle.com/competitions/open-problems-single-cell-perturbations\" target=\"_blank\">https://kaggle.com/competitions/open-problems-single-cell-perturbations</a><br>\nDeutsch, Francine M., Dorothy LeBaron, and Maury March Fryer. \"What is in a smile?.\" Psychology of Women Quarterly 11.3 (1987): 341-352.<br>\nBertsimas, Dimitris, and Colin Pawlowski. \"Tensor completion with noisy side information.\" Machine Learning 112.10 (2023): 3945-3976.<br>\nLotfollahi, Mohammad, et al. \"Predicting cellular responses to complex perturbations in high‐throughput screens.\" Molecular Systems Biology (2023): e11517.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2556430": "Business context:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n\nData context: \nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\nCode link:\nhttps://github.com/900Step/Open-challenge\n\nOverview\nThis solution considers the impact of drugs on the gene expression of unknown cell types from three aspects: dimensional change, feature construction, and regression algorithms. Due to the high dimensionality of genes, PCA is used to reduce the dimensions of the response variable. Natural language processing, combined with SMILE chemical structures, is employed to encode the drugs. For the regression algorithm, we attempt to use a simple regression model (SVR) for prediction. A tensor model is utilized to conduct stratified analysis on plate-class data. The regression effect of experiments on linear combinations of some public datasets.\n\nData Processing\n\nDue to the project providing both public and private test sets, we do not perform any additional train-test splitting during the modeling process.\n\nResponse variable\nOur objective is to predict highly-dimensional gene expression data using existing cell type and drug information. Initially, this requires altering the data dimensions and structure. Given over 18,000 genes as response variables and the limited sample size, it becomes challenging to draw reliable conclusions. Further-more, this scenario often leads to model overfitting and weak interpretability. Hence, analogous to the independent variables, it is imperative to apply dimensionality reduction to the response variables. We employ the Principal Component Analysis (PCA) method, preserving 10-20 dimensions to maintain a cumulative contribution rate of 95%, thus constructing new response variables.\n\nFeature construction\nThe column ”SMILE” (Simplified Molecular Input Line Entry System [2]) encompasses the chemical properties of drugs, for instance, Clotrimazole, with its chemical encoding as\n”Clc1ccccc1C(c1ccccc1)(c1ccccc1)n1ccnc1”. We utilize natural language processing techniques\nto encode the chemical expressions of drugs using strings. For example, ”c1ccccc1” represents\na benzene ring, and ”Cl” signifies a chlorine atom. Certain punctuation marks also bear specific meanings, such as ”=” for a double bond and ”#” for a triple bond. Through encoding chemical structures, we extract chemical structural features pertinent to the drugs. The S and R configurations in chiral structures are denoted by ”@” and ”@@”, respectively.\n\nAfter these transformations, we convert the boolean values of ”control” to {0,1}, remove the original three columns of drug information from the dataset, and thus obtain a preliminarily processed dataset.\n\nModeling\n\nSVR\nUsing the SVM-based regression algorithm from scikit-learn for prediction, it can be anticipated that due to the low dimensionality of the independent variables and the complex relationships of the response variables, there will be a significant prediction error. This method does not incorporate any specificity processing based on the model's background.\n\nTensor\nWe adapted Algorithm 3 from the article \"Tensor Completion with Noisy Side Information\" for tensor-based analysis. This approach involved transforming the original dataset into a 3-dimensional tensor object with the dimensions (146, 18211, 6), corresponding to (drug, gene, cell type). To address missing values in the gene expression data, a \"warm start\" matrix is employed. This matrix is initially populated with zeroes or randomized values for the absent data points. This technique provides a preliminary structure, aiding in more effectively estimating and filling in the missing gene expression values. \n\nCPA method (didn’t been realized but we think it is a good method)\nI explored the CPA (Component-Based Pathway Analysis) method, which aims to uncover latent variables that influence response prediction. This algorithm converts gene expression data, along with all other features, into the 'adata' data type. By using a more complex data structure, it enables the full utilization of the information. Our goal is to garner more gene-related information, such as gene length, expression characteristics, and whether certain genes exhibit synergistic or antagonistic effects with specific cells or drugs. Enriching the feature information of the existing data could significantly reduce the prediction error rate for differential expression in the forecasting process.\n\nIn Kaggle, there are some public datasets related to specific competitions. Building upon the corresponding files, I attempted to apply a linear combination with weighting to them. The results of this approach were superior to those of each individual file used in the process.\n/kaggle/input/open-problems-2-submits-collection\n/kaggle/input/op2-603\n/kaggle/input/op2-604\n/kaggle/input/op2-607\n/kaggle/input/op2-720",
    "2556437": "reference:\nDaniel Burkhardt, Andrew Benz, Richard Lieberman, Scott Gigante, Ashley Chow, Ryan Holbrook, Robrecht Cannoodt, Malte Luecken. (2023). Open Problems – Single-Cell Perturbations. Kaggle. https://kaggle.com/competitions/open-problems-single-cell-perturbations\nDeutsch, Francine M., Dorothy LeBaron, and Maury March Fryer. \"What is in a smile?.\" Psychology of Women Quarterly 11.3 (1987): 341-352.\nBertsimas, Dimitris, and Colin Pawlowski. \"Tensor completion with noisy side information.\" Machine Learning 112.10 (2023): 3945-3976.\nLotfollahi, Mohammad, et al. \"Predicting cellular responses to complex perturbations in high‐throughput screens.\" Molecular Systems Biology (2023): e11517."
  },
  "source": "meta"
}