{
  "id": 459623,
  "title": "7th Place Solution for the Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/459623",
  "author_name": "Zhijian Li",
  "post_date": "2023-12-06T02:31:57.590000",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thank the organizers for hosting this interesting competition and congrats to the Winners.<br>\nIt was a big surprise for me to see my resolution finally achieve 7th place. </p>\n<h2>Context</h2>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></p>\n<h2>Overview of the approach</h2>\n<p>My approach has two major steps: i) learn embeddings for all cell types, small molecular, and genes,  ii) use embedding as features to train a deep learning model to predict the target while accounting for overfitting. Throughout the competition, I only used a fully connected network with three layers. My best submission is also based on a single FC model.</p>\n<h3>Learning embeddings</h3>\n<p>In this step, the goal is to learn a specific embedding for each cell type, molecular, and gene. I was strongly inspired by <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6\" target=\"_blank\">this paper</a>, where the authors used deep tensor factorization to learn a dense, information-rich representation for cell type, experimental assay, and genomic position.  </p>\n<p>Basically, I used the same approach but played around with the model architectures and parameters, such as the number of latent factors, and the number of dimensions of the network, as well as how to combine the features (e.g., concatenate vs. additive). </p>\n<p>In the end, my mode is the following:</p>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n\n        self.cell_types = cell_types\n        self.compounds = compounds\n        self.genes = genes\n\n        self.n_cell_types = (cell_types)\n        self.n_compounds = (compounds)\n        self.n_genes = (genes)\n\n        self.n_cell_type_factors = n_cell_type_factors\n        self.n_compounds_factors = n_compounds_factors\n        self.n_gene_factors = n_gene_factors\n\n        self.cell_type_embedding = torch.nn.Embedding(self.n_cell_types, self.n_cell_type_factors)\n        self.compound_embedding = torch.nn.Embedding(self.n_compounds, self.n_compounds_factors)\n        self.gene_embedding = torch.nn.Embedding(self.n_genes, self.n_gene_factors)\n\n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n        self.n_factors = n_cell_type_factors + n_compounds_factors + n_gene_factors\n\n        self.model = nn.Sequential(nn.Linear(self.n_factors, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, ))\n\n     ():\n        cell_type_vec = self.cell_type_embedding(cell_type_indices)\n        compound_vec = self.compound_embedding(compound_indices)\n        gene_vec = self.gene_embedding(gene_indices)\n\n        x = torch.concat([cell_type_vec, compound_vec, gene_vec], dim=)\n        x = self.model(x)\n\n         x\n</code></pre>\n<p>To train this model, I used all the data from <code>de_train.parquet</code> and converted the table as follows:</p>\n<pre><code>\ndf = pd.read_parquet()\ndf = df.sort_values([, ])\ndf = df.drop([, , ], axis=)\ndf = pd.melt(df, id_vars=[, ], var_name=, value_name=)\n</code></pre>\n<p>The training data looks like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Ffcf2b04f2255dc9fe2f0382320346feb%2FScreenshot%202023-12-05%20at%2021.02.43.png?generation=1701828192377214&amp;alt=media\" alt=\"\"></p>\n<p>Here, I trained the model with 100 epochs without validation, because I will only the embedding layer for Step 2.</p>\n<p>To check if the model learns meaningful embedding for cell types, molecules, and genes, I also visualized the embedding with UMAP. For example, below is the 2D UMAP of gene embeddings:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F59782ef056d394c5deb99376c0af2a05%2FScreenshot%202023-12-05%20at%2021.05.28.png?generation=1701828348397738&amp;alt=media\" alt=\"\"></p>\n<p>By eyeballing, it looks like there are some structures for the genes.</p>\n<h3>Predicting target</h3>\n<p>Once I obtained the embeddings, I trained another model to predict the target, and this model was also used to generate the final submission. Again, I used a FC network as follows:</p>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n\n        self.n_input = n_input\n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n\n        self.model = nn.Sequential(nn.Linear(self.n_input, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, ))\n\n     ():\n        x = self.model(x)\n\n         x\n</code></pre>\n<p>To prevent overfitting, I used each of the cell types as validation data, which means my model was based on 4-fold cross-validation. For the compounds, I followed the notebook here to select the private test compounds for validation:</p>\n<pre><code> key, cell_type  cell_type_names.items():\n    (cell_type)\n\n    \n    df_train = df[(df[] != key) | ~df[].isin(privte_ids)]\n    df_valid = df[(df[] == key) &amp; df[].isin(privte_ids)]\n\n    df_train = df_train.sort_values([, ])\n    df_valid = df_valid.sort_values()\n\n    df_train = convert_to_long_df(df_train)\n    df_valid = convert_to_long_df(df_valid)\n\n    df_train.to_csv()\n    df_valid.to_csv()\n</code></pre>\n<p>For model training, I used the following strategies:</p>\n<pre><code>    \n    criterion = torch.nn.MSELoss()\n    optimizer = AdamW(model.parameters(), lr=args.lr, weight_decay=)\n    scheduler = ReduceLROnPlateau(optimizer, , min_lr=)\n</code></pre>\n<p>I trained one model by using each of the cell types as validation data, and for the final submission, I just averaged the predictions.</p>\n<h2>What didn't work for me</h2>\n<p>During the competition, I spent a lot of time including the features based on prior knowledge, such as single-cell data, molecular structure embedding, and gene embedding for other large models, such as geneFormer. However, they all didn't work out. So my final model was just based on the features learned from Step 1.</p>\n<h2>Source</h2>\n<p><a href=\"https://github.com/lzj1769/7th_place_solution_Single-Cell-Perturbations\" target=\"_blank\">https://github.com/lzj1769/7th_place_solution_Single-Cell-Perturbations</a><br>\n<a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6</a><br>\n<a href=\"https://github.com/jmschrei/avocado\" target=\"_blank\">https://github.com/jmschrei/avocado</a></p>",
  "messages": [
    {
      "id": 2550395,
      "postDate": "2023-12-06T02:31:57.590Z",
      "content": "<p>Thank the organizers for hosting this interesting competition and congrats to the Winners.<br>\nIt was a big surprise for me to see my resolution finally achieve 7th place. </p>\n<h2>Context</h2>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></p>\n<h2>Overview of the approach</h2>\n<p>My approach has two major steps: i) learn embeddings for all cell types, small molecular, and genes,  ii) use embedding as features to train a deep learning model to predict the target while accounting for overfitting. Throughout the competition, I only used a fully connected network with three layers. My best submission is also based on a single FC model.</p>\n<h3>Learning embeddings</h3>\n<p>In this step, the goal is to learn a specific embedding for each cell type, molecular, and gene. I was strongly inspired by <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6\" target=\"_blank\">this paper</a>, where the authors used deep tensor factorization to learn a dense, information-rich representation for cell type, experimental assay, and genomic position.  </p>\n<p>Basically, I used the same approach but played around with the model architectures and parameters, such as the number of latent factors, and the number of dimensions of the network, as well as how to combine the features (e.g., concatenate vs. additive). </p>\n<p>In the end, my mode is the following:</p>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n\n        self.cell_types = cell_types\n        self.compounds = compounds\n        self.genes = genes\n\n        self.n_cell_types = (cell_types)\n        self.n_compounds = (compounds)\n        self.n_genes = (genes)\n\n        self.n_cell_type_factors = n_cell_type_factors\n        self.n_compounds_factors = n_compounds_factors\n        self.n_gene_factors = n_gene_factors\n\n        self.cell_type_embedding = torch.nn.Embedding(self.n_cell_types, self.n_cell_type_factors)\n        self.compound_embedding = torch.nn.Embedding(self.n_compounds, self.n_compounds_factors)\n        self.gene_embedding = torch.nn.Embedding(self.n_genes, self.n_gene_factors)\n\n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n        self.n_factors = n_cell_type_factors + n_compounds_factors + n_gene_factors\n\n        self.model = nn.Sequential(nn.Linear(self.n_factors, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, ))\n\n     ():\n        cell_type_vec = self.cell_type_embedding(cell_type_indices)\n        compound_vec = self.compound_embedding(compound_indices)\n        gene_vec = self.gene_embedding(gene_indices)\n\n        x = torch.concat([cell_type_vec, compound_vec, gene_vec], dim=)\n        x = self.model(x)\n\n         x\n</code></pre>\n<p>To train this model, I used all the data from <code>de_train.parquet</code> and converted the table as follows:</p>\n<pre><code>\ndf = pd.read_parquet()\ndf = df.sort_values([, ])\ndf = df.drop([, , ], axis=)\ndf = pd.melt(df, id_vars=[, ], var_name=, value_name=)\n</code></pre>\n<p>The training data looks like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Ffcf2b04f2255dc9fe2f0382320346feb%2FScreenshot%202023-12-05%20at%2021.02.43.png?generation=1701828192377214&amp;alt=media\" alt=\"\"></p>\n<p>Here, I trained the model with 100 epochs without validation, because I will only the embedding layer for Step 2.</p>\n<p>To check if the model learns meaningful embedding for cell types, molecules, and genes, I also visualized the embedding with UMAP. For example, below is the 2D UMAP of gene embeddings:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F59782ef056d394c5deb99376c0af2a05%2FScreenshot%202023-12-05%20at%2021.05.28.png?generation=1701828348397738&amp;alt=media\" alt=\"\"></p>\n<p>By eyeballing, it looks like there are some structures for the genes.</p>\n<h3>Predicting target</h3>\n<p>Once I obtained the embeddings, I trained another model to predict the target, and this model was also used to generate the final submission. Again, I used a FC network as follows:</p>\n<pre><code> (torch.nn.Module):\n     ():\n        ().__init__()\n\n        self.n_input = n_input\n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n\n        self.model = nn.Sequential(nn.Linear(self.n_input, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, ))\n\n     ():\n        x = self.model(x)\n\n         x\n</code></pre>\n<p>To prevent overfitting, I used each of the cell types as validation data, which means my model was based on 4-fold cross-validation. For the compounds, I followed the notebook here to select the private test compounds for validation:</p>\n<pre><code> key, cell_type  cell_type_names.items():\n    (cell_type)\n\n    \n    df_train = df[(df[] != key) | ~df[].isin(privte_ids)]\n    df_valid = df[(df[] == key) &amp; df[].isin(privte_ids)]\n\n    df_train = df_train.sort_values([, ])\n    df_valid = df_valid.sort_values()\n\n    df_train = convert_to_long_df(df_train)\n    df_valid = convert_to_long_df(df_valid)\n\n    df_train.to_csv()\n    df_valid.to_csv()\n</code></pre>\n<p>For model training, I used the following strategies:</p>\n<pre><code>    \n    criterion = torch.nn.MSELoss()\n    optimizer = AdamW(model.parameters(), lr=args.lr, weight_decay=)\n    scheduler = ReduceLROnPlateau(optimizer, , min_lr=)\n</code></pre>\n<p>I trained one model by using each of the cell types as validation data, and for the final submission, I just averaged the predictions.</p>\n<h2>What didn't work for me</h2>\n<p>During the competition, I spent a lot of time including the features based on prior knowledge, such as single-cell data, molecular structure embedding, and gene embedding for other large models, such as geneFormer. However, they all didn't work out. So my final model was just based on the features learned from Step 1.</p>\n<h2>Source</h2>\n<p><a href=\"https://github.com/lzj1769/7th_place_solution_Single-Cell-Perturbations\" target=\"_blank\">https://github.com/lzj1769/7th_place_solution_Single-Cell-Perturbations</a><br>\n<a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6\" target=\"_blank\">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6</a><br>\n<a href=\"https://github.com/jmschrei/avocado\" target=\"_blank\">https://github.com/jmschrei/avocado</a></p>",
      "rawMarkdown": "Thank the organizers for hosting this interesting competition and congrats to the Winners.\nIt was a big surprise for me to see my resolution finally achieve 7th place. \n\n## Context\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n## Overview of the approach\n\nMy approach has two major steps: i) learn embeddings for all cell types, small molecular, and genes,  ii) use embedding as features to train a deep learning model to predict the target while accounting for overfitting. Throughout the competition, I only used a fully connected network with three layers. My best submission is also based on a single FC model.\n\n### Learning embeddings\n\nIn this step, the goal is to learn a specific embedding for each cell type, molecular, and gene. I was strongly inspired by [this paper](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6), where the authors used deep tensor factorization to learn a dense, information-rich representation for cell type, experimental assay, and genomic position.  \n\nBasically, I used the same approach but played around with the model architectures and parameters, such as the number of latent factors, and the number of dimensions of the network, as well as how to combine the features (e.g., concatenate vs. additive). \n\nIn the end, my mode is the following:\n```python\nclass DeepTensorFactorization(torch.nn.Module):\n    def __init__(self, \n                 cell_types, \n                 compounds, \n                 genes, \n                 n_cell_type_factors: int=4, \n                 n_compounds_factors: int=16, \n                 n_gene_factors: int=128,\n                 n_hiddens: int=2048,\n                 dropout: float=0.1):\n        super().__init__()\n        \n        self.cell_types = cell_types\n        self.compounds = compounds\n        self.genes = genes\n        \n        self.n_cell_types = len(cell_types)\n        self.n_compounds = len(compounds)\n        self.n_genes = len(genes)\n        \n        self.n_cell_type_factors = n_cell_type_factors\n        self.n_compounds_factors = n_compounds_factors\n        self.n_gene_factors = n_gene_factors\n        \n        self.cell_type_embedding = torch.nn.Embedding(self.n_cell_types, self.n_cell_type_factors)\n        self.compound_embedding = torch.nn.Embedding(self.n_compounds, self.n_compounds_factors)\n        self.gene_embedding = torch.nn.Embedding(self.n_genes, self.n_gene_factors)\n        \n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n        self.n_factors = n_cell_type_factors + n_compounds_factors + n_gene_factors\n        \n        self.model = nn.Sequential(nn.Linear(self.n_factors, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, 1))\n        \n    def forward(self, cell_type_indices, compound_indices, gene_indices):\n        cell_type_vec = self.cell_type_embedding(cell_type_indices)\n        compound_vec = self.compound_embedding(compound_indices)\n        gene_vec = self.gene_embedding(gene_indices)\n        \n        x = torch.concat([cell_type_vec, compound_vec, gene_vec], dim=1)\n        x = self.model(x)\n        \n        return x\n```\nTo train this model, I used all the data from `de_train.parquet` and converted the table as follows:\n```python\n# convert to long df\ndf = pd.read_parquet('de_train.parquet')\ndf = df.sort_values(['cell_type', 'sm_name'])\ndf = df.drop(['sm_lincs_id', 'SMILES', 'control'], axis=1)\ndf = pd.melt(df, id_vars=['cell_type', 'sm_name'], var_name='gene', value_name='target')\n```\nThe training data looks like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Ffcf2b04f2255dc9fe2f0382320346feb%2FScreenshot%202023-12-05%20at%2021.02.43.png?generation=1701828192377214&alt=media)\n\nHere, I trained the model with 100 epochs without validation, because I will only the embedding layer for Step 2.\n\nTo check if the model learns meaningful embedding for cell types, molecules, and genes, I also visualized the embedding with UMAP. For example, below is the 2D UMAP of gene embeddings:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F59782ef056d394c5deb99376c0af2a05%2FScreenshot%202023-12-05%20at%2021.05.28.png?generation=1701828348397738&alt=media)\n\n\nBy eyeballing, it looks like there are some structures for the genes.\n\n### Predicting target  \n\nOnce I obtained the embeddings, I trained another model to predict the target, and this model was also used to generate the final submission. Again, I used a FC network as follows:\n```python\nclass PerturbNet(torch.nn.Module):\n    def __init__(self, \n                 n_input: int=148,\n                 n_hiddens: int=2048,\n                 dropout: float=0.5):\n        super().__init__()\n        \n        self.n_input = n_input\n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n\n        self.model = nn.Sequential(nn.Linear(self.n_input, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, 1))\n        \n    def forward(self, x):\n        x = self.model(x)\n        \n        return x\n```\n\nTo prevent overfitting, I used each of the cell types as validation data, which means my model was based on 4-fold cross-validation. For the compounds, I followed the notebook here to select the private test compounds for validation:\n\n```python\nfor key, cell_type in cell_type_names.items():\n    print(cell_type)\n    \n    # split data for training and validation, here we used the private test compounds for validation\n    df_train = df[(df['cell_type'] != key) | ~df['sm_lincs_id'].isin(privte_ids)]\n    df_valid = df[(df['cell_type'] == key) & df['sm_lincs_id'].isin(privte_ids)]\n    \n    df_train = df_train.sort_values(['cell_type', 'sm_name'])\n    df_valid = df_valid.sort_values('sm_name')\n    \n    df_train = convert_to_long_df(df_train)\n    df_valid = convert_to_long_df(df_valid)\n    \n    df_train.to_csv(f'../../results/PerturbNet/splited_data/train_{cell_type}.csv')\n    df_valid.to_csv(f'../../results/PerturbNet/splited_data/valid_{cell_type}.csv')\n```\n\nFor model training, I used the following strategies:\n```python\n    # Setup loss and optimizer\n    criterion = torch.nn.MSELoss()\n    optimizer = AdamW(model.parameters(), lr=args.lr, weight_decay=1e-4)\n    scheduler = ReduceLROnPlateau(optimizer, 'min', min_lr=1e-5)\n```\n\nI trained one model by using each of the cell types as validation data, and for the final submission, I just averaged the predictions.\n\n## What didn't work for me\n\nDuring the competition, I spent a lot of time including the features based on prior knowledge, such as single-cell data, molecular structure embedding, and gene embedding for other large models, such as geneFormer. However, they all didn't work out. So my final model was just based on the features learned from Step 1.\n\n\n## Source\nhttps://github.com/lzj1769/7th_place_solution_Single-Cell-Perturbations\nhttps://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6\nhttps://github.com/jmschrei/avocado",
      "votes": 14
    },
    {
      "id": 2550665,
      "postDate": "2023-12-06T07:39:54.457Z",
      "content": "<p>Congratulations!!! Dr <a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> 🎉🎉🎉!!! You can't imagine how excited I was when I saw your plan!!! It's unbelievable that our thoughts are almost identical.  I think this method of encoding cells, compounds, and genes before making predictions can be more easily promoted, so I also have been spending time on it.</p>\n<p>The difference is that I tried other encoding methods to encode cells, compounds, and genes. In addition, I expanded the dataset by combining the cell annotation method (<a href=\"https://github.com/SongqiZhou/scMMT\" target=\"_blank\">scMMT</a>, a method I wrote that is currently under review) with the 160k PMBCs dataset. However, due to a lack of scale control, the final expanded dataset seems to have had some negative impact on me. At the beginning, it seemed to be effective, but I didn't grasp the scale well, expanded too much, and ultimately had a negative impact.</p>\n<p>I also attempted to encode cell types using the principal components of RNA expression levels for each cell type in a multimodal dataset(As shown in the following figure), but it didn't work well and I didn't find a obvious better encoding method for cell types than onehot encoding.  So I used a combination of onehot and PCA1 methods. I see that you also used the \"sc.pl.embedding\" operation but I didn't, perhaps this is the reason why my cell coding is not ideal. I used to think they contained an equal amount of information, so I didn't try to it. I guess it's the effect of cell count, and I'm still searching for the real reason.</p>\n<p>In short, your method has been of great help to me, but due to my limited skills, I need to take the time to understand your coding methods. I will study your method carefully and hope to have the opportunity to learn and exchange more ideas with you in the future!!!🙏🙏🙏</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8305359%2F9454a952e8cb5186d3c02066cd4f9299%2F_20231206163251.png?generation=1701851598654845&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratulations!!! Dr @zhijianli 🎉🎉🎉!!! You can't imagine how excited I was when I saw your plan!!! It's unbelievable that our thoughts are almost identical.  I think this method of encoding cells, compounds, and genes before making predictions can be more easily promoted, so I also have been spending time on it.\n\nThe difference is that I tried other encoding methods to encode cells, compounds, and genes. In addition, I expanded the dataset by combining the cell annotation method ([scMMT](https://github.com/SongqiZhou/scMMT), a method I wrote that is currently under review) with the 160k PMBCs dataset. However, due to a lack of scale control, the final expanded dataset seems to have had some negative impact on me. At the beginning, it seemed to be effective, but I didn't grasp the scale well, expanded too much, and ultimately had a negative impact.\n\nI also attempted to encode cell types using the principal components of RNA expression levels for each cell type in a multimodal dataset(As shown in the following figure), but it didn't work well and I didn't find a obvious better encoding method for cell types than onehot encoding.  So I used a combination of onehot and PCA1 methods. I see that you also used the \"sc.pl.embedding\" operation but I didn't, perhaps this is the reason why my cell coding is not ideal. I used to think they contained an equal amount of information, so I didn't try to it. I guess it's the effect of cell count, and I'm still searching for the real reason.\n\nIn short, your method has been of great help to me, but due to my limited skills, I need to take the time to understand your coding methods. I will study your method carefully and hope to have the opportunity to learn and exchange more ideas with you in the future!!!🙏🙏🙏\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8305359%2F9454a952e8cb5186d3c02066cd4f9299%2F_20231206163251.png?generation=1701851598654845&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 2552954,
          "postDate": "2023-12-07T22:31:39.457Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/songqizhou\" target=\"_blank\">@songqizhou</a> </p>\n<p>Thanks for your comments, and I am glad that we converged on a similar approach.</p>\n<p>Regarding 'encode cell types using the principal components of RNA expression levels for each cell type in a multimodal dataset', this also didn't work out for me. I tried two different ways, i.e., 1) generate PCA for each single cell and then use average PCs the each cell type, 2) first get pseudo-bulk profiles for each cell type and then do the PCA. It turns out they just didn't boost my model.</p>\n<p>Sorry for the bad organization of my code repository; I am working on it to have a clean version, which should be much easier to understand.</p>\n<p>Finally, good luck with your submission! </p>",
          "rawMarkdown": "Hi @songqizhou \n\nThanks for your comments, and I am glad that we converged on a similar approach.\n\nRegarding 'encode cell types using the principal components of RNA expression levels for each cell type in a multimodal dataset', this also didn't work out for me. I tried two different ways, i.e., 1) generate PCA for each single cell and then use average PCs the each cell type, 2) first get pseudo-bulk profiles for each cell type and then do the PCA. It turns out they just didn't boost my model.\n\nSorry for the bad organization of my code repository; I am working on it to have a clean version, which should be much easier to understand.\n\nFinally, good luck with your submission! \n",
          "votes": 4,
          "replies": [
            {
              "id": 2553457,
              "postDate": "2023-12-08T09:03:10.393Z",
              "content": "<p>Thank you very much, Dr Li! Your code is as beautiful as poetry. Best regards to you!</p>",
              "rawMarkdown": "Thank you very much, Dr Li! Your code is as beautiful as poetry. Best regards to you!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2559137,
      "postDate": "2023-12-12T16:12:44.867Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2558237,
      "postDate": "2023-12-12T02:42:25.130Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2550665,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-12-06T07:39:54.457000",
      "content": "<p>Congratulations!!! Dr <a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> 🎉🎉🎉!!! You can't imagine how excited I was when I saw your plan!!! It's unbelievable that our thoughts are almost identical.  I think this method of encoding cells, compounds, and genes before making predictions can be more easily promoted, so I also have been spending time on it.</p>\n<p>The difference is that I tried other encoding methods to encode cells, compounds, and genes. In addition, I expanded the dataset by combining the cell annotation method (<a href=\"https://github.com/SongqiZhou/scMMT\" target=\"_blank\">scMMT</a>, a method I wrote that is currently under review) with the 160k PMBCs dataset. However, due to a lack of scale control, the final expanded dataset seems to have had some negative impact on me. At the beginning, it seemed to be effective, but I didn't grasp the scale well, expanded too much, and ultimately had a negative impact.</p>\n<p>I also attempted to encode cell types using the principal components of RNA expression levels for each cell type in a multimodal dataset(As shown in the following figure), but it didn't work well and I didn't find a obvious better encoding method for cell types than onehot encoding.  So I used a combination of onehot and PCA1 methods. I see that you also used the \"sc.pl.embedding\" operation but I didn't, perhaps this is the reason why my cell coding is not ideal. I used to think they contained an equal amount of information, so I didn't try to it. I guess it's the effect of cell count, and I'm still searching for the real reason.</p>\n<p>In short, your method has been of great help to me, but due to my limited skills, I need to take the time to understand your coding methods. I will study your method carefully and hope to have the opportunity to learn and exchange more ideas with you in the future!!!🙏🙏🙏</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8305359%2F9454a952e8cb5186d3c02066cd4f9299%2F_20231206163251.png?generation=1701851598654845&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2552954,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2023-12-07T22:31:39.457000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/songqizhou\" target=\"_blank\">@songqizhou</a> </p>\n<p>Thanks for your comments, and I am glad that we converged on a similar approach.</p>\n<p>Regarding 'encode cell types using the principal components of RNA expression levels for each cell type in a multimodal dataset', this also didn't work out for me. I tried two different ways, i.e., 1) generate PCA for each single cell and then use average PCs the each cell type, 2) first get pseudo-bulk profiles for each cell type and then do the PCA. It turns out they just didn't boost my model.</p>\n<p>Sorry for the bad organization of my code repository; I am working on it to have a clean version, which should be much easier to understand.</p>\n<p>Finally, good luck with your submission! </p>",
          "votes": 4,
          "replies": [
            {
              "id": 2553457,
              "author_name": "BarryZhou",
              "author_url": "",
              "post_date": "2023-12-08T09:03:10.393000",
              "content": "<p>Thank you very much, Dr Li! Your code is as beautiful as poetry. Best regards to you!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2559137,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-12T16:12:44.867000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2558237,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-12T02:42:25.130000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2550395": "Thank the organizers for hosting this interesting competition and congrats to the Winners.\nIt was a big surprise for me to see my resolution finally achieve 7th place. \n\n## Context\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n## Overview of the approach\n\nMy approach has two major steps: i) learn embeddings for all cell types, small molecular, and genes,  ii) use embedding as features to train a deep learning model to predict the target while accounting for overfitting. Throughout the competition, I only used a fully connected network with three layers. My best submission is also based on a single FC model.\n\n### Learning embeddings\n\nIn this step, the goal is to learn a specific embedding for each cell type, molecular, and gene. I was strongly inspired by [this paper](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6), where the authors used deep tensor factorization to learn a dense, information-rich representation for cell type, experimental assay, and genomic position.  \n\nBasically, I used the same approach but played around with the model architectures and parameters, such as the number of latent factors, and the number of dimensions of the network, as well as how to combine the features (e.g., concatenate vs. additive). \n\nIn the end, my mode is the following:\n```python\nclass DeepTensorFactorization(torch.nn.Module):\n    def __init__(self, \n                 cell_types, \n                 compounds, \n                 genes, \n                 n_cell_type_factors: int=4, \n                 n_compounds_factors: int=16, \n                 n_gene_factors: int=128,\n                 n_hiddens: int=2048,\n                 dropout: float=0.1):\n        super().__init__()\n        \n        self.cell_types = cell_types\n        self.compounds = compounds\n        self.genes = genes\n        \n        self.n_cell_types = len(cell_types)\n        self.n_compounds = len(compounds)\n        self.n_genes = len(genes)\n        \n        self.n_cell_type_factors = n_cell_type_factors\n        self.n_compounds_factors = n_compounds_factors\n        self.n_gene_factors = n_gene_factors\n        \n        self.cell_type_embedding = torch.nn.Embedding(self.n_cell_types, self.n_cell_type_factors)\n        self.compound_embedding = torch.nn.Embedding(self.n_compounds, self.n_compounds_factors)\n        self.gene_embedding = torch.nn.Embedding(self.n_genes, self.n_gene_factors)\n        \n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n        self.n_factors = n_cell_type_factors + n_compounds_factors + n_gene_factors\n        \n        self.model = nn.Sequential(nn.Linear(self.n_factors, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, 1))\n        \n    def forward(self, cell_type_indices, compound_indices, gene_indices):\n        cell_type_vec = self.cell_type_embedding(cell_type_indices)\n        compound_vec = self.compound_embedding(compound_indices)\n        gene_vec = self.gene_embedding(gene_indices)\n        \n        x = torch.concat([cell_type_vec, compound_vec, gene_vec], dim=1)\n        x = self.model(x)\n        \n        return x\n```\nTo train this model, I used all the data from `de_train.parquet` and converted the table as follows:\n```python\n# convert to long df\ndf = pd.read_parquet('de_train.parquet')\ndf = df.sort_values(['cell_type', 'sm_name'])\ndf = df.drop(['sm_lincs_id', 'SMILES', 'control'], axis=1)\ndf = pd.melt(df, id_vars=['cell_type', 'sm_name'], var_name='gene', value_name='target')\n```\nThe training data looks like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2Ffcf2b04f2255dc9fe2f0382320346feb%2FScreenshot%202023-12-05%20at%2021.02.43.png?generation=1701828192377214&alt=media)\n\nHere, I trained the model with 100 epochs without validation, because I will only the embedding layer for Step 2.\n\nTo check if the model learns meaningful embedding for cell types, molecules, and genes, I also visualized the embedding with UMAP. For example, below is the 2D UMAP of gene embeddings:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F642483%2F59782ef056d394c5deb99376c0af2a05%2FScreenshot%202023-12-05%20at%2021.05.28.png?generation=1701828348397738&alt=media)\n\n\nBy eyeballing, it looks like there are some structures for the genes.\n\n### Predicting target  \n\nOnce I obtained the embeddings, I trained another model to predict the target, and this model was also used to generate the final submission. Again, I used a FC network as follows:\n```python\nclass PerturbNet(torch.nn.Module):\n    def __init__(self, \n                 n_input: int=148,\n                 n_hiddens: int=2048,\n                 dropout: float=0.5):\n        super().__init__()\n        \n        self.n_input = n_input\n        self.n_hiddens = n_hiddens\n        self.dropout = dropout\n\n        self.model = nn.Sequential(nn.Linear(self.n_input, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, self.n_hiddens),\n                                   nn.BatchNorm1d(self.n_hiddens),\n                                   nn.ReLU(),\n                                   nn.Dropout(self.dropout),\n                                   nn.Linear(self.n_hiddens, 1))\n        \n    def forward(self, x):\n        x = self.model(x)\n        \n        return x\n```\n\nTo prevent overfitting, I used each of the cell types as validation data, which means my model was based on 4-fold cross-validation. For the compounds, I followed the notebook here to select the private test compounds for validation:\n\n```python\nfor key, cell_type in cell_type_names.items():\n    print(cell_type)\n    \n    # split data for training and validation, here we used the private test compounds for validation\n    df_train = df[(df['cell_type'] != key) | ~df['sm_lincs_id'].isin(privte_ids)]\n    df_valid = df[(df['cell_type'] == key) & df['sm_lincs_id'].isin(privte_ids)]\n    \n    df_train = df_train.sort_values(['cell_type', 'sm_name'])\n    df_valid = df_valid.sort_values('sm_name')\n    \n    df_train = convert_to_long_df(df_train)\n    df_valid = convert_to_long_df(df_valid)\n    \n    df_train.to_csv(f'../../results/PerturbNet/splited_data/train_{cell_type}.csv')\n    df_valid.to_csv(f'../../results/PerturbNet/splited_data/valid_{cell_type}.csv')\n```\n\nFor model training, I used the following strategies:\n```python\n    # Setup loss and optimizer\n    criterion = torch.nn.MSELoss()\n    optimizer = AdamW(model.parameters(), lr=args.lr, weight_decay=1e-4)\n    scheduler = ReduceLROnPlateau(optimizer, 'min', min_lr=1e-5)\n```\n\nI trained one model by using each of the cell types as validation data, and for the final submission, I just averaged the predictions.\n\n## What didn't work for me\n\nDuring the competition, I spent a lot of time including the features based on prior knowledge, such as single-cell data, molecular structure embedding, and gene embedding for other large models, such as geneFormer. However, they all didn't work out. So my final model was just based on the features learned from Step 1.\n\n\n## Source\nhttps://github.com/lzj1769/7th_place_solution_Single-Cell-Perturbations\nhttps://genomebiology.biomedcentral.com/articles/10.1186/s13059-020-01977-6\nhttps://github.com/jmschrei/avocado",
    "2550665": "Congratulations!!! Dr @zhijianli 🎉🎉🎉!!! You can't imagine how excited I was when I saw your plan!!! It's unbelievable that our thoughts are almost identical.  I think this method of encoding cells, compounds, and genes before making predictions can be more easily promoted, so I also have been spending time on it.\n\nThe difference is that I tried other encoding methods to encode cells, compounds, and genes. In addition, I expanded the dataset by combining the cell annotation method ([scMMT](https://github.com/SongqiZhou/scMMT), a method I wrote that is currently under review) with the 160k PMBCs dataset. However, due to a lack of scale control, the final expanded dataset seems to have had some negative impact on me. At the beginning, it seemed to be effective, but I didn't grasp the scale well, expanded too much, and ultimately had a negative impact.\n\nI also attempted to encode cell types using the principal components of RNA expression levels for each cell type in a multimodal dataset(As shown in the following figure), but it didn't work well and I didn't find a obvious better encoding method for cell types than onehot encoding.  So I used a combination of onehot and PCA1 methods. I see that you also used the \"sc.pl.embedding\" operation but I didn't, perhaps this is the reason why my cell coding is not ideal. I used to think they contained an equal amount of information, so I didn't try to it. I guess it's the effect of cell count, and I'm still searching for the real reason.\n\nIn short, your method has been of great help to me, but due to my limited skills, I need to take the time to understand your coding methods. I will study your method carefully and hope to have the opportunity to learn and exchange more ideas with you in the future!!!🙏🙏🙏\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8305359%2F9454a952e8cb5186d3c02066cd4f9299%2F_20231206163251.png?generation=1701851598654845&alt=media)",
    "2559137": "",
    "2558237": ""
  }
}