{
  "id": 447256,
  "title": "Did you use any biological prior knowledge and how it works?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/447256",
  "author_name": "Zhijian Li",
  "post_date": "2023-10-15T03:33:41.163000",
  "votes": 16,
  "comment_count": 23,
  "views": 0,
  "content": "<p>So far, my model is solely based on training data as a regression method. Currently, the local CV is 0.708, and the public LB is 0.567. </p>\n<p>I am wondering if anyone find useful prior knowledge to boost the performance.</p>\n<p>Any ideas or thoughts are welcome.</p>",
  "messages": [
    {
      "id": 2482426,
      "postDate": "2023-10-15T03:33:41.163Z",
      "content": "<p>So far, my model is solely based on training data as a regression method. Currently, the local CV is 0.708, and the public LB is 0.567. </p>\n<p>I am wondering if anyone find useful prior knowledge to boost the performance.</p>\n<p>Any ideas or thoughts are welcome.</p>",
      "rawMarkdown": "So far, my model is solely based on training data as a regression method. Currently, the local CV is 0.708, and the public LB is 0.567. \n\nI am wondering if anyone find useful prior knowledge to boost the performance.\n\nAny ideas or thoughts are welcome.\n\n\n",
      "votes": 16
    },
    {
      "id": 2483469,
      "postDate": "2023-10-15T17:43:22.700Z",
      "content": "<p>I can relate to that. Currently, my local CV stands at 0.598, while public LB  is 0.571. I made an attempt to integrate chemical compound data, but unfortunately, it didn't yield favorable results. Now, I'm considering experimenting with ATAC-seq data. While I'm not entirely sure how to incorporate it into my model, I'm eager to give it a try and see how it performs.</p>",
      "rawMarkdown": "I can relate to that. Currently, my local CV stands at 0.598, while public LB  is 0.571. I made an attempt to integrate chemical compound data, but unfortunately, it didn't yield favorable results. Now, I'm considering experimenting with ATAC-seq data. While I'm not entirely sure how to incorporate it into my model, I'm eager to give it a try and see how it performs.",
      "votes": 5,
      "replies": [
        {
          "id": 2485983,
          "postDate": "2023-10-17T15:37:45.070Z",
          "content": "<p>My feeling is that the chemical side would be more informative than the biological side, given the huge difference in sequence lengths (chemical molecule and gene).</p>",
          "rawMarkdown": "My feeling is that the chemical side would be more informative than the biological side, given the huge difference in sequence lengths (chemical molecule and gene).",
          "replies": [
            {
              "id": 2486007,
              "postDate": "2023-10-17T16:10:51.430Z",
              "content": "<p>Actually, I tried to include the features generated by ChemBERT for compounds, but I had no luck and the results became worse. </p>\n<p>Maybe I should try other approaches to extract features for molecular.</p>",
              "rawMarkdown": "Actually, I tried to include the features generated by ChemBERT for compounds, but I had no luck and the results became worse. \n\nMaybe I should try other approaches to extract features for molecular.\n\n"
            },
            {
              "id": 2486010,
              "postDate": "2023-10-17T16:14:14.330Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2486012,
              "postDate": "2023-10-17T16:19:10.763Z",
              "content": "<p>Agree with you but i believe ATAC seq can provide some boost to the model. But its really hard to incorporate it into the model given its long length. </p>",
              "rawMarkdown": "Agree with you but i believe ATAC seq can provide some boost to the model. But its really hard to incorporate it into the model given its long length. ",
              "votes": 1
            },
            {
              "id": 2487334,
              "postDate": "2023-10-18T14:21:49.740Z",
              "content": "<p>As for me - not very clear what can the way to use - ATAC-seq data, since we do not have it for the test part - so cannot make features directly, only may be some aggregates by cell type , or by compound , but hardly it is not yet in the data already </p>",
              "rawMarkdown": "As for me - not very clear what can the way to use - ATAC-seq data, since we do not have it for the test part - so cannot make features directly, only may be some aggregates by cell type , or by compound , but hardly it is not yet in the data already "
            },
            {
              "id": 2487356,
              "postDate": "2023-10-18T14:30:24.743Z",
              "content": "<p>an idea that came to me was generating a joint embedding for RNA-seq and ATAC-seq. might use those winner models from the 2021 single-cell competition.</p>",
              "rawMarkdown": "an idea that came to me was generating a joint embedding for RNA-seq and ATAC-seq. might use those winner models from the 2021 single-cell competition.",
              "votes": 1
            },
            {
              "id": 2487412,
              "postDate": "2023-10-18T15:14:21.457Z",
              "content": "<p>Nice idea, but still how to attach that to modeling - we do NOT have neither RNA,neither ATAC for the test part - make some aggregates by cell-type, and compound ? A kind of target encoded features based on these embeddings aggregates </p>",
              "rawMarkdown": "Nice idea, but still how to attach that to modeling - we do NOT have neither RNA,neither ATAC for the test part - make some aggregates by cell-type, and compound ? A kind of target encoded features based on these embeddings aggregates \n"
            },
            {
              "id": 2487460,
              "postDate": "2023-10-18T15:49:54.083Z",
              "content": "<p>I used chemBERT to generate a 512 dimension embedding for each molecular and then used the top 30 PCs as input. The performance decreased.</p>",
              "rawMarkdown": "I used chemBERT to generate a 512 dimension embedding for each molecular and then used the top 30 PCs as input. The performance decreased.",
              "votes": 2
            },
            {
              "id": 2489944,
              "postDate": "2023-10-20T10:20:37.067Z",
              "content": "<p>a viable solution for me is to average the latent space for each cell type and use it as our cell embedding. I haven't tried that yet.</p>",
              "rawMarkdown": "a viable solution for me is to average the latent space for each cell type and use it as our cell embedding. I haven't tried that yet.",
              "votes": 1
            },
            {
              "id": 2491250,
              "postDate": "2023-10-21T13:06:39.750Z",
              "content": "<p>The ATC-seq data could help indicate each cell type's epigenomic and, therefore, which genes are more likely to be expressed for each cell type. I am unsure how to best add it to the dataset, maybe as embedding representing the cell types. </p>",
              "rawMarkdown": "The ATC-seq data could help indicate each cell type's epigenomic and, therefore, which genes are more likely to be expressed for each cell type. I am unsure how to best add it to the dataset, maybe as embedding representing the cell types. ",
              "votes": 2
            }
          ]
        },
        {
          "id": 2493953,
          "postDate": "2023-10-23T16:44:17.450Z",
          "content": "<p>Dear Kishan, thank you for your great contribution of notebook Nueral_Network_Regression and NLP_regression! Recently I am learning from your notebooks to construct models based on your 152-dimensional one hot embedding in Nueral_Network_Regression notebook. But no matter how I change my model, I can just get a 0.6xx LB score. I am depressed about my result. I am glad to see you reach a 0.571 LB score, and I want to ask if you are using the 152-dimensional input or the NLP input to reach such a score? I hope you can answer me to help me find the direction to change my input or improve my models.🙋‍♂️</p>",
          "rawMarkdown": "Dear Kishan, thank you for your great contribution of notebook Nueral_Network_Regression and NLP_regression! Recently I am learning from your notebooks to construct models based on your 152-dimensional one hot embedding in Nueral_Network_Regression notebook. But no matter how I change my model, I can just get a 0.6xx LB score. I am depressed about my result. I am glad to see you reach a 0.571 LB score, and I want to ask if you are using the 152-dimensional input or the NLP input to reach such a score? I hope you can answer me to help me find the direction to change my input or improve my models.🙋‍♂️",
          "votes": 1,
          "replies": [
            {
              "id": 2496218,
              "postDate": "2023-10-23T20:29:13.087Z",
              "content": "<p>Thank you for your kind words, <a href=\"https://www.kaggle.com/lc23333\" target=\"_blank\">@lc23333</a>. Don't overthink; instead, try exploring different approaches. I encourage you to experiment  with more complex models or ensembling model for the 152-dimensional input.</p>\n<p>However, I want to stress this that solely focusing on a high leaderboard score might lead to overfitting on the public test data. I still think relying solely on pure regression without incorporating chemical or biological data may not be the best approach for solving this problem and it has very high chances of overfitting on public test and the goal is to generalize. So, don't compete for leaderboard and try solving the problem and in the process learning from it . </p>\n<p>Also, don't get caught up in achieving better score solely from one approach. Keep learning and experimenting! </p>\n<p>P.s. It's no secret that anyone can achieve very good score on public test using neural network complex models from only 152 dimensional input. I also did the same. But the key is neural nets are very good at learning train data consequently, overfitting! </p>",
              "rawMarkdown": "Thank you for your kind words, @lc23333. Don't overthink; instead, try exploring different approaches. I encourage you to experiment  with more complex models or ensembling model for the 152-dimensional input.\n\nHowever, I want to stress this that solely focusing on a high leaderboard score might lead to overfitting on the public test data. I still think relying solely on pure regression without incorporating chemical or biological data may not be the best approach for solving this problem and it has very high chances of overfitting on public test and the goal is to generalize. So, don't compete for leaderboard and try solving the problem and in the process learning from it . \n \nAlso, don't get caught up in achieving better score solely from one approach. Keep learning and experimenting! \n\nP.s. It's no secret that anyone can achieve very good score on public test using neural network complex models from only 152 dimensional input. I also did the same. But the key is neural nets are very good at learning train data consequently, overfitting! ",
              "votes": 3
            },
            {
              "id": 2496349,
              "postDate": "2023-10-24T01:57:05.840Z",
              "content": "<p>Thank you for your generous share! I understand your concern about the overfitting. I will try to solve the task instead of achieving high scores. Sincerely hope we can have a good journey in the project.😆</p>",
              "rawMarkdown": "Thank you for your generous share! I understand your concern about the overfitting. I will try to solve the task instead of achieving high scores. Sincerely hope we can have a good journey in the project.😆",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2482474,
      "postDate": "2023-10-15T04:32:15.563Z",
      "content": "<p>Nothing works for me so far… I think we are approaching the limit for regression without successfully using biological prior</p>",
      "rawMarkdown": "Nothing works for me so far... I think we are approaching the limit for regression without successfully using biological prior",
      "votes": 6,
      "replies": [
        {
          "id": 2487465,
          "postDate": "2023-10-18T15:51:11.570Z",
          "content": "<p>Same feeling for me.</p>",
          "rawMarkdown": "Same feeling for me.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2490630,
      "postDate": "2023-10-20T21:17:54.473Z",
      "content": "<p>I tried ChemBERTa embeddings like others, and the performance decreased for me, too. However, some better models on the leaderboard might have been overfitting. This method may still lead to improved performance on the global dataset. <br>\nAs for the biology prior, I feel the best way would be to integrate the training data with datasets of chemical toxicity or the LINCS dataset, as Laura Sisson suggested.<br>\nAnother strategy that could be interesting is to use gene ontology analysis to determine if the chemical treatments impact specific biological pathways and to add that information to the training set. I have started that work; we will see :). </p>",
      "rawMarkdown": "I tried ChemBERTa embeddings like others, and the performance decreased for me, too. However, some better models on the leaderboard might have been overfitting. This method may still lead to improved performance on the global dataset. \nAs for the biology prior, I feel the best way would be to integrate the training data with datasets of chemical toxicity or the LINCS dataset, as Laura Sisson suggested.\nAnother strategy that could be interesting is to use gene ontology analysis to determine if the chemical treatments impact specific biological pathways and to add that information to the training set. I have started that work; we will see :). ",
      "votes": 3
    },
    {
      "id": 2486963,
      "postDate": "2023-10-18T09:03:49.597Z",
      "content": "<p>I trained a GNN similar to GEARS using a gene coexpression network and gene ontology, but neither worked fine. The problem might lie in the number of isolated nodes. I generated them based on the gene expression in adata_train.parquet. this is a bit annoying.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F5670e1fc81872392259a4b5f803feba4%2F2023-10-18%2017.00.32.png?generation=1697619655280220&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I trained a GNN similar to GEARS using a gene coexpression network and gene ontology, but neither worked fine. The problem might lie in the number of isolated nodes. I generated them based on the gene expression in adata_train.parquet. this is a bit annoying.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F5670e1fc81872392259a4b5f803feba4%2F2023-10-18%2017.00.32.png?generation=1697619655280220&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 2486964,
          "postDate": "2023-10-18T09:04:53.580Z",
          "content": "<p>or maybe GNN is too hard to fit on such a small dataset?</p>",
          "rawMarkdown": "or maybe GNN is too hard to fit on such a small dataset?",
          "votes": 1,
          "replies": [
            {
              "id": 2493269,
              "postDate": "2023-10-23T09:08:03.873Z",
              "content": "<p>so many GNNs were used in competitions recently.</p>",
              "rawMarkdown": "so many GNNs were used in competitions recently.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2487206,
          "postDate": "2023-10-18T12:46:06.023Z",
          "content": "<blockquote>\n  <p>I trained a GNN similar to GEARS using a gene coexpression network and gene ontology, but neither worked fine. The problem might lie in the number of isolated nodes. I generated them based on the gene expression in adata_train.parquet. this is a bit annoying.</p>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F5670e1fc81872392259a4b5f803feba4%2F2023-10-18%2017.00.32.png?generation=1697619655280220&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<p>It is probably something else that break your model.</p>\n<p>Don't know about GEARS, but I also use a GNN. My understanding is that you don't even need any edges (it just reduce to a regular MLP when without edges) for the model to work reasonably well</p>",
          "rawMarkdown": "> I trained a GNN similar to GEARS using a gene coexpression network and gene ontology, but neither worked fine. The problem might lie in the number of isolated nodes. I generated them based on the gene expression in adata_train.parquet. this is a bit annoying.\n> \n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F5670e1fc81872392259a4b5f803feba4%2F2023-10-18%2017.00.32.png?generation=1697619655280220&alt=media)\n\nIt is probably something else that break your model.\n\nDon't know about GEARS, but I also use a GNN. My understanding is that you don't even need any edges (it just reduce to a regular MLP when without edges) for the model to work reasonably well\n\n",
          "replies": [
            {
              "id": 2487362,
              "postDate": "2023-10-18T14:31:44.330Z",
              "content": "<p>let me dive into it. I don't know much about GNNs, so probably have constructed it in a bad way…</p>",
              "rawMarkdown": "let me dive into it. I don't know much about GNNs, so probably have constructed it in a bad way..."
            }
          ]
        }
      ]
    },
    {
      "id": 2486011,
      "postDate": "2023-10-17T16:16:05.953Z",
      "content": "<p>Adding biological prior knowledge, such as gene interactions or pathway information, can significantly enhance your model's performance in genomics tasks. Have you explored integrating such knowledge to further improve your results? Upvote if you agree! 🧬🔬📈</p>",
      "rawMarkdown": "Adding biological prior knowledge, such as gene interactions or pathway information, can significantly enhance your model's performance in genomics tasks. Have you explored integrating such knowledge to further improve your results? Upvote if you agree! 🧬🔬📈"
    }
  ],
  "comments": [
    {
      "id": 2483469,
      "author_name": "Kishan Vavdara",
      "author_url": "",
      "post_date": "2023-10-15T17:43:22.700000",
      "content": "<p>I can relate to that. Currently, my local CV stands at 0.598, while public LB  is 0.571. I made an attempt to integrate chemical compound data, but unfortunately, it didn't yield favorable results. Now, I'm considering experimenting with ATAC-seq data. While I'm not entirely sure how to incorporate it into my model, I'm eager to give it a try and see how it performs.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2485983,
          "author_name": "Zekai Li",
          "author_url": "",
          "post_date": "2023-10-17T15:37:45.070000",
          "content": "<p>My feeling is that the chemical side would be more informative than the biological side, given the huge difference in sequence lengths (chemical molecule and gene).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2486007,
              "author_name": "Zhijian Li",
              "author_url": "",
              "post_date": "2023-10-17T16:10:51.430000",
              "content": "<p>Actually, I tried to include the features generated by ChemBERT for compounds, but I had no luck and the results became worse. </p>\n<p>Maybe I should try other approaches to extract features for molecular.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2486010,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-17T16:14:14.330000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2486012,
              "author_name": "Kishan Vavdara",
              "author_url": "",
              "post_date": "2023-10-17T16:19:10.763000",
              "content": "<p>Agree with you but i believe ATAC seq can provide some boost to the model. But its really hard to incorporate it into the model given its long length. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2487334,
              "author_name": "Alexander Chervov",
              "author_url": "",
              "post_date": "2023-10-18T14:21:49.740000",
              "content": "<p>As for me - not very clear what can the way to use - ATAC-seq data, since we do not have it for the test part - so cannot make features directly, only may be some aggregates by cell type , or by compound , but hardly it is not yet in the data already </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2487356,
              "author_name": "Daniel Shao",
              "author_url": "",
              "post_date": "2023-10-18T14:30:24.743000",
              "content": "<p>an idea that came to me was generating a joint embedding for RNA-seq and ATAC-seq. might use those winner models from the 2021 single-cell competition.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2487412,
              "author_name": "Alexander Chervov",
              "author_url": "",
              "post_date": "2023-10-18T15:14:21.457000",
              "content": "<p>Nice idea, but still how to attach that to modeling - we do NOT have neither RNA,neither ATAC for the test part - make some aggregates by cell-type, and compound ? A kind of target encoded features based on these embeddings aggregates </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2487460,
              "author_name": "Zhijian Li",
              "author_url": "",
              "post_date": "2023-10-18T15:49:54.083000",
              "content": "<p>I used chemBERT to generate a 512 dimension embedding for each molecular and then used the top 30 PCs as input. The performance decreased.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2489944,
              "author_name": "Daniel Shao",
              "author_url": "",
              "post_date": "2023-10-20T10:20:37.067000",
              "content": "<p>a viable solution for me is to average the latent space for each cell type and use it as our cell embedding. I haven't tried that yet.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2491250,
              "author_name": "Will",
              "author_url": "",
              "post_date": "2023-10-21T13:06:39.750000",
              "content": "<p>The ATC-seq data could help indicate each cell type's epigenomic and, therefore, which genes are more likely to be expressed for each cell type. I am unsure how to best add it to the dataset, maybe as embedding representing the cell types. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2493953,
          "author_name": "Edward",
          "author_url": "",
          "post_date": "2023-10-23T16:44:17.450000",
          "content": "<p>Dear Kishan, thank you for your great contribution of notebook Nueral_Network_Regression and NLP_regression! Recently I am learning from your notebooks to construct models based on your 152-dimensional one hot embedding in Nueral_Network_Regression notebook. But no matter how I change my model, I can just get a 0.6xx LB score. I am depressed about my result. I am glad to see you reach a 0.571 LB score, and I want to ask if you are using the 152-dimensional input or the NLP input to reach such a score? I hope you can answer me to help me find the direction to change my input or improve my models.🙋‍♂️</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2496218,
              "author_name": "Kishan Vavdara",
              "author_url": "",
              "post_date": "2023-10-23T20:29:13.087000",
              "content": "<p>Thank you for your kind words, <a href=\"https://www.kaggle.com/lc23333\" target=\"_blank\">@lc23333</a>. Don't overthink; instead, try exploring different approaches. I encourage you to experiment  with more complex models or ensembling model for the 152-dimensional input.</p>\n<p>However, I want to stress this that solely focusing on a high leaderboard score might lead to overfitting on the public test data. I still think relying solely on pure regression without incorporating chemical or biological data may not be the best approach for solving this problem and it has very high chances of overfitting on public test and the goal is to generalize. So, don't compete for leaderboard and try solving the problem and in the process learning from it . </p>\n<p>Also, don't get caught up in achieving better score solely from one approach. Keep learning and experimenting! </p>\n<p>P.s. It's no secret that anyone can achieve very good score on public test using neural network complex models from only 152 dimensional input. I also did the same. But the key is neural nets are very good at learning train data consequently, overfitting! </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2496349,
              "author_name": "Edward",
              "author_url": "",
              "post_date": "2023-10-24T01:57:05.840000",
              "content": "<p>Thank you for your generous share! I understand your concern about the overfitting. I will try to solve the task instead of achieving high scores. Sincerely hope we can have a good journey in the project.😆</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2482474,
      "author_name": "qihuaz",
      "author_url": "",
      "post_date": "2023-10-15T04:32:15.563000",
      "content": "<p>Nothing works for me so far… I think we are approaching the limit for regression without successfully using biological prior</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2487465,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2023-10-18T15:51:11.570000",
          "content": "<p>Same feeling for me.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2490630,
      "author_name": "Will",
      "author_url": "",
      "post_date": "2023-10-20T21:17:54.473000",
      "content": "<p>I tried ChemBERTa embeddings like others, and the performance decreased for me, too. However, some better models on the leaderboard might have been overfitting. This method may still lead to improved performance on the global dataset. <br>\nAs for the biology prior, I feel the best way would be to integrate the training data with datasets of chemical toxicity or the LINCS dataset, as Laura Sisson suggested.<br>\nAnother strategy that could be interesting is to use gene ontology analysis to determine if the chemical treatments impact specific biological pathways and to add that information to the training set. I have started that work; we will see :). </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2486963,
      "author_name": "Daniel Shao",
      "author_url": "",
      "post_date": "2023-10-18T09:03:49.597000",
      "content": "<p>I trained a GNN similar to GEARS using a gene coexpression network and gene ontology, but neither worked fine. The problem might lie in the number of isolated nodes. I generated them based on the gene expression in adata_train.parquet. this is a bit annoying.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F5670e1fc81872392259a4b5f803feba4%2F2023-10-18%2017.00.32.png?generation=1697619655280220&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2486964,
          "author_name": "Daniel Shao",
          "author_url": "",
          "post_date": "2023-10-18T09:04:53.580000",
          "content": "<p>or maybe GNN is too hard to fit on such a small dataset?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2493269,
              "author_name": "Lingxi",
              "author_url": "",
              "post_date": "2023-10-23T09:08:03.873000",
              "content": "<p>so many GNNs were used in competitions recently.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2487206,
          "author_name": "qihuaz",
          "author_url": "",
          "post_date": "2023-10-18T12:46:06.023000",
          "content": "<blockquote>\n  <p>I trained a GNN similar to GEARS using a gene coexpression network and gene ontology, but neither worked fine. The problem might lie in the number of isolated nodes. I generated them based on the gene expression in adata_train.parquet. this is a bit annoying.</p>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F5670e1fc81872392259a4b5f803feba4%2F2023-10-18%2017.00.32.png?generation=1697619655280220&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<p>It is probably something else that break your model.</p>\n<p>Don't know about GEARS, but I also use a GNN. My understanding is that you don't even need any edges (it just reduce to a regular MLP when without edges) for the model to work reasonably well</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2487362,
              "author_name": "Daniel Shao",
              "author_url": "",
              "post_date": "2023-10-18T14:31:44.330000",
              "content": "<p>let me dive into it. I don't know much about GNNs, so probably have constructed it in a bad way…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2486011,
      "author_name": "Bhavesh Padharia",
      "author_url": "",
      "post_date": "2023-10-17T16:16:05.953000",
      "content": "<p>Adding biological prior knowledge, such as gene interactions or pathway information, can significantly enhance your model's performance in genomics tasks. Have you explored integrating such knowledge to further improve your results? Upvote if you agree! 🧬🔬📈</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2482426": "So far, my model is solely based on training data as a regression method. Currently, the local CV is 0.708, and the public LB is 0.567. \n\nI am wondering if anyone find useful prior knowledge to boost the performance.\n\nAny ideas or thoughts are welcome.\n\n\n",
    "2483469": "I can relate to that. Currently, my local CV stands at 0.598, while public LB  is 0.571. I made an attempt to integrate chemical compound data, but unfortunately, it didn't yield favorable results. Now, I'm considering experimenting with ATAC-seq data. While I'm not entirely sure how to incorporate it into my model, I'm eager to give it a try and see how it performs.",
    "2482474": "Nothing works for me so far... I think we are approaching the limit for regression without successfully using biological prior",
    "2490630": "I tried ChemBERTa embeddings like others, and the performance decreased for me, too. However, some better models on the leaderboard might have been overfitting. This method may still lead to improved performance on the global dataset. \nAs for the biology prior, I feel the best way would be to integrate the training data with datasets of chemical toxicity or the LINCS dataset, as Laura Sisson suggested.\nAnother strategy that could be interesting is to use gene ontology analysis to determine if the chemical treatments impact specific biological pathways and to add that information to the training set. I have started that work; we will see :). ",
    "2486963": "I trained a GNN similar to GEARS using a gene coexpression network and gene ontology, but neither worked fine. The problem might lie in the number of isolated nodes. I generated them based on the gene expression in adata_train.parquet. this is a bit annoying.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7863678%2F5670e1fc81872392259a4b5f803feba4%2F2023-10-18%2017.00.32.png?generation=1697619655280220&alt=media)",
    "2486011": "Adding biological prior knowledge, such as gene interactions or pathway information, can significantly enhance your model's performance in genomics tasks. Have you explored integrating such knowledge to further improve your results? Upvote if you agree! 🧬🔬📈"
  }
}