{
  "id": 364669,
  "title": "Are there any good bio-driven methods which can lead great improvement?",
  "url": "/competitions/open-problems-multimodal/discussion/364669",
  "author_name": "",
  "post_date": "2022-11-07T19:20:50.629523200Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi all, I am thinking about how to reach 0.814. I have tried bio-driven idea including:</p>\n<ol>\n<li><p>For cite-seq, use imputed significant protein labeled genes to train, and I have tried MAGIC and Deep Impute.</p></li>\n<li><p>For cite-seq, use cells to construct cell-cell sim network and use GNN to predict protein.</p></li>\n<li><p>For cite-seq, use correlation to select high correlated gene-protein relation, and then predict the protein one by one.</p></li>\n<li><p>For cite-seq, encode gender, time and cell type information into the input file, and then perform the training process and prediction process.</p></li>\n<li><p>For multiome, use gene activity score rather than original peak information.</p></li>\n</ol>\n<p>All of them cannot improve my current score, so I am very confused about incuding extra information to this competition. I am very happy to learn some bio-driven ideas. </p>",
  "messages": [
    {
      "id": "2020777",
      "postDate": "11/07/2022 19:20:50",
      "content": "<p>Hi all, I am thinking about how to reach 0.814. I have tried bio-driven idea including:</p>\n<ol>\n<li><p>For cite-seq, use imputed significant protein labeled genes to train, and I have tried MAGIC and Deep Impute.</p></li>\n<li><p>For cite-seq, use cells to construct cell-cell sim network and use GNN to predict protein.</p></li>\n<li><p>For cite-seq, use correlation to select high correlated gene-protein relation, and then predict the protein one by one.</p></li>\n<li><p>For cite-seq, encode gender, time and cell type information into the input file, and then perform the training process and prediction process.</p></li>\n<li><p>For multiome, use gene activity score rather than original peak information.</p></li>\n</ol>\n<p>All of them cannot improve my current score, so I am very confused about incuding extra information to this competition. I am very happy to learn some bio-driven ideas. </p>",
      "rawMarkdown": "Hi all, I am thinking about how to reach 0.814. I have tried bio-driven idea including:\n\n1. For cite-seq, use imputed significant protein labeled genes to train, and I have tried MAGIC and Deep Impute.\n\n2. For cite-seq, use cells to construct cell-cell sim network and use GNN to predict protein.\n\n3. For cite-seq, use correlation to select high correlated gene-protein relation, and then predict the protein one by one.\n\n4. For cite-seq, encode gender, time and cell type information into the input file, and then perform the training process and prediction process.\n\n5. For multiome, use gene activity score rather than original peak information.\n\nAll of them cannot improve my current score, so I am very confused about incuding extra information to this competition. I am very happy to learn some bio-driven ideas.",
      "votes": null
    },
    {
      "id": "2020896",
      "postDate": "11/07/2022 20:49:43",
      "content": "<p>Zhou, Z., Ye, C., Wang, J. et al. Surface protein imputation from single cell transcriptomes by deep neural networks. Nat Commun 11, 651 (2020). <br>\nThis is a paper for cite-seq prediction, it uses R package saver-x to denoise the count matrix and then use multi-branch NN to predict. I have not tried this, just as a share. </p>",
      "rawMarkdown": "Zhou, Z., Ye, C., Wang, J. et al. Surface protein imputation from single cell transcriptomes by deep neural networks. Nat Commun 11, 651 (2020). \nThis is a paper for cite-seq prediction, it uses R package saver-x to denoise the count matrix and then use multi-branch NN to predict. I have not tried this, just as a share.",
      "votes": null
    },
    {
      "id": "2020898",
      "postDate": "11/07/2022 20:51:31",
      "content": "<p>What does 1. here means? Would you like to explain it more? </p>",
      "rawMarkdown": "What does 1. here means? Would you like to explain it more?",
      "votes": null
    },
    {
      "id": "2020949",
      "postDate": "11/07/2022 21:45:59",
      "content": "<p>Thanks for your sharing. However, according to their results, the cor even cannot exceed the public methods. </p>",
      "rawMarkdown": "Thanks for your sharing. However, according to their results, the cor even cannot exceed the public methods.",
      "votes": null
    },
    {
      "id": "2020951",
      "postDate": "11/07/2022 21:46:55",
      "content": "<p>You can search \"single cell gene expression imputation\" for more information. Here is one paper: <a href=\"https://academic.oup.com/nar/article-abstract/50/9/4877/6582166\" target=\"_blank\">https://academic.oup.com/nar/article-abstract/50/9/4877/6582166</a></p>",
      "rawMarkdown": "You can search \"single cell gene expression imputation\" for more information. Here is one paper: https://academic.oup.com/nar/article-abstract/50/9/4877/6582166",
      "votes": null
    },
    {
      "id": "2021105",
      "postDate": "11/08/2022 02:25:26",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F6ff77898959d04bbd1f690a3bcff4df9%2Fcorrelation.png?generation=1667874310402899&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F48e8e073775313e36817686679fd5976%2Fcorrelation%20of%20each%20marker.png?generation=1667874324356259&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F6ff77898959d04bbd1f690a3bcff4df9%2Fcorrelation.png?generation=1667874310402899&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F48e8e073775313e36817686679fd5976%2Fcorrelation%20of%20each%20marker.png?generation=1667874324356259&alt=media)",
      "votes": null
    },
    {
      "id": "2021106",
      "postDate": "11/08/2022 02:26:13",
      "content": "<p>Within dataset correlation is pretty high, it seems to me though. </p>",
      "rawMarkdown": "Within dataset correlation is pretty high, it seems to me though.",
      "votes": null
    },
    {
      "id": "2021233",
      "postDate": "11/08/2022 05:38:08",
      "content": "<p>Single Cell Manifold preserving feature selection to select important features, to be kept before dimensionality reduction. <br>\n<a href=\"https://scmer.readthedocs.io/en/latest/modules.html#scmer.UmapL1\" target=\"_blank\">https://scmer.readthedocs.io/en/latest/modules.html#scmer.UmapL1</a></p>",
      "rawMarkdown": "Single Cell Manifold preserving feature selection to select important features, to be kept before dimensionality reduction. \nhttps://scmer.readthedocs.io/en/latest/modules.html#scmer.UmapL1",
      "votes": null
    },
    {
      "id": "2021234",
      "postDate": "11/08/2022 05:39:35",
      "content": "<p>also self-assembling manifolds to select seq genes, and stochastic gates to find features that directly improve score  when regressed on labels; </p>\n<p>my teammate and I have the relevant features selected, but we don't have time to fine-tune the prediction network, so we don't know the best score we can get; </p>",
      "rawMarkdown": "also self-assembling manifolds to select seq genes, and stochastic gates to find features that directly improve score  when regressed on labels; \n\nmy teammate and I have the relevant features selected, but we don't have time to fine-tune the prediction network, so we don't know the best score we can get;",
      "votes": null
    },
    {
      "id": "2021749",
      "postDate": "11/08/2022 13:03:33",
      "content": "<p><a href=\"https://genomebiology.biomedcentral.com/counter/pdf/10.1186/s13059-019-1898-6.pdf\" target=\"_blank\">https://genomebiology.biomedcentral.com/counter/pdf/10.1186/s13059-019-1898-6.pdf</a> Hi, I think according to this paper, at least for scRNA-seq, FA and PCA are top methods to handle it, so I do not think UMAPs could be a good choice, and normally UMAPs need PCs as initial points I think.</p>",
      "rawMarkdown": "https://genomebiology.biomedcentral.com/counter/pdf/10.1186/s13059-019-1898-6.pdf Hi, I think according to this paper, at least for scRNA-seq, FA and PCA are top methods to handle it, so I do not think UMAPs could be a good choice, and normally UMAPs need PCs as initial points I think.",
      "votes": null
    },
    {
      "id": "2021763",
      "postDate": "11/08/2022 13:21:21",
      "content": "<p>I stopped at 0.812 with single model.</p>",
      "rawMarkdown": "I stopped at 0.812 with single model.",
      "votes": null
    },
    {
      "id": "2021869",
      "postDate": "11/08/2022 14:30:26",
      "content": "<p>Yes it seems UMAP isn't a good choice, because UMAP and tsne both focus on preserving local structure when choosing its lower dimensional embedding, whereas PCA preserves global structure. Global structure is more important to maximize correlation score for this competition.</p>\n<p>However, scmer has an advantage in that it can act like supervised feature selection, Scmer minimizes the KL divergence between the latent embedding's cell-to-cell matrix and that of the original matrix, but it can also be set to minimize the divergence between the latent embedding (of the CITE-seq data) and the target protein manifold.</p>",
      "rawMarkdown": "Yes it seems UMAP isn't a good choice, because UMAP and tsne both focus on preserving local structure when choosing its lower dimensional embedding, whereas PCA preserves global structure. Global structure is more important to maximize correlation score for this competition.\n\nHowever, scmer has an advantage in that it can act like supervised feature selection, Scmer minimizes the KL divergence between the latent embedding's cell-to-cell matrix and that of the original matrix, but it can also be set to minimize the divergence between the latent embedding (of the CITE-seq data) and the target protein manifold.",
      "votes": null
    },
    {
      "id": "2029090",
      "postDate": "11/14/2022 12:36:09",
      "content": "<p>it's awesome</p>",
      "rawMarkdown": "it's awesome",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2020896,
      "author_name": "jinyang18",
      "author_url": "",
      "post_date": "11/07/2022 20:49:43",
      "content": "<p>Zhou, Z., Ye, C., Wang, J. et al. Surface protein imputation from single cell transcriptomes by deep neural networks. Nat Commun 11, 651 (2020). <br>\nThis is a paper for cite-seq prediction, it uses R package saver-x to denoise the count matrix and then use multi-branch NN to predict. I have not tried this, just as a share. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2020949,
          "author_name": "llttyy",
          "author_url": "",
          "post_date": "11/07/2022 21:45:59",
          "content": "<p>Thanks for your sharing. However, according to their results, the cor even cannot exceed the public methods. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2021105,
          "author_name": "jinyang18",
          "author_url": "",
          "post_date": "11/08/2022 02:25:26",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F6ff77898959d04bbd1f690a3bcff4df9%2Fcorrelation.png?generation=1667874310402899&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F48e8e073775313e36817686679fd5976%2Fcorrelation%20of%20each%20marker.png?generation=1667874324356259&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2021106,
          "author_name": "jinyang18",
          "author_url": "",
          "post_date": "11/08/2022 02:26:13",
          "content": "<p>Within dataset correlation is pretty high, it seems to me though. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2020898,
      "author_name": "jinyang18",
      "author_url": "",
      "post_date": "11/07/2022 20:51:31",
      "content": "<p>What does 1. here means? Would you like to explain it more? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2020951,
          "author_name": "llttyy",
          "author_url": "",
          "post_date": "11/07/2022 21:46:55",
          "content": "<p>You can search \"single cell gene expression imputation\" for more information. Here is one paper: <a href=\"https://academic.oup.com/nar/article-abstract/50/9/4877/6582166\" target=\"_blank\">https://academic.oup.com/nar/article-abstract/50/9/4877/6582166</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2021233,
      "author_name": "kraigyuhengtou",
      "author_url": "",
      "post_date": "11/08/2022 05:38:08",
      "content": "<p>Single Cell Manifold preserving feature selection to select important features, to be kept before dimensionality reduction. <br>\n<a href=\"https://scmer.readthedocs.io/en/latest/modules.html#scmer.UmapL1\" target=\"_blank\">https://scmer.readthedocs.io/en/latest/modules.html#scmer.UmapL1</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2021234,
          "author_name": "kraigyuhengtou",
          "author_url": "",
          "post_date": "11/08/2022 05:39:35",
          "content": "<p>also self-assembling manifolds to select seq genes, and stochastic gates to find features that directly improve score  when regressed on labels; </p>\n<p>my teammate and I have the relevant features selected, but we don't have time to fine-tune the prediction network, so we don't know the best score we can get; </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2021749,
          "author_name": "llttyy",
          "author_url": "",
          "post_date": "11/08/2022 13:03:33",
          "content": "<p><a href=\"https://genomebiology.biomedcentral.com/counter/pdf/10.1186/s13059-019-1898-6.pdf\" target=\"_blank\">https://genomebiology.biomedcentral.com/counter/pdf/10.1186/s13059-019-1898-6.pdf</a> Hi, I think according to this paper, at least for scRNA-seq, FA and PCA are top methods to handle it, so I do not think UMAPs could be a good choice, and normally UMAPs need PCs as initial points I think.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2021869,
          "author_name": "kraigyuhengtou",
          "author_url": "",
          "post_date": "11/08/2022 14:30:26",
          "content": "<p>Yes it seems UMAP isn't a good choice, because UMAP and tsne both focus on preserving local structure when choosing its lower dimensional embedding, whereas PCA preserves global structure. Global structure is more important to maximize correlation score for this competition.</p>\n<p>However, scmer has an advantage in that it can act like supervised feature selection, Scmer minimizes the KL divergence between the latent embedding's cell-to-cell matrix and that of the original matrix, but it can also be set to minimize the divergence between the latent embedding (of the CITE-seq data) and the target protein manifold.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2021763,
      "author_name": "feitengli",
      "author_url": "",
      "post_date": "11/08/2022 13:21:21",
      "content": "<p>I stopped at 0.812 with single model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2029090,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2022 12:36:09",
      "content": "<p>it's awesome</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2020777": "Hi all, I am thinking about how to reach 0.814. I have tried bio-driven idea including:\n\n1. For cite-seq, use imputed significant protein labeled genes to train, and I have tried MAGIC and Deep Impute.\n\n2. For cite-seq, use cells to construct cell-cell sim network and use GNN to predict protein.\n\n3. For cite-seq, use correlation to select high correlated gene-protein relation, and then predict the protein one by one.\n\n4. For cite-seq, encode gender, time and cell type information into the input file, and then perform the training process and prediction process.\n\n5. For multiome, use gene activity score rather than original peak information.\n\nAll of them cannot improve my current score, so I am very confused about incuding extra information to this competition. I am very happy to learn some bio-driven ideas.",
    "2020896": "Zhou, Z., Ye, C., Wang, J. et al. Surface protein imputation from single cell transcriptomes by deep neural networks. Nat Commun 11, 651 (2020). \nThis is a paper for cite-seq prediction, it uses R package saver-x to denoise the count matrix and then use multi-branch NN to predict. I have not tried this, just as a share.",
    "2020898": "What does 1. here means? Would you like to explain it more?",
    "2020949": "Thanks for your sharing. However, according to their results, the cor even cannot exceed the public methods.",
    "2020951": "You can search \"single cell gene expression imputation\" for more information. Here is one paper: https://academic.oup.com/nar/article-abstract/50/9/4877/6582166",
    "2021105": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F6ff77898959d04bbd1f690a3bcff4df9%2Fcorrelation.png?generation=1667874310402899&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11080160%2F48e8e073775313e36817686679fd5976%2Fcorrelation%20of%20each%20marker.png?generation=1667874324356259&alt=media)",
    "2021106": "Within dataset correlation is pretty high, it seems to me though.",
    "2021233": "Single Cell Manifold preserving feature selection to select important features, to be kept before dimensionality reduction. \nhttps://scmer.readthedocs.io/en/latest/modules.html#scmer.UmapL1",
    "2021234": "also self-assembling manifolds to select seq genes, and stochastic gates to find features that directly improve score  when regressed on labels; \n\nmy teammate and I have the relevant features selected, but we don't have time to fine-tune the prediction network, so we don't know the best score we can get;",
    "2021749": "https://genomebiology.biomedcentral.com/counter/pdf/10.1186/s13059-019-1898-6.pdf Hi, I think according to this paper, at least for scRNA-seq, FA and PCA are top methods to handle it, so I do not think UMAPs could be a good choice, and normally UMAPs need PCs as initial points I think.",
    "2021763": "I stopped at 0.812 with single model.",
    "2021869": "Yes it seems UMAP isn't a good choice, because UMAP and tsne both focus on preserving local structure when choosing its lower dimensional embedding, whereas PCA preserves global structure. Global structure is more important to maximize correlation score for this competition.\n\nHowever, scmer has an advantage in that it can act like supervised feature selection, Scmer minimizes the KL divergence between the latent embedding's cell-to-cell matrix and that of the original matrix, but it can also be set to minimize the divergence between the latent embedding (of the CITE-seq data) and the target protein manifold.",
    "2029090": "it's awesome"
  },
  "source": "meta"
}