{
  "id": 363052,
  "title": "It is a time series...",
  "url": "/competitions/open-problems-multimodal/discussion/363052",
  "author_name": "AmbrosM",
  "post_date": "2022-10-30T18:24:30.389000",
  "votes": 57,
  "comment_count": 15,
  "views": 0,
  "content": "<p>In this competition, we are working with a time series. The cells in the experiment can be classified into seven cell types. If we plot the ratio of cell types by day, we see that during the course of the experiment the hematopoietic stem cells (HSC, green in the diagram) slowly change into other cell types. We begin with 50 % HSC, and on day 7 only 20 % HSC remain. We don't know the counts for day 10, but we may guess that there will be even fewer hematopoietic stem cells left.</p>\n<p><img src=\"https://i.imgur.com/cW5PbjY.png\" alt=\"CITEseq cell types by day\"></p>\n<p>If we cross-validate with a <code>GroupKFold</code> by days, we see that for both subproblems the final validation day (day 4 for CITEseq and day 7 for Multiome) has the lowest cv score, perhaps because of the diversity of the cell types and because we're extrapolating into the future.</p>\n<p>As is well known, the private leaderboard is calculated on predictions for a future day (day 7 for CITEseq and day 10 for Multiome). Is it pessimistic if I conjecture that the scores for day 7 of CITEseq and day 10 of Multiome will be even lower?</p>\n<p><img src=\"https://i.imgur.com/wqzYsyY.png\" alt=\"CITEseq cv scores by day\"></p>\n<p>Source code for the upper diagrams is in the <a href=\"https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense\" target=\"_blank\">EDA which makes sense</a>.</p>",
  "messages": [
    {
      "id": 2010409,
      "postDate": "2022-10-30T18:24:30.390Z",
      "content": "<p>In this competition, we are working with a time series. The cells in the experiment can be classified into seven cell types. If we plot the ratio of cell types by day, we see that during the course of the experiment the hematopoietic stem cells (HSC, green in the diagram) slowly change into other cell types. We begin with 50 % HSC, and on day 7 only 20 % HSC remain. We don't know the counts for day 10, but we may guess that there will be even fewer hematopoietic stem cells left.</p>\n<p><img src=\"https://i.imgur.com/cW5PbjY.png\" alt=\"CITEseq cell types by day\"></p>\n<p>If we cross-validate with a <code>GroupKFold</code> by days, we see that for both subproblems the final validation day (day 4 for CITEseq and day 7 for Multiome) has the lowest cv score, perhaps because of the diversity of the cell types and because we're extrapolating into the future.</p>\n<p>As is well known, the private leaderboard is calculated on predictions for a future day (day 7 for CITEseq and day 10 for Multiome). Is it pessimistic if I conjecture that the scores for day 7 of CITEseq and day 10 of Multiome will be even lower?</p>\n<p><img src=\"https://i.imgur.com/wqzYsyY.png\" alt=\"CITEseq cv scores by day\"></p>\n<p>Source code for the upper diagrams is in the <a href=\"https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense\" target=\"_blank\">EDA which makes sense</a>.</p>",
      "rawMarkdown": "In this competition, we are working with a time series. The cells in the experiment can be classified into seven cell types. If we plot the ratio of cell types by day, we see that during the course of the experiment the hematopoietic stem cells (HSC, green in the diagram) slowly change into other cell types. We begin with 50 % HSC, and on day 7 only 20 % HSC remain. We don't know the counts for day 10, but we may guess that there will be even fewer hematopoietic stem cells left.\n\n![CITEseq cell types by day](https://i.imgur.com/cW5PbjY.png)\n\nIf we cross-validate with a `GroupKFold` by days, we see that for both subproblems the final validation day (day 4 for CITEseq and day 7 for Multiome) has the lowest cv score, perhaps because of the diversity of the cell types and because we're extrapolating into the future.\n\nAs is well known, the private leaderboard is calculated on predictions for a future day (day 7 for CITEseq and day 10 for Multiome). Is it pessimistic if I conjecture that the scores for day 7 of CITEseq and day 10 of Multiome will be even lower?\n\n![CITEseq cv scores by day](https://i.imgur.com/wqzYsyY.png)\n\nSource code for the upper diagrams is in the [EDA which makes sense](https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense).\n",
      "votes": 57
    },
    {
      "id": 2010870,
      "postDate": "2022-10-31T07:42:05.990Z",
      "content": "<p>Thank you for sharing great insights.</p>\n<p>This is a just idea without validation , but let me share.<br>\nThere should be some genes(for CITEseq) or chromatin states(for Multiome) contributing to the accuracy of cell type prediction.<br>\nI assume many competitors are using dimension reduction techniques like SVD but those kinds of features should not be included in dimension reduction targets and should be used as raw columns if you want to predict day 7 of CITEseq and day 10 of Multiome.</p>\n<p>If you want to find contributing features for cell type prediction, you can create several GDBT models in order to overcome memory issues.<br>\neg. 2300 features  * 10 models and find top N important features of each models.</p>\n<p>This idea is hard to validate with only Public LB score but I hope it will help your Private LB scores</p>",
      "rawMarkdown": "Thank you for sharing great insights.\n\nThis is a just idea without validation , but let me share.\nThere should be some genes(for CITEseq) or chromatin states(for Multiome) contributing to the accuracy of cell type prediction.\nI assume many competitors are using dimension reduction techniques like SVD but those kinds of features should not be included in dimension reduction targets and should be used as raw columns if you want to predict day 7 of CITEseq and day 10 of Multiome.\n\nIf you want to find contributing features for cell type prediction, you can create several GDBT models in order to overcome memory issues.\neg. 2300 features  * 10 models and find top N important features of each models.\n\nThis idea is hard to validate with only Public LB score but I hope it will help your Private LB scores",
      "votes": 7,
      "replies": [
        {
          "id": 2017167,
          "postDate": "2022-11-04T15:00:27.500Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2015069,
      "postDate": "2022-11-03T04:44:06.787Z",
      "content": "<p>Just adding this information here from some comments in the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344607#1960755\" target=\"_blank\">Welcome thread</a></p>\n<p>\"Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.\"</p>\n<p>from/host - \"the cells were cultured with StemSpan SFEM media supplemented with CC100 and thrombopoietin (TPO) over 10 days. Cells were incubated at 37ºC and media was changed every 2-3 days\"</p>\n<p>It is unclear if the train data has any days with media changed, (maybe in Multiome) and what effects that might show.  Presumably the future days 7 and 10 will have had media changed in both CITESeq and Multiome. May throw another spanner in the works.</p>",
      "rawMarkdown": "Just adding this information here from some comments in the [Welcome thread](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344607#1960755)\n\n\"Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.\"\n\nfrom/host - \"the cells were cultured with StemSpan SFEM media supplemented with CC100 and thrombopoietin (TPO) over 10 days. Cells were incubated at 37ºC and media was changed every 2-3 days\"\n\nIt is unclear if the train data has any days with media changed, (maybe in Multiome) and what effects that might show.  Presumably the future days 7 and 10 will have had media changed in both CITESeq and Multiome. May throw another spanner in the works.",
      "votes": 1
    },
    {
      "id": 2010658,
      "postDate": "2022-10-31T01:51:01.393Z",
      "content": "<p>I think one solution for this problem is to encode time series into our training step, but it is hard to say it can work or not because the private testing dataset only contains data from one day.</p>",
      "rawMarkdown": "I think one solution for this problem is to encode time series into our training step, but it is hard to say it can work or not because the private testing dataset only contains data from one day.",
      "votes": 1
    },
    {
      "id": 2030878,
      "postDate": "2022-11-15T17:14:00.130Z",
      "content": "<p>Inspiring work! Thanks for sharing. I never thought this would be a time series. </p>",
      "rawMarkdown": "Inspiring work! Thanks for sharing. I never thought this would be a time series. "
    },
    {
      "id": 2030799,
      "postDate": "2022-11-15T16:14:33.230Z",
      "content": "<p>Very interesting</p>",
      "rawMarkdown": "Very interesting"
    },
    {
      "id": 2030150,
      "postDate": "2022-11-15T08:21:45.067Z",
      "content": "<p>Nice Nice, tkx</p>",
      "rawMarkdown": "Nice Nice, tkx"
    },
    {
      "id": 2011545,
      "postDate": "2022-10-31T16:25:28.760Z",
      "content": "<p>Thanks for the topic ! </p>\n<p>May I allow myself to make some comments, with the hope someone will develop your idea further.</p>\n<p>Yes, with the change of the day - we see cell types are changing , and also proliferation capacities are changing, and may be other things are changing…<br>\nBUT:<br>\nour models should relate say RNA to Proteins  (CITE-seq),<br>\nand if changes in RNA and Proteins are synchronous - then there is no problem for the prediction models.<br>\nWHY <br>\n(it may happen):<br>\nThe naïve model is :   PROTEIN = RNA  (for the same gene ),<br>\nand for say bacteria - that might be even quite not bad prediction. <br>\nAnd so for that model - we do not care about variations of Proteins - because  there would be always the same as RNA.</p>\n<p>So to explore your idea further - we need to compare <br>\nhow something like Protein / RNA (for each CD gene) are varying with the day.<br>\nBut not solely Protein, or solely RNA. </p>",
      "rawMarkdown": "Thanks for the topic ! \n\nMay I allow myself to make some comments, with the hope someone will develop your idea further.\n\nYes, with the change of the day - we see cell types are changing , and also proliferation capacities are changing, and may be other things are changing...\nBUT:\nour models should relate say RNA to Proteins  (CITE-seq),\nand if changes in RNA and Proteins are synchronous - then there is no problem for the prediction models.\nWHY \n(it may happen):\nThe naïve model is :   PROTEIN = RNA  (for the same gene ),\nand for say bacteria - that might be even quite not bad prediction. \nAnd so for that model - we do not care about variations of Proteins - because  there would be always the same as RNA.\n\nSo to explore your idea further - we need to compare \nhow something like Protein / RNA (for each CD gene) are varying with the day.\nBut not solely Protein, or solely RNA. \n\n\n",
      "replies": [
        {
          "id": 2013233,
          "postDate": "2022-11-01T18:26:25.473Z",
          "content": "<p>Indeed, for the cite part, I used UMAP to plot the association of RNA levels by day and found this variation (presumably the color is the day, the regions are different neighborhoods):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F374429%2F70425cc81dbb206b20af036c0c7a9307%2FScreen%20Shot%202022-11-01%20at%2010.57.00%20AM.png?generation=1667325508632842&amp;alt=media\" alt=\"Sample RNA Levels\"><br>\nI also used UMAP to plot the association of protein levels by day and found this variation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F374429%2Fa9151bf3d2baf66c0f892cff7514b5b8%2FScreen%20Shot%202022-11-01%20at%2010.57.16%20AM.png?generation=1667325568797091&amp;alt=media\" alt=\"Sample Protein Levels\"><br>\nThat shows that very different groups of RNA and protein get expressed on different days, but as you say, a model should be robust to these changes, i.e. you can't call your model predictive if it only works every other Monday.</p>",
          "rawMarkdown": "Indeed, for the cite part, I used UMAP to plot the association of RNA levels by day and found this variation (presumably the color is the day, the regions are different neighborhoods):\n![Sample RNA Levels](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F374429%2F70425cc81dbb206b20af036c0c7a9307%2FScreen%20Shot%202022-11-01%20at%2010.57.00%20AM.png?generation=1667325508632842&alt=media)\nI also used UMAP to plot the association of protein levels by day and found this variation:\n![Sample Protein Levels](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F374429%2Fa9151bf3d2baf66c0f892cff7514b5b8%2FScreen%20Shot%202022-11-01%20at%2010.57.16%20AM.png?generation=1667325568797091&alt=media)\nThat shows that very different groups of RNA and protein get expressed on different days, but as you say, a model should be robust to these changes, i.e. you can't call your model predictive if it only works every other Monday.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2025297,
      "postDate": "2022-11-11T06:15:25.747Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2028758,
      "postDate": "2022-11-14T08:02:23.630Z",
      "content": "<p>Thank you so much !!!</p>",
      "rawMarkdown": "Thank you so much !!!"
    },
    {
      "id": 2028518,
      "postDate": "2022-11-14T02:59:24.127Z",
      "content": "<p>Thanks for sharing !!</p>",
      "rawMarkdown": "Thanks for sharing !!"
    },
    {
      "id": 2014848,
      "postDate": "2022-11-02T22:45:43.397Z",
      "content": "<p>Inspiring work, thanks a lot.</p>",
      "rawMarkdown": "Inspiring work, thanks a lot."
    },
    {
      "id": 2012386,
      "postDate": "2022-11-01T07:40:01.723Z",
      "content": "<p>thanks a lot</p>",
      "rawMarkdown": "thanks a lot"
    },
    {
      "id": 2010426,
      "postDate": "2022-10-30T18:35:24.233Z",
      "content": "<p>Very interesting, thanks for sharing.</p>",
      "rawMarkdown": "Very interesting, thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 2010870,
      "author_name": "Shinya(Hyper-Positive-Yancy)",
      "author_url": "",
      "post_date": "2022-10-31T07:42:05.990000",
      "content": "<p>Thank you for sharing great insights.</p>\n<p>This is a just idea without validation , but let me share.<br>\nThere should be some genes(for CITEseq) or chromatin states(for Multiome) contributing to the accuracy of cell type prediction.<br>\nI assume many competitors are using dimension reduction techniques like SVD but those kinds of features should not be included in dimension reduction targets and should be used as raw columns if you want to predict day 7 of CITEseq and day 10 of Multiome.</p>\n<p>If you want to find contributing features for cell type prediction, you can create several GDBT models in order to overcome memory issues.<br>\neg. 2300 features  * 10 models and find top N important features of each models.</p>\n<p>This idea is hard to validate with only Public LB score but I hope it will help your Private LB scores</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2017167,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-04T15:00:27.500000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2015069,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2022-11-03T04:44:06.787000",
      "content": "<p>Just adding this information here from some comments in the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344607#1960755\" target=\"_blank\">Welcome thread</a></p>\n<p>\"Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.\"</p>\n<p>from/host - \"the cells were cultured with StemSpan SFEM media supplemented with CC100 and thrombopoietin (TPO) over 10 days. Cells were incubated at 37ºC and media was changed every 2-3 days\"</p>\n<p>It is unclear if the train data has any days with media changed, (maybe in Multiome) and what effects that might show.  Presumably the future days 7 and 10 will have had media changed in both CITESeq and Multiome. May throw another spanner in the works.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2010658,
      "author_name": "TESUZI",
      "author_url": "",
      "post_date": "2022-10-31T01:51:01.393000",
      "content": "<p>I think one solution for this problem is to encode time series into our training step, but it is hard to say it can work or not because the private testing dataset only contains data from one day.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2030878,
      "author_name": "Zhilin Jin",
      "author_url": "",
      "post_date": "2022-11-15T17:14:00.130000",
      "content": "<p>Inspiring work! Thanks for sharing. I never thought this would be a time series. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2030799,
      "author_name": "Frenkist",
      "author_url": "",
      "post_date": "2022-11-15T16:14:33.230000",
      "content": "<p>Very interesting</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2030150,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-15T08:21:45.067000",
      "content": "<p>Nice Nice, tkx</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2011545,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2022-10-31T16:25:28.760000",
      "content": "<p>Thanks for the topic ! </p>\n<p>May I allow myself to make some comments, with the hope someone will develop your idea further.</p>\n<p>Yes, with the change of the day - we see cell types are changing , and also proliferation capacities are changing, and may be other things are changing…<br>\nBUT:<br>\nour models should relate say RNA to Proteins  (CITE-seq),<br>\nand if changes in RNA and Proteins are synchronous - then there is no problem for the prediction models.<br>\nWHY <br>\n(it may happen):<br>\nThe naïve model is :   PROTEIN = RNA  (for the same gene ),<br>\nand for say bacteria - that might be even quite not bad prediction. <br>\nAnd so for that model - we do not care about variations of Proteins - because  there would be always the same as RNA.</p>\n<p>So to explore your idea further - we need to compare <br>\nhow something like Protein / RNA (for each CD gene) are varying with the day.<br>\nBut not solely Protein, or solely RNA. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2013233,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2022-11-01T18:26:25.473000",
          "content": "<p>Indeed, for the cite part, I used UMAP to plot the association of RNA levels by day and found this variation (presumably the color is the day, the regions are different neighborhoods):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F374429%2F70425cc81dbb206b20af036c0c7a9307%2FScreen%20Shot%202022-11-01%20at%2010.57.00%20AM.png?generation=1667325508632842&amp;alt=media\" alt=\"Sample RNA Levels\"><br>\nI also used UMAP to plot the association of protein levels by day and found this variation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F374429%2Fa9151bf3d2baf66c0f892cff7514b5b8%2FScreen%20Shot%202022-11-01%20at%2010.57.16%20AM.png?generation=1667325568797091&amp;alt=media\" alt=\"Sample Protein Levels\"><br>\nThat shows that very different groups of RNA and protein get expressed on different days, but as you say, a model should be robust to these changes, i.e. you can't call your model predictive if it only works every other Monday.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2025297,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-11T06:15:25.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2028758,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T08:02:23.630000",
      "content": "<p>Thank you so much !!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2028518,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T02:59:24.127000",
      "content": "<p>Thanks for sharing !!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2014848,
      "author_name": "Priyanshu S. Prajapati",
      "author_url": "",
      "post_date": "2022-11-02T22:45:43.397000",
      "content": "<p>Inspiring work, thanks a lot.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2012386,
      "author_name": "Liilili",
      "author_url": "",
      "post_date": "2022-11-01T07:40:01.723000",
      "content": "<p>thanks a lot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2010426,
      "author_name": "aspiring",
      "author_url": "",
      "post_date": "2022-10-30T18:35:24.233000",
      "content": "<p>Very interesting, thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2010409": "In this competition, we are working with a time series. The cells in the experiment can be classified into seven cell types. If we plot the ratio of cell types by day, we see that during the course of the experiment the hematopoietic stem cells (HSC, green in the diagram) slowly change into other cell types. We begin with 50 % HSC, and on day 7 only 20 % HSC remain. We don't know the counts for day 10, but we may guess that there will be even fewer hematopoietic stem cells left.\n\n![CITEseq cell types by day](https://i.imgur.com/cW5PbjY.png)\n\nIf we cross-validate with a `GroupKFold` by days, we see that for both subproblems the final validation day (day 4 for CITEseq and day 7 for Multiome) has the lowest cv score, perhaps because of the diversity of the cell types and because we're extrapolating into the future.\n\nAs is well known, the private leaderboard is calculated on predictions for a future day (day 7 for CITEseq and day 10 for Multiome). Is it pessimistic if I conjecture that the scores for day 7 of CITEseq and day 10 of Multiome will be even lower?\n\n![CITEseq cv scores by day](https://i.imgur.com/wqzYsyY.png)\n\nSource code for the upper diagrams is in the [EDA which makes sense](https://www.kaggle.com/code/ambrosm/msci-eda-which-makes-sense).\n",
    "2010870": "Thank you for sharing great insights.\n\nThis is a just idea without validation , but let me share.\nThere should be some genes(for CITEseq) or chromatin states(for Multiome) contributing to the accuracy of cell type prediction.\nI assume many competitors are using dimension reduction techniques like SVD but those kinds of features should not be included in dimension reduction targets and should be used as raw columns if you want to predict day 7 of CITEseq and day 10 of Multiome.\n\nIf you want to find contributing features for cell type prediction, you can create several GDBT models in order to overcome memory issues.\neg. 2300 features  * 10 models and find top N important features of each models.\n\nThis idea is hard to validate with only Public LB score but I hope it will help your Private LB scores",
    "2015069": "Just adding this information here from some comments in the [Welcome thread](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344607#1960755)\n\n\"Just to confirm: the cells were NOT stimulated in any way after plating? It would be helpful, if you could describe a little bit more about what happened to the cells after they were obtained from the donors over those different time periods.\"\n\nfrom/host - \"the cells were cultured with StemSpan SFEM media supplemented with CC100 and thrombopoietin (TPO) over 10 days. Cells were incubated at 37ºC and media was changed every 2-3 days\"\n\nIt is unclear if the train data has any days with media changed, (maybe in Multiome) and what effects that might show.  Presumably the future days 7 and 10 will have had media changed in both CITESeq and Multiome. May throw another spanner in the works.",
    "2010658": "I think one solution for this problem is to encode time series into our training step, but it is hard to say it can work or not because the private testing dataset only contains data from one day.",
    "2030878": "Inspiring work! Thanks for sharing. I never thought this would be a time series. ",
    "2030799": "Very interesting",
    "2030150": "Nice Nice, tkx",
    "2011545": "Thanks for the topic ! \n\nMay I allow myself to make some comments, with the hope someone will develop your idea further.\n\nYes, with the change of the day - we see cell types are changing , and also proliferation capacities are changing, and may be other things are changing...\nBUT:\nour models should relate say RNA to Proteins  (CITE-seq),\nand if changes in RNA and Proteins are synchronous - then there is no problem for the prediction models.\nWHY \n(it may happen):\nThe naïve model is :   PROTEIN = RNA  (for the same gene ),\nand for say bacteria - that might be even quite not bad prediction. \nAnd so for that model - we do not care about variations of Proteins - because  there would be always the same as RNA.\n\nSo to explore your idea further - we need to compare \nhow something like Protein / RNA (for each CD gene) are varying with the day.\nBut not solely Protein, or solely RNA. \n\n\n",
    "2025297": "",
    "2028758": "Thank you so much !!!",
    "2028518": "Thanks for sharing !!",
    "2014848": "Inspiring work, thanks a lot.",
    "2012386": "thanks a lot",
    "2010426": "Very interesting, thanks for sharing."
  }
}