{
  "id": 346888,
  "title": "Understanding the Competition + Some Domain Knowledge",
  "url": "/competitions/open-problems-multimodal/discussion/346888",
  "author_name": "Ravi Shah",
  "post_date": "2022-08-21T21:52:04.997000",
  "votes": 132,
  "comment_count": 17,
  "views": 0,
  "content": "<h1>The Goal:</h1>\n<ul>\n<li>We can think of this competition as having two parts: Multiome and CITEseq.</li>\n<li>For both parts, we need a model that can make a vector prediction given a vector input.</li>\n<li>Each cell id represents a sample</li>\n<li>For Multiome, we need to use transformed numerical DNA data to predict transformed numerical RNA data.</li>\n<li>For CITEseq, we need to use transformed numerical RNA data to predict transformed numerical Protein data</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fbf8c54a125ecfd986cd30c5ecc0724a2%2Fcentral_dogma2.PNG?generation=1661118946987012&amp;alt=media\" alt=\"\"><br>\n(The numbers are just random obviously use the values in the actual dataset. Also this only shows for one sample, you will actually be doing this for many samples)<br>\n<a href=\"https://en.wikipedia.org/wiki/Central_dogma_of_molecular_biology\" target=\"_blank\">learn more about central dogma of molecular biology</a></p>\n<h1>Submission</h1>\n<ul>\n<li>format your predictions similar to the targets tables</li>\n<li>index the cell_id / gene_id from the label matrix to fill out the evaluations_ids.csv</li>\n<li>map the evaulations_ids.csv to the sample_submission.csv using the row_id</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1997e2ec55923d44b4d0a53221311456%2Fsub_pic.PNG?generation=1661117789631290&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1908671,
      "postDate": "2022-08-21T21:52:04.997Z",
      "content": "<h1>The Goal:</h1>\n<ul>\n<li>We can think of this competition as having two parts: Multiome and CITEseq.</li>\n<li>For both parts, we need a model that can make a vector prediction given a vector input.</li>\n<li>Each cell id represents a sample</li>\n<li>For Multiome, we need to use transformed numerical DNA data to predict transformed numerical RNA data.</li>\n<li>For CITEseq, we need to use transformed numerical RNA data to predict transformed numerical Protein data</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fbf8c54a125ecfd986cd30c5ecc0724a2%2Fcentral_dogma2.PNG?generation=1661118946987012&amp;alt=media\" alt=\"\"><br>\n(The numbers are just random obviously use the values in the actual dataset. Also this only shows for one sample, you will actually be doing this for many samples)<br>\n<a href=\"https://en.wikipedia.org/wiki/Central_dogma_of_molecular_biology\" target=\"_blank\">learn more about central dogma of molecular biology</a></p>\n<h1>Submission</h1>\n<ul>\n<li>format your predictions similar to the targets tables</li>\n<li>index the cell_id / gene_id from the label matrix to fill out the evaluations_ids.csv</li>\n<li>map the evaulations_ids.csv to the sample_submission.csv using the row_id</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1997e2ec55923d44b4d0a53221311456%2Fsub_pic.PNG?generation=1661117789631290&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "# The Goal:\n\n- We can think of this competition as having two parts: Multiome and CITEseq.\n- For both parts, we need a model that can make a vector prediction given a vector input.\n- Each cell id represents a sample\n- For Multiome, we need to use transformed numerical DNA data to predict transformed numerical RNA data.\n- For CITEseq, we need to use transformed numerical RNA data to predict transformed numerical Protein data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fbf8c54a125ecfd986cd30c5ecc0724a2%2Fcentral_dogma2.PNG?generation=1661118946987012&alt=media)\n(The numbers are just random obviously use the values in the actual dataset. Also this only shows for one sample, you will actually be doing this for many samples)\n[learn more about central dogma of molecular biology](https://en.wikipedia.org/wiki/Central_dogma_of_molecular_biology)\n\n# Submission\n\n- format your predictions similar to the targets tables\n- index the cell_id / gene_id from the label matrix to fill out the evaluations_ids.csv\n- map the evaulations_ids.csv to the sample_submission.csv using the row_id\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1997e2ec55923d44b4d0a53221311456%2Fsub_pic.PNG?generation=1661117789631290&alt=media)",
      "votes": 129
    },
    {
      "id": 1923314,
      "postDate": "2022-09-02T06:05:28.470Z",
      "content": "<p>Very useful information. upvoted <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> </p>",
      "rawMarkdown": "Very useful information. upvoted @ravishah1 ",
      "votes": 10
    },
    {
      "id": 1919914,
      "postDate": "2022-08-30T18:25:21.307Z",
      "content": "<p>Such a  comprehensive explanation! Thanks for sharing</p>",
      "rawMarkdown": "Such a  comprehensive explanation! Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1919088,
      "postDate": "2022-08-30T06:02:11.200Z",
      "content": "<p>Thanks for sharing, very helpful <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> !</p>",
      "rawMarkdown": "Thanks for sharing, very helpful @ravishah1 !",
      "votes": 1
    },
    {
      "id": 1916491,
      "postDate": "2022-08-27T23:33:55.683Z",
      "content": "<p>This is so concise and so helpful.  This should absolutely be a gold level discussion/thread</p>",
      "rawMarkdown": "This is so concise and so helpful.  This should absolutely be a gold level discussion/thread",
      "votes": 1
    },
    {
      "id": 1930653,
      "postDate": "2022-09-08T05:20:12.470Z",
      "content": "<p>Simple but powerful. Thanks for sharing!</p>",
      "rawMarkdown": "Simple but powerful. Thanks for sharing!"
    },
    {
      "id": 1913728,
      "postDate": "2022-08-25T13:21:48.737Z",
      "content": "<p>Thank you!<br>\nBut I can't understand where those vectors in the given data… <br>\nIn the cite part we have three files: train_cite_inputs.h5, train_cite_targets.h5 and test_cite_inputs.h5. Can you explain what these files mean? </p>",
      "rawMarkdown": "Thank you!\nBut I can't understand where those vectors in the given data... \nIn the cite part we have three files: train_cite_inputs.h5, train_cite_targets.h5 and test_cite_inputs.h5. Can you explain what these files mean? ",
      "votes": 2,
      "replies": [
        {
          "id": 1913755,
          "postDate": "2022-08-25T13:43:05.403Z",
          "content": "<p>Gene -&gt; RNA-&gt; Protein. (transcription ;  translation )</p>\n<p>You train on: \"RNA\"<br>\nYou must predict - Target: \"Protein\" (only selected proteins - that ones called CD**) </p>\n<p><a href=\"https://en.wikipedia.org/wiki/Cluster_of_differentiation\" target=\"_blank\">https://en.wikipedia.org/wiki/Cluster_of_differentiation</a><br>\nPS<br>\nyou may ask biological/bioinformatics questions in <a href=\"https://t.me/sberlogabio\" target=\"_blank\">https://t.me/sberlogabio</a></p>",
          "rawMarkdown": "Gene -> RNA-> Protein. (transcription ;  translation )\n\nYou train on: \"RNA\"\nYou must predict - Target: \"Protein\" (only selected proteins - that ones called CD**) \n\nhttps://en.wikipedia.org/wiki/Cluster_of_differentiation\nPS\nyou may ask biological/bioinformatics questions in https://t.me/sberlogabio\n\n\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 1960479,
      "postDate": "2022-09-28T16:04:04.617Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1910943,
      "postDate": "2022-08-23T19:46:37.710Z",
      "content": "<p>Great insight! Thanks for sharing this. </p>",
      "rawMarkdown": "Great insight! Thanks for sharing this. ",
      "votes": 1
    },
    {
      "id": 1910329,
      "postDate": "2022-08-23T10:51:22.323Z",
      "content": "<p>Thanks. It is clear</p>",
      "rawMarkdown": "Thanks. It is clear",
      "votes": 1
    },
    {
      "id": 1908906,
      "postDate": "2022-08-22T06:10:51.417Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1920107,
      "postDate": "2022-08-30T22:17:34.740Z",
      "content": "<p>Thanks for making things clearer.</p>",
      "rawMarkdown": "Thanks for making things clearer.",
      "votes": 2
    },
    {
      "id": 2029790,
      "postDate": "2022-11-14T23:16:20.170Z",
      "content": "<p>Thanks. Very nice.</p>",
      "rawMarkdown": "Thanks. Very nice."
    },
    {
      "id": 2029052,
      "postDate": "2022-11-14T11:57:39.180Z",
      "content": "<p>This is really useful, thanks!</p>",
      "rawMarkdown": "This is really useful, thanks!"
    },
    {
      "id": 2028680,
      "postDate": "2022-11-14T06:57:45.107Z",
      "content": "<p>This is great! Thank you.    </p>",
      "rawMarkdown": "This is great! Thank you.\t"
    },
    {
      "id": 2010309,
      "postDate": "2022-10-30T17:16:44.160Z",
      "content": "<p>Super helpful, thanks!</p>",
      "rawMarkdown": "Super helpful, thanks!"
    },
    {
      "id": 1989502,
      "postDate": "2022-10-16T03:21:22.397Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!"
    }
  ],
  "comments": [
    {
      "id": 1923314,
      "author_name": "Batros Jamali",
      "author_url": "",
      "post_date": "2022-09-02T06:05:28.470000",
      "content": "<p>Very useful information. upvoted <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> </p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 1919914,
      "author_name": "Chris Ried",
      "author_url": "",
      "post_date": "2022-08-30T18:25:21.307000",
      "content": "<p>Such a  comprehensive explanation! Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1919088,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2022-08-30T06:02:11.200000",
      "content": "<p>Thanks for sharing, very helpful <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1916491,
      "author_name": "KirkDCO",
      "author_url": "",
      "post_date": "2022-08-27T23:33:55.683000",
      "content": "<p>This is so concise and so helpful.  This should absolutely be a gold level discussion/thread</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1930653,
      "author_name": "Hongjun Jang",
      "author_url": "",
      "post_date": "2022-09-08T05:20:12.470000",
      "content": "<p>Simple but powerful. Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913728,
      "author_name": "Veronika Chernova",
      "author_url": "",
      "post_date": "2022-08-25T13:21:48.737000",
      "content": "<p>Thank you!<br>\nBut I can't understand where those vectors in the given data… <br>\nIn the cite part we have three files: train_cite_inputs.h5, train_cite_targets.h5 and test_cite_inputs.h5. Can you explain what these files mean? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1913755,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2022-08-25T13:43:05.403000",
          "content": "<p>Gene -&gt; RNA-&gt; Protein. (transcription ;  translation )</p>\n<p>You train on: \"RNA\"<br>\nYou must predict - Target: \"Protein\" (only selected proteins - that ones called CD**) </p>\n<p><a href=\"https://en.wikipedia.org/wiki/Cluster_of_differentiation\" target=\"_blank\">https://en.wikipedia.org/wiki/Cluster_of_differentiation</a><br>\nPS<br>\nyou may ask biological/bioinformatics questions in <a href=\"https://t.me/sberlogabio\" target=\"_blank\">https://t.me/sberlogabio</a></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1960479,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-28T16:04:04.617000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1910943,
      "author_name": "Chirag Desai",
      "author_url": "",
      "post_date": "2022-08-23T19:46:37.710000",
      "content": "<p>Great insight! Thanks for sharing this. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1910329,
      "author_name": "vreabvgyuol",
      "author_url": "",
      "post_date": "2022-08-23T10:51:22.323000",
      "content": "<p>Thanks. It is clear</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1908906,
      "author_name": "Saman",
      "author_url": "",
      "post_date": "2022-08-22T06:10:51.417000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1920107,
      "author_name": "Vijay Katta",
      "author_url": "",
      "post_date": "2022-08-30T22:17:34.740000",
      "content": "<p>Thanks for making things clearer.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2029790,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T23:16:20.170000",
      "content": "<p>Thanks. Very nice.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2029052,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T11:57:39.180000",
      "content": "<p>This is really useful, thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2028680,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T06:57:45.107000",
      "content": "<p>This is great! Thank you.    </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2010309,
      "author_name": "Hiroyuki Onishi",
      "author_url": "",
      "post_date": "2022-10-30T17:16:44.160000",
      "content": "<p>Super helpful, thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1989502,
      "author_name": "kasumi66",
      "author_url": "",
      "post_date": "2022-10-16T03:21:22.397000",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1908671": "# The Goal:\n\n- We can think of this competition as having two parts: Multiome and CITEseq.\n- For both parts, we need a model that can make a vector prediction given a vector input.\n- Each cell id represents a sample\n- For Multiome, we need to use transformed numerical DNA data to predict transformed numerical RNA data.\n- For CITEseq, we need to use transformed numerical RNA data to predict transformed numerical Protein data\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fbf8c54a125ecfd986cd30c5ecc0724a2%2Fcentral_dogma2.PNG?generation=1661118946987012&alt=media)\n(The numbers are just random obviously use the values in the actual dataset. Also this only shows for one sample, you will actually be doing this for many samples)\n[learn more about central dogma of molecular biology](https://en.wikipedia.org/wiki/Central_dogma_of_molecular_biology)\n\n# Submission\n\n- format your predictions similar to the targets tables\n- index the cell_id / gene_id from the label matrix to fill out the evaluations_ids.csv\n- map the evaulations_ids.csv to the sample_submission.csv using the row_id\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1997e2ec55923d44b4d0a53221311456%2Fsub_pic.PNG?generation=1661117789631290&alt=media)",
    "1923314": "Very useful information. upvoted @ravishah1 ",
    "1919914": "Such a  comprehensive explanation! Thanks for sharing",
    "1919088": "Thanks for sharing, very helpful @ravishah1 !",
    "1916491": "This is so concise and so helpful.  This should absolutely be a gold level discussion/thread",
    "1930653": "Simple but powerful. Thanks for sharing!",
    "1913728": "Thank you!\nBut I can't understand where those vectors in the given data... \nIn the cite part we have three files: train_cite_inputs.h5, train_cite_targets.h5 and test_cite_inputs.h5. Can you explain what these files mean? ",
    "1960479": "",
    "1910943": "Great insight! Thanks for sharing this. ",
    "1910329": "Thanks. It is clear",
    "1908906": "Thanks for sharing",
    "1920107": "Thanks for making things clearer.",
    "2029790": "Thanks. Very nice.",
    "2029052": "This is really useful, thanks!",
    "2028680": "This is great! Thank you.\t",
    "2010309": "Super helpful, thanks!",
    "1989502": "Thanks for sharing!!"
  }
}