{
  "id": 353523,
  "title": "Multiome Shapes and scoring",
  "url": "/competitions/open-problems-multimodal/discussion/353523",
  "author_name": "",
  "post_date": "2022-09-18T19:38:38.949554300Z",
  "votes": 39,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The following graph is intended to help visualize at once all the data names and shapes involved in the Multiome part of the competition.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fd81e9456a5a04d89da5f917feab412c7%2FMultiOmeShape.png?generation=1663530730501366&amp;alt=media\" alt=\"\"></p>\n<p>We have to predict Multiome Test_targets. However, we only need to submit part of it. 30% of the rows and for each row, 15% of the columns. It is the second part of the submission (58,931,360 row_ids in the submission file)</p>\n<p>We can use this graph and the one for <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353522\" target=\"_blank\">CiteSeq</a> to compare more specifically the Test_targets Data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fa490de664914da4481720d4415a68cde%2FTestTargetShapes.png?generation=1663562079905748&amp;alt=media\" alt=\"\"></p>\n<p>The organizer decided to sample Multiome Test_target, due to its large size. Only 30% of the rows (cells) are used for the scoring and for each of these selected rows, only 15% of the columns (genes) are selected.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F0eb956102a53b8b105466e255f749367%2FTestTargetRatios.png?generation=1663563409305863&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Before sampling, the 2D Multiome Test_target matrix is 192 times larger than its CiteSeq equivalent.</li>\n<li>After sampling the ratio dropped to 8.7</li>\n</ul>\n<p>However <strong>the sampling has some very important implications on the scoring</strong></p>\n<p>The submission file has 8.7 times more data for Multiome than for CiteSeq and we can at first glance think that Multiome is way more important for the scoring. However what really matters is the number of rows, because each row correlation has the same weight, even if the row is much longer!<br>\nIf we look at the number of rows after sampling, suddenly CiteSeq is more important!! CiteSeq has all its 48,663 rows while Multiome is left with only 16,780. <br>\nConsequently CiteSeq represent ~74% of the score and Multiome ~26% if we look at the number of rows (and number of correlations) used.<br>\nThis can be checked roughly with the CV scores. CiteSeq CV is around 89 and Multiome 66.5. The LB score is closer to the CiteSeq CV score.</p>\n<p><strong><em>Amendment: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a></em></strong><br>\nAs <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> explained below, there was an issue in the initial data and the organizer is now ignoring 7,476 rows in the CITEseq submission. You can see the implications on the scoring weights <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360180\" target=\"_blank\">here.</a></p>",
  "messages": [
    {
      "id": "1945054",
      "postDate": "09/18/2022 19:38:38",
      "content": "<p>The following graph is intended to help visualize at once all the data names and shapes involved in the Multiome part of the competition.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fd81e9456a5a04d89da5f917feab412c7%2FMultiOmeShape.png?generation=1663530730501366&amp;alt=media\" alt=\"\"></p>\n<p>We have to predict Multiome Test_targets. However, we only need to submit part of it. 30% of the rows and for each row, 15% of the columns. It is the second part of the submission (58,931,360 row_ids in the submission file)</p>\n<p>We can use this graph and the one for <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353522\" target=\"_blank\">CiteSeq</a> to compare more specifically the Test_targets Data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fa490de664914da4481720d4415a68cde%2FTestTargetShapes.png?generation=1663562079905748&amp;alt=media\" alt=\"\"></p>\n<p>The organizer decided to sample Multiome Test_target, due to its large size. Only 30% of the rows (cells) are used for the scoring and for each of these selected rows, only 15% of the columns (genes) are selected.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F0eb956102a53b8b105466e255f749367%2FTestTargetRatios.png?generation=1663563409305863&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Before sampling, the 2D Multiome Test_target matrix is 192 times larger than its CiteSeq equivalent.</li>\n<li>After sampling the ratio dropped to 8.7</li>\n</ul>\n<p>However <strong>the sampling has some very important implications on the scoring</strong></p>\n<p>The submission file has 8.7 times more data for Multiome than for CiteSeq and we can at first glance think that Multiome is way more important for the scoring. However what really matters is the number of rows, because each row correlation has the same weight, even if the row is much longer!<br>\nIf we look at the number of rows after sampling, suddenly CiteSeq is more important!! CiteSeq has all its 48,663 rows while Multiome is left with only 16,780. <br>\nConsequently CiteSeq represent ~74% of the score and Multiome ~26% if we look at the number of rows (and number of correlations) used.<br>\nThis can be checked roughly with the CV scores. CiteSeq CV is around 89 and Multiome 66.5. The LB score is closer to the CiteSeq CV score.</p>\n<p><strong><em>Amendment: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a></em></strong><br>\nAs <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> explained below, there was an issue in the initial data and the organizer is now ignoring 7,476 rows in the CITEseq submission. You can see the implications on the scoring weights <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360180\" target=\"_blank\">here.</a></p>",
      "rawMarkdown": "The following graph is intended to help visualize at once all the data names and shapes involved in the Multiome part of the competition.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fd81e9456a5a04d89da5f917feab412c7%2FMultiOmeShape.png?generation=1663530730501366&alt=media)\n\nWe have to predict Multiome Test_targets. However, we only need to submit part of it. 30% of the rows and for each row, 15% of the columns. It is the second part of the submission (58,931,360 row_ids in the submission file)\n\nWe can use this graph and the one for [CiteSeq](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353522) to compare more specifically the Test_targets Data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fa490de664914da4481720d4415a68cde%2FTestTargetShapes.png?generation=1663562079905748&alt=media)\n\nThe organizer decided to sample Multiome Test_target, due to its large size. Only 30% of the rows (cells) are used for the scoring and for each of these selected rows, only 15% of the columns (genes) are selected.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F0eb956102a53b8b105466e255f749367%2FTestTargetRatios.png?generation=1663563409305863&alt=media)\n\n- Before sampling, the 2D Multiome Test_target matrix is 192 times larger than its CiteSeq equivalent.\n- After sampling the ratio dropped to 8.7\n\nHowever **the sampling has some very important implications on the scoring**\n\nThe submission file has 8.7 times more data for Multiome than for CiteSeq and we can at first glance think that Multiome is way more important for the scoring. However what really matters is the number of rows, because each row correlation has the same weight, even if the row is much longer!\nIf we look at the number of rows after sampling, suddenly CiteSeq is more important!! CiteSeq has all its 48,663 rows while Multiome is left with only 16,780. \nConsequently CiteSeq represent ~74% of the score and Multiome ~26% if we look at the number of rows (and number of correlations) used.\nThis can be checked roughly with the CV scores. CiteSeq CV is around 89 and Multiome 66.5. The LB score is closer to the CiteSeq CV score.\n\n***Amendment: [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933)***\nAs @ambrosm explained below, there was an issue in the initial data and the organizer is now ignoring 7,476 rows in the CITEseq submission. You can see the implications on the scoring weights [here.](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360180)",
      "votes": null
    },
    {
      "id": "1945405",
      "postDate": "09/19/2022 05:54:39",
      "content": "<p>With the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">data update of 2022-09-10</a>, the importance of CITEseq was reduced: The first 7476 rows of CITEseq test are ignored in scoring so that CITEseq effectively only has 48663 - 7476 = 41187 rows.</p>",
      "rawMarkdown": "With the [data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933), the importance of CITEseq was reduced: The first 7476 rows of CITEseq test are ignored in scoring so that CITEseq effectively only has 48663 - 7476 = 41187 rows.",
      "votes": null
    },
    {
      "id": "1945573",
      "postDate": "09/19/2022 08:22:07",
      "content": "<p>I see, I started working on the competition after this date and missed this issue. Will adjust my numbers asap. thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> for this and everything so valuable you have published !</p>",
      "rawMarkdown": "I see, I started working on the competition after this date and missed this issue. Will adjust my numbers asap. thanks @ambrosm for this and everything so valuable you have published !",
      "votes": null
    },
    {
      "id": "1980004",
      "postDate": "10/09/2022 23:17:23",
      "content": "<p>I would upvote this twice if I could.  I refer to it frequently!</p>",
      "rawMarkdown": "I would upvote this twice if I could.  I refer to it frequently!",
      "votes": null
    },
    {
      "id": "1981305",
      "postDate": "10/10/2022 18:06:21",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> . This is very encouraging. I did this because I was confused 😄 and then I wanted to share with everybody.</p>",
      "rawMarkdown": "Thank you very much @kirkdco . This is very encouraging. I did this because I was confused 😄 and then I wanted to share with everybody.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1945405,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "09/19/2022 05:54:39",
      "content": "<p>With the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">data update of 2022-09-10</a>, the importance of CITEseq was reduced: The first 7476 rows of CITEseq test are ignored in scoring so that CITEseq effectively only has 48663 - 7476 = 41187 rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1945573,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "09/19/2022 08:22:07",
          "content": "<p>I see, I started working on the competition after this date and missed this issue. Will adjust my numbers asap. thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> for this and everything so valuable you have published !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1980004,
      "author_name": "kirkdco",
      "author_url": "",
      "post_date": "10/09/2022 23:17:23",
      "content": "<p>I would upvote this twice if I could.  I refer to it frequently!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1981305,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "10/10/2022 18:06:21",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a> . This is very encouraging. I did this because I was confused 😄 and then I wanted to share with everybody.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1945054": "The following graph is intended to help visualize at once all the data names and shapes involved in the Multiome part of the competition.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fd81e9456a5a04d89da5f917feab412c7%2FMultiOmeShape.png?generation=1663530730501366&alt=media)\n\nWe have to predict Multiome Test_targets. However, we only need to submit part of it. 30% of the rows and for each row, 15% of the columns. It is the second part of the submission (58,931,360 row_ids in the submission file)\n\nWe can use this graph and the one for [CiteSeq](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/353522) to compare more specifically the Test_targets Data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2Fa490de664914da4481720d4415a68cde%2FTestTargetShapes.png?generation=1663562079905748&alt=media)\n\nThe organizer decided to sample Multiome Test_target, due to its large size. Only 30% of the rows (cells) are used for the scoring and for each of these selected rows, only 15% of the columns (genes) are selected.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F0eb956102a53b8b105466e255f749367%2FTestTargetRatios.png?generation=1663563409305863&alt=media)\n\n- Before sampling, the 2D Multiome Test_target matrix is 192 times larger than its CiteSeq equivalent.\n- After sampling the ratio dropped to 8.7\n\nHowever **the sampling has some very important implications on the scoring**\n\nThe submission file has 8.7 times more data for Multiome than for CiteSeq and we can at first glance think that Multiome is way more important for the scoring. However what really matters is the number of rows, because each row correlation has the same weight, even if the row is much longer!\nIf we look at the number of rows after sampling, suddenly CiteSeq is more important!! CiteSeq has all its 48,663 rows while Multiome is left with only 16,780. \nConsequently CiteSeq represent ~74% of the score and Multiome ~26% if we look at the number of rows (and number of correlations) used.\nThis can be checked roughly with the CV scores. CiteSeq CV is around 89 and Multiome 66.5. The LB score is closer to the CiteSeq CV score.\n\n***Amendment: [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933)***\nAs @ambrosm explained below, there was an issue in the initial data and the organizer is now ignoring 7,476 rows in the CITEseq submission. You can see the implications on the scoring weights [here.](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/360180)",
    "1945405": "With the [data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933), the importance of CITEseq was reduced: The first 7476 rows of CITEseq test are ignored in scoring so that CITEseq effectively only has 48663 - 7476 = 41187 rows.",
    "1945573": "I see, I started working on the competition after this date and missed this issue. Will adjust my numbers asap. thanks @ambrosm for this and everything so valuable you have published !",
    "1980004": "I would upvote this twice if I could.  I refer to it frequently!",
    "1981305": "Thank you very much @kirkdco . This is very encouraging. I did this because I was confused 😄 and then I wanted to share with everybody."
  },
  "source": "meta"
}