{
  "id": 360180,
  "title": "Cite / Multiome weights in the score",
  "url": "/competitions/open-problems-multimodal/discussion/360180",
  "author_name": "",
  "post_date": "2022-10-15T11:18:12.299652900Z",
  "votes": 18,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I would like first to thank <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> for pointing out to me the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a> and its implications in the scoring.</p>\n<p>This competition is actually <strong>2 competitions in one : Cite and Multiome.</strong> We only see the combined score in the leaderboard.<br>\nHowever we want to know what are the weights for each part in the scoring. Are they balanced? should we spend more time and effort on Multiome or Cite?<br>\nThe first idea would be to look at the ratio of data in the submission file : the first 6,812,820 values are for Cite and the remaining 58,931,360 values are for Multiome. So we may think at first glance that Multiome is 8.6 time more important than Cite. That is however <strong>completely wrong.</strong><br>\nWhat matters in the scoring is the number of biological cells. Each cell is a unit of scoring and the predictions on each cell are compared to the ground truth through a correlation score. <br>\n<strong>Each cell is a row</strong> in the test data. So let's count the rows!<br>\nHere the situation is complicated by the fact that there was an error in the data at the start of the competition and the error has been fixed: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a><br>\nWe will start by calculating the weights before the data update. We will then look at the implications of the data update.</p>\n<p><strong>1. BEFORE THE DATA UPDATE</strong><br>\nThe situation before the data update is relatively simple. <br>\nCite test has 48,663 rows.<br>\nMultiome test has 55,935 rows but only 30% are used. So  Multiome test has 16,780 rows in the score.<br>\nBefore the data update the weights are as follow :<br>\nCite : 48,663/(48,663 + 16,780) = 74.4%<br>\nMultiome : 16,780/(48,663+16,780) = 25.6%</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F59c621a5208c26c274b04a2e336c8758%2Fbeforeupdate.png?generation=1665818116598299&amp;alt=media\" alt=\"\"></p>\n<p>We can see that contrary to what we may think from the submission file, Cite is actually the most important in the scoring with the following weights:<br>\n<strong>Cite : 74.4% and Multiome 25.6%</strong></p>\n<p><em>The reason why Cite is more important despite occupying such a small ratio in the submission file is that Mutiome test has much longer rows. There are 3,512 columns after sampling in Multiome test while there are only 140 columns in Cite test. Multiome takes a lot of the submission file because it has many columns!</em></p>\n<p><strong>2. AFTER THE DATA UPDATE</strong><br>\nThe data Update had some (unintended) consequences on the weights. Here  We will carefully <strong>differentiate the Public and Private Leaderboard.</strong> We didn't have to do this before the data update as everything was the same in the Public and Private Leaderboard. But this is no longer the case.<br>\nIf you carefully read <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a> you notice that there is now 7,476 rows in Cite Test Public which are ignored. The key word here is <strong>Public.</strong> Regarding the Private Test, nothing changed and the weights we calculated are still valid.<br>\nThe Public/Private split is approximately 42%/58%</p>\n<p>Let's first look at the Private numbers :<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F85fb582b2f8fbf39ef39326bc745b3d0%2Fprivateafterupdate.png?generation=1665823538853702&amp;alt=media\" alt=\"\"></p>\n<p>We can see that the ratios are unchanged after the data update:<br>\n<strong>Private LB: Cite 74.4% and Multiome 25.6%</strong></p>\n<p>Let's now look at the Public numbers:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F14b5185761a8bbd6b97a9b6ed1bb73a0%2Fpublicafterupdate.png?generation=1665823568408391&amp;alt=media\" alt=\"\"><br>\nThere are 7,476 rows in the public Cite which are ignored in the scoring. This reduces the importance of Cite and the ratios are :<br>\n<strong>Public LB: Cite 64.8% and Multiome 35.2%</strong></p>\n<p><strong>CONCLUSION</strong><br>\nThe weights before and after the data update are as follow:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3e33995150dc09d104aa73fabcc905a0%2FSummary.png?generation=1665820339125058&amp;alt=media\" alt=\"\"></p>\n<p>There are 2 main implications.</p>\n<p><strong>Cite is more important in the Private Leaderbord than in the Public.</strong><br>\n If your Cite model is better than your Multiome (relatively to everybody else), expect a Shakeup. If your Multiome is better than your Cite, expect a Shakedown.</p>\n<p><strong>The scores in the Private Leaderboard will be higher than the score in the Public Leaderboard.</strong><br>\nSuppose your Cite score is 89.4 and your Multiome score is 66.5 (these values are reasonable if you look at your CV scores)<br>\nThen your Public and Private scores will be as follow:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3abdac0e960f9330f540b0cf73300e78%2FScoreExample.png?generation=1665821816330587&amp;alt=media\" alt=\"\"></p>\n<p>There was a drop in the scores after the update. This is due to 2 reasons. </p>\n<ul>\n<li>The weights changed (in the reverse way they will change when the Private LB will be unveiled)</li>\n<li>The perfect rows coming from the leak are no longer taken into account.</li>\n</ul>",
  "messages": [
    {
      "id": "1988523",
      "postDate": "10/15/2022 11:18:12",
      "content": "<p>I would like first to thank <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> for pointing out to me the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a> and its implications in the scoring.</p>\n<p>This competition is actually <strong>2 competitions in one : Cite and Multiome.</strong> We only see the combined score in the leaderboard.<br>\nHowever we want to know what are the weights for each part in the scoring. Are they balanced? should we spend more time and effort on Multiome or Cite?<br>\nThe first idea would be to look at the ratio of data in the submission file : the first 6,812,820 values are for Cite and the remaining 58,931,360 values are for Multiome. So we may think at first glance that Multiome is 8.6 time more important than Cite. That is however <strong>completely wrong.</strong><br>\nWhat matters in the scoring is the number of biological cells. Each cell is a unit of scoring and the predictions on each cell are compared to the ground truth through a correlation score. <br>\n<strong>Each cell is a row</strong> in the test data. So let's count the rows!<br>\nHere the situation is complicated by the fact that there was an error in the data at the start of the competition and the error has been fixed: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a><br>\nWe will start by calculating the weights before the data update. We will then look at the implications of the data update.</p>\n<p><strong>1. BEFORE THE DATA UPDATE</strong><br>\nThe situation before the data update is relatively simple. <br>\nCite test has 48,663 rows.<br>\nMultiome test has 55,935 rows but only 30% are used. So  Multiome test has 16,780 rows in the score.<br>\nBefore the data update the weights are as follow :<br>\nCite : 48,663/(48,663 + 16,780) = 74.4%<br>\nMultiome : 16,780/(48,663+16,780) = 25.6%</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F59c621a5208c26c274b04a2e336c8758%2Fbeforeupdate.png?generation=1665818116598299&amp;alt=media\" alt=\"\"></p>\n<p>We can see that contrary to what we may think from the submission file, Cite is actually the most important in the scoring with the following weights:<br>\n<strong>Cite : 74.4% and Multiome 25.6%</strong></p>\n<p><em>The reason why Cite is more important despite occupying such a small ratio in the submission file is that Mutiome test has much longer rows. There are 3,512 columns after sampling in Multiome test while there are only 140 columns in Cite test. Multiome takes a lot of the submission file because it has many columns!</em></p>\n<p><strong>2. AFTER THE DATA UPDATE</strong><br>\nThe data Update had some (unintended) consequences on the weights. Here  We will carefully <strong>differentiate the Public and Private Leaderboard.</strong> We didn't have to do this before the data update as everything was the same in the Public and Private Leaderboard. But this is no longer the case.<br>\nIf you carefully read <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933\" target=\"_blank\">The data update of 2022-09-10</a> you notice that there is now 7,476 rows in Cite Test Public which are ignored. The key word here is <strong>Public.</strong> Regarding the Private Test, nothing changed and the weights we calculated are still valid.<br>\nThe Public/Private split is approximately 42%/58%</p>\n<p>Let's first look at the Private numbers :<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F85fb582b2f8fbf39ef39326bc745b3d0%2Fprivateafterupdate.png?generation=1665823538853702&amp;alt=media\" alt=\"\"></p>\n<p>We can see that the ratios are unchanged after the data update:<br>\n<strong>Private LB: Cite 74.4% and Multiome 25.6%</strong></p>\n<p>Let's now look at the Public numbers:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F14b5185761a8bbd6b97a9b6ed1bb73a0%2Fpublicafterupdate.png?generation=1665823568408391&amp;alt=media\" alt=\"\"><br>\nThere are 7,476 rows in the public Cite which are ignored in the scoring. This reduces the importance of Cite and the ratios are :<br>\n<strong>Public LB: Cite 64.8% and Multiome 35.2%</strong></p>\n<p><strong>CONCLUSION</strong><br>\nThe weights before and after the data update are as follow:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3e33995150dc09d104aa73fabcc905a0%2FSummary.png?generation=1665820339125058&amp;alt=media\" alt=\"\"></p>\n<p>There are 2 main implications.</p>\n<p><strong>Cite is more important in the Private Leaderbord than in the Public.</strong><br>\n If your Cite model is better than your Multiome (relatively to everybody else), expect a Shakeup. If your Multiome is better than your Cite, expect a Shakedown.</p>\n<p><strong>The scores in the Private Leaderboard will be higher than the score in the Public Leaderboard.</strong><br>\nSuppose your Cite score is 89.4 and your Multiome score is 66.5 (these values are reasonable if you look at your CV scores)<br>\nThen your Public and Private scores will be as follow:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3abdac0e960f9330f540b0cf73300e78%2FScoreExample.png?generation=1665821816330587&amp;alt=media\" alt=\"\"></p>\n<p>There was a drop in the scores after the update. This is due to 2 reasons. </p>\n<ul>\n<li>The weights changed (in the reverse way they will change when the Private LB will be unveiled)</li>\n<li>The perfect rows coming from the leak are no longer taken into account.</li>\n</ul>",
      "rawMarkdown": "I would like first to thank @ambrosm for pointing out to me the [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933) and its implications in the scoring.\n\nThis competition is actually **2 competitions in one : Cite and Multiome.** We only see the combined score in the leaderboard.\nHowever we want to know what are the weights for each part in the scoring. Are they balanced? should we spend more time and effort on Multiome or Cite?\nThe first idea would be to look at the ratio of data in the submission file : the first 6,812,820 values are for Cite and the remaining 58,931,360 values are for Multiome. So we may think at first glance that Multiome is 8.6 time more important than Cite. That is however **completely wrong.**\nWhat matters in the scoring is the number of biological cells. Each cell is a unit of scoring and the predictions on each cell are compared to the ground truth through a correlation score. \n**Each cell is a row** in the test data. So let's count the rows!\nHere the situation is complicated by the fact that there was an error in the data at the start of the competition and the error has been fixed: [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933)\nWe will start by calculating the weights before the data update. We will then look at the implications of the data update.\n\n**1. BEFORE THE DATA UPDATE**\nThe situation before the data update is relatively simple. \nCite test has 48,663 rows.\nMultiome test has 55,935 rows but only 30% are used. So  Multiome test has 16,780 rows in the score.\nBefore the data update the weights are as follow :\nCite : 48,663/(48,663 + 16,780) = 74.4%\nMultiome : 16,780/(48,663+16,780) = 25.6%\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F59c621a5208c26c274b04a2e336c8758%2Fbeforeupdate.png?generation=1665818116598299&alt=media)\n\nWe can see that contrary to what we may think from the submission file, Cite is actually the most important in the scoring with the following weights:\n**Cite : 74.4% and Multiome 25.6%**\n\n*The reason why Cite is more important despite occupying such a small ratio in the submission file is that Mutiome test has much longer rows. There are 3,512 columns after sampling in Multiome test while there are only 140 columns in Cite test. Multiome takes a lot of the submission file because it has many columns!*\n\n**2. AFTER THE DATA UPDATE**\nThe data Update had some (unintended) consequences on the weights. Here  We will carefully **differentiate the Public and Private Leaderboard.** We didn't have to do this before the data update as everything was the same in the Public and Private Leaderboard. But this is no longer the case.\nIf you carefully read [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933) you notice that there is now 7,476 rows in Cite Test Public which are ignored. The key word here is **Public.** Regarding the Private Test, nothing changed and the weights we calculated are still valid.\nThe Public/Private split is approximately 42%/58%\n\nLet's first look at the Private numbers :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F85fb582b2f8fbf39ef39326bc745b3d0%2Fprivateafterupdate.png?generation=1665823538853702&alt=media)\n\nWe can see that the ratios are unchanged after the data update:\n**Private LB: Cite 74.4% and Multiome 25.6%**\n\nLet's now look at the Public numbers:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F14b5185761a8bbd6b97a9b6ed1bb73a0%2Fpublicafterupdate.png?generation=1665823568408391&alt=media)\nThere are 7,476 rows in the public Cite which are ignored in the scoring. This reduces the importance of Cite and the ratios are :\n**Public LB: Cite 64.8% and Multiome 35.2%**\n\n**CONCLUSION**\nThe weights before and after the data update are as follow:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3e33995150dc09d104aa73fabcc905a0%2FSummary.png?generation=1665820339125058&alt=media)\n\nThere are 2 main implications.\n\n**Cite is more important in the Private Leaderbord than in the Public.**\n If your Cite model is better than your Multiome (relatively to everybody else), expect a Shakeup. If your Multiome is better than your Cite, expect a Shakedown.\n\n**The scores in the Private Leaderboard will be higher than the score in the Public Leaderboard.**\nSuppose your Cite score is 89.4 and your Multiome score is 66.5 (these values are reasonable if you look at your CV scores)\nThen your Public and Private scores will be as follow:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3abdac0e960f9330f540b0cf73300e78%2FScoreExample.png?generation=1665821816330587&alt=media)\n\nThere was a drop in the scores after the update. This is due to 2 reasons. \n- The weights changed (in the reverse way they will change when the Private LB will be unveiled)\n- The perfect rows coming from the leak are no longer taken into account.",
      "votes": null
    },
    {
      "id": "1990773",
      "postDate": "10/16/2022 18:26:38",
      "content": "<p><a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a>  this is super helpful a novice question can u tell me the math on how u reached 81.33 :D . I am just not getting that number </p>",
      "rawMarkdown": "gehallak  this is super helpful a novice question can u tell me the math on how u reached 81.33 :D . I am just not getting that number",
      "votes": null
    },
    {
      "id": "1990857",
      "postDate": "10/16/2022 19:00:11",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> .  I am glad you found it useful. I applied the weights. 64.8% for Cite and 35.2% for Multiome.<br>\n89.40 * 0.648+ 66.50 * 0.352<br>\nTo be more precise, with one more decimal: 89.40 * 0.6478 + 66.50 * 0.3522=81.33</p>",
      "rawMarkdown": "Hi @gauravbrills .  I am glad you found it useful. I applied the weights. 64.8% for Cite and 35.2% for Multiome.\n89.40 * 0.648+ 66.50 * 0.352\nTo be more precise, with one more decimal: 89.40 * 0.6478 + 66.50 * 0.3522=81.33",
      "votes": null
    },
    {
      "id": "1990899",
      "postDate": "10/16/2022 19:56:05",
      "content": "<p>Ahh me silly multiplied in the wrong ratio 😃 thanks</p>",
      "rawMarkdown": "Ahh me silly multiplied in the wrong ratio 😃 thanks",
      "votes": null
    },
    {
      "id": "1991032",
      "postDate": "10/16/2022 23:08:06",
      "content": "<p><a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a> works well now just estimated scores now way better but it does give some correlation . May I ask the sample scores u selected are using which CV strategy ?</p>",
      "rawMarkdown": "gehallak works well now just estimated scores now way better but it does give some correlation . May I ask the sample scores u selected are using which CV strategy ?",
      "votes": null
    },
    {
      "id": "1991257",
      "postDate": "10/17/2022 04:23:39",
      "content": "<p><a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a>, Yes the correlation is good with these weights. For example, from the improvement on my Multiome CV score, I was able to estimate the improvement in the ranking on the LB. It is not easy because we don't see enough decimals on the LB but possible if you suppose the scores are evenly incrementing.<br>\nI tried Neural Network, Catboost, XGB, LGBM and Linear Regression. But the examples are fictitious.</p>",
      "rawMarkdown": "gauravbrills, Yes the correlation is good with these weights. For example, from the improvement on my Multiome CV score, I was able to estimate the improvement in the ranking on the LB. It is not easy because we don't see enough decimals on the LB but possible if you suppose the scores are evenly incrementing.\nI tried Neural Network, Catboost, XGB, LGBM and Linear Regression. But the examples are fictitious.",
      "votes": null
    },
    {
      "id": "1991279",
      "postDate": "10/17/2022 04:44:33",
      "content": "<p>I suppose CITE weights 0.661, MULTI weights 0.339 in submission correlation score<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359222\" target=\"_blank\">It is discussed here</a></p>",
      "rawMarkdown": "I suppose CITE weights 0.661, MULTI weights 0.339 in submission correlation score\n[It is discussed here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359222)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1990773,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "10/16/2022 18:26:38",
      "content": "<p><a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a>  this is super helpful a novice question can u tell me the math on how u reached 81.33 :D . I am just not getting that number </p>",
      "votes": null,
      "replies": [
        {
          "id": 1990857,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "10/16/2022 19:00:11",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> .  I am glad you found it useful. I applied the weights. 64.8% for Cite and 35.2% for Multiome.<br>\n89.40 * 0.648+ 66.50 * 0.352<br>\nTo be more precise, with one more decimal: 89.40 * 0.6478 + 66.50 * 0.3522=81.33</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1990899,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "10/16/2022 19:56:05",
          "content": "<p>Ahh me silly multiplied in the wrong ratio 😃 thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1991032,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "10/16/2022 23:08:06",
          "content": "<p><a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a> works well now just estimated scores now way better but it does give some correlation . May I ask the sample scores u selected are using which CV strategy ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1991257,
          "author_name": "gehallak",
          "author_url": "",
          "post_date": "10/17/2022 04:23:39",
          "content": "<p><a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a>, Yes the correlation is good with these weights. For example, from the improvement on my Multiome CV score, I was able to estimate the improvement in the ranking on the LB. It is not easy because we don't see enough decimals on the LB but possible if you suppose the scores are evenly incrementing.<br>\nI tried Neural Network, Catboost, XGB, LGBM and Linear Regression. But the examples are fictitious.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1991279,
      "author_name": "kaggledummie007",
      "author_url": "",
      "post_date": "10/17/2022 04:44:33",
      "content": "<p>I suppose CITE weights 0.661, MULTI weights 0.339 in submission correlation score<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359222\" target=\"_blank\">It is discussed here</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1988523": "I would like first to thank @ambrosm for pointing out to me the [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933) and its implications in the scoring.\n\nThis competition is actually **2 competitions in one : Cite and Multiome.** We only see the combined score in the leaderboard.\nHowever we want to know what are the weights for each part in the scoring. Are they balanced? should we spend more time and effort on Multiome or Cite?\nThe first idea would be to look at the ratio of data in the submission file : the first 6,812,820 values are for Cite and the remaining 58,931,360 values are for Multiome. So we may think at first glance that Multiome is 8.6 time more important than Cite. That is however **completely wrong.**\nWhat matters in the scoring is the number of biological cells. Each cell is a unit of scoring and the predictions on each cell are compared to the ground truth through a correlation score. \n**Each cell is a row** in the test data. So let's count the rows!\nHere the situation is complicated by the fact that there was an error in the data at the start of the competition and the error has been fixed: [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933)\nWe will start by calculating the weights before the data update. We will then look at the implications of the data update.\n\n**1. BEFORE THE DATA UPDATE**\nThe situation before the data update is relatively simple. \nCite test has 48,663 rows.\nMultiome test has 55,935 rows but only 30% are used. So  Multiome test has 16,780 rows in the score.\nBefore the data update the weights are as follow :\nCite : 48,663/(48,663 + 16,780) = 74.4%\nMultiome : 16,780/(48,663+16,780) = 25.6%\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F59c621a5208c26c274b04a2e336c8758%2Fbeforeupdate.png?generation=1665818116598299&alt=media)\n\nWe can see that contrary to what we may think from the submission file, Cite is actually the most important in the scoring with the following weights:\n**Cite : 74.4% and Multiome 25.6%**\n\n*The reason why Cite is more important despite occupying such a small ratio in the submission file is that Mutiome test has much longer rows. There are 3,512 columns after sampling in Multiome test while there are only 140 columns in Cite test. Multiome takes a lot of the submission file because it has many columns!*\n\n**2. AFTER THE DATA UPDATE**\nThe data Update had some (unintended) consequences on the weights. Here  We will carefully **differentiate the Public and Private Leaderboard.** We didn't have to do this before the data update as everything was the same in the Public and Private Leaderboard. But this is no longer the case.\nIf you carefully read [The data update of 2022-09-10](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350933) you notice that there is now 7,476 rows in Cite Test Public which are ignored. The key word here is **Public.** Regarding the Private Test, nothing changed and the weights we calculated are still valid.\nThe Public/Private split is approximately 42%/58%\n\nLet's first look at the Private numbers :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F85fb582b2f8fbf39ef39326bc745b3d0%2Fprivateafterupdate.png?generation=1665823538853702&alt=media)\n\nWe can see that the ratios are unchanged after the data update:\n**Private LB: Cite 74.4% and Multiome 25.6%**\n\nLet's now look at the Public numbers:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F14b5185761a8bbd6b97a9b6ed1bb73a0%2Fpublicafterupdate.png?generation=1665823568408391&alt=media)\nThere are 7,476 rows in the public Cite which are ignored in the scoring. This reduces the importance of Cite and the ratios are :\n**Public LB: Cite 64.8% and Multiome 35.2%**\n\n**CONCLUSION**\nThe weights before and after the data update are as follow:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3e33995150dc09d104aa73fabcc905a0%2FSummary.png?generation=1665820339125058&alt=media)\n\nThere are 2 main implications.\n\n**Cite is more important in the Private Leaderbord than in the Public.**\n If your Cite model is better than your Multiome (relatively to everybody else), expect a Shakeup. If your Multiome is better than your Cite, expect a Shakedown.\n\n**The scores in the Private Leaderboard will be higher than the score in the Public Leaderboard.**\nSuppose your Cite score is 89.4 and your Multiome score is 66.5 (these values are reasonable if you look at your CV scores)\nThen your Public and Private scores will be as follow:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9718188%2F3abdac0e960f9330f540b0cf73300e78%2FScoreExample.png?generation=1665821816330587&alt=media)\n\nThere was a drop in the scores after the update. This is due to 2 reasons. \n- The weights changed (in the reverse way they will change when the Private LB will be unveiled)\n- The perfect rows coming from the leak are no longer taken into account.",
    "1990773": "gehallak  this is super helpful a novice question can u tell me the math on how u reached 81.33 :D . I am just not getting that number",
    "1990857": "Hi @gauravbrills .  I am glad you found it useful. I applied the weights. 64.8% for Cite and 35.2% for Multiome.\n89.40 * 0.648+ 66.50 * 0.352\nTo be more precise, with one more decimal: 89.40 * 0.6478 + 66.50 * 0.3522=81.33",
    "1990899": "Ahh me silly multiplied in the wrong ratio 😃 thanks",
    "1991032": "gehallak works well now just estimated scores now way better but it does give some correlation . May I ask the sample scores u selected are using which CV strategy ?",
    "1991257": "gauravbrills, Yes the correlation is good with these weights. For example, from the improvement on my Multiome CV score, I was able to estimate the improvement in the ranking on the LB. It is not easy because we don't see enough decimals on the LB but possible if you suppose the scores are evenly incrementing.\nI tried Neural Network, Catboost, XGB, LGBM and Linear Regression. But the examples are fictitious.",
    "1991279": "I suppose CITE weights 0.661, MULTI weights 0.339 in submission correlation score\n[It is discussed here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/359222)"
  },
  "source": "meta"
}