{
  "id": 350402,
  "title": "What is the reason to use only 30% of test data for prediction / submission?",
  "url": "/competitions/open-problems-multimodal/discussion/350402",
  "author_name": "",
  "post_date": "2022-09-05T14:59:24.547646800Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The subject to discuss is prediction for MULTIOME part of data.</p>\n<p><a href=\"http://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">It is explicitly stated</a> that we should predict using only 30% of rows in test dataset:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9150719%2F018bfe51a19e73bd00940a9a024b34f0%2F1.jpg?generation=1662389609974489&amp;alt=media\" alt=\"\"></p>\n<p>And it really is.<br>\nFile “test_multi_inputs.h5” contains 55.935 rows each one for unique cell. But “sample_submission.csv” and “evaluation_ids.csv” links only to 16.780 cell_id’s for MULTIOME data.</p>\n<p>I am not experienced with Kaggle and DS but this looks strange for me.<br>\nI have not seen something like that in any over competition or lesson.<br>\nUsually we have to make predictions using all the test data.<br>\nAt least all rows of it.</p>\n<p>So I have 2 questions.</p>\n<ol>\n<li>Why not to limit test dataset to these 30% of rows and use other 70% for training?</li>\n<li>Is this a hint that the prediction should take into account all data in the test dataset, and not just use 30% of the rows?<br>\nWhat algorithms can do this?</li>\n</ol>",
  "messages": [
    {
      "id": "1927329",
      "postDate": "09/05/2022 14:59:24",
      "content": "<p>The subject to discuss is prediction for MULTIOME part of data.</p>\n<p><a href=\"http://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">It is explicitly stated</a> that we should predict using only 30% of rows in test dataset:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9150719%2F018bfe51a19e73bd00940a9a024b34f0%2F1.jpg?generation=1662389609974489&amp;alt=media\" alt=\"\"></p>\n<p>And it really is.<br>\nFile “test_multi_inputs.h5” contains 55.935 rows each one for unique cell. But “sample_submission.csv” and “evaluation_ids.csv” links only to 16.780 cell_id’s for MULTIOME data.</p>\n<p>I am not experienced with Kaggle and DS but this looks strange for me.<br>\nI have not seen something like that in any over competition or lesson.<br>\nUsually we have to make predictions using all the test data.<br>\nAt least all rows of it.</p>\n<p>So I have 2 questions.</p>\n<ol>\n<li>Why not to limit test dataset to these 30% of rows and use other 70% for training?</li>\n<li>Is this a hint that the prediction should take into account all data in the test dataset, and not just use 30% of the rows?<br>\nWhat algorithms can do this?</li>\n</ol>",
      "rawMarkdown": "The subject to discuss is prediction for MULTIOME part of data.\n\n[It is explicitly stated](http://www.kaggle.com/competitions/open-problems-multimodal/data) that we should predict using only 30% of rows in test dataset:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9150719%2F018bfe51a19e73bd00940a9a024b34f0%2F1.jpg?generation=1662389609974489&alt=media)\n\nAnd it really is.\nFile “test_multi_inputs.h5” contains 55.935 rows each one for unique cell. But “sample_submission.csv” and “evaluation_ids.csv” links only to 16.780 cell_id’s for MULTIOME data.\n\nI am not experienced with Kaggle and DS but this looks strange for me.\nI have not seen something like that in any over competition or lesson.\nUsually we have to make predictions using all the test data.\nAt least all rows of it.\n\nSo I have 2 questions.\n1. Why not to limit test dataset to these 30% of rows and use other 70% for training?\n2. Is this a hint that the prediction should take into account all data in the test dataset, and not just use 30% of the rows?\nWhat algorithms can do this?",
      "votes": null
    },
    {
      "id": "1928185",
      "postDate": "09/06/2022 10:02:54",
      "content": "<p>As I understand for the 1st question, if the 70% will be used for training, the model generalization (between donors and days) would not be measured as supposed by organizers.<br>\nFor the 2nd, for example, you can use all the data for correction for batch effect. It's the first idea, that comes to my mind.</p>",
      "rawMarkdown": "As I understand for the 1st question, if the 70% will be used for training, the model generalization (between donors and days) would not be measured as supposed by organizers.\nFor the 2nd, for example, you can use all the data for correction for batch effect. It's the first idea, that comes to my mind.",
      "votes": null
    },
    {
      "id": "1928519",
      "postDate": "09/06/2022 13:59:44",
      "content": "<p>I don't think it is about defining the training and testing set - it is to reduce the number of predicted data points to a sampled subset that need to be submitted for scoring.</p>\n<p>On the other hand, what I don't quite see is how this scoring is balanced. There are only 140 predicted outcomes for cite (the RNA to PROTEIN part), while there are over 20,000 predictions for multi (DNA to RNA).</p>\n<p>As I understand it, these are actually two challenges merged into one score. It will be highly biased (by more than a factor of a hundred) toward the prediction of RNA from DNA, up to a point that one might want to focus only on this part. </p>\n<p>Or am I missing something?</p>",
      "rawMarkdown": "I don't think it is about defining the training and testing set - it is to reduce the number of predicted data points to a sampled subset that need to be submitted for scoring.\n\nOn the other hand, what I don't quite see is how this scoring is balanced. There are only 140 predicted outcomes for cite (the RNA to PROTEIN part), while there are over 20,000 predictions for multi (DNA to RNA).\n\nAs I understand it, these are actually two challenges merged into one score. It will be highly biased (by more than a factor of a hundred) toward the prediction of RNA from DNA, up to a point that one might want to focus only on this part. \n\nOr am I missing something?",
      "votes": null
    },
    {
      "id": "1929151",
      "postDate": "09/06/2022 21:20:22",
      "content": "<p>you're missing that the pearson correlation is per row, so it does not matter the number of columns :)</p>",
      "rawMarkdown": "you're missing that the pearson correlation is per row, so it does not matter the number of columns :)",
      "votes": null
    },
    {
      "id": "1929679",
      "postDate": "09/07/2022 09:17:23",
      "content": "<p>No - that is not the case. </p>\n<p>I ran the below tests. </p>\n<p>It is clear that the predictions on cite weigh ten times more than the predictions on multi. </p>\n<p>Here is why:</p>\n<p>(1) Count the number of occurrences of of \"ENSG\" in evaluation_ids.csv (these are the gene ids).<br>\nThere are 58,931,360 gene ids out within the 65,744,180 (89.6%) evaluated data points, but only 6,812,820 protein ids.</p>\n<p>(2) To check this on the real evaluation of the kaggle site, I ran the simple test case of predicting all outcomes by their means. The score is 0.741. Then I randomized only the order of multi data predictions (means) and left the cite data intact. The score was 0.584. Then I randomized only the cite data and left the multi data intact. The score was 0.133.</p>\n<p>From this it is clear that the cite data dominates the evaluation!</p>",
      "rawMarkdown": "No - that is not the case. \n\nI ran the below tests. \n\nIt is clear that the predictions on cite weigh ten times more than the predictions on multi. \n\nHere is why:\n\n(1) Count the number of occurrences of of \"ENSG\" in evaluation_ids.csv (these are the gene ids).\nThere are 58,931,360 gene ids out within the 65,744,180 (89.6%) evaluated data points, but only 6,812,820 protein ids.\n\n(2) To check this on the real evaluation of the kaggle site, I ran the simple test case of predicting all outcomes by their means. The score is 0.741. Then I randomized only the order of multi data predictions (means) and left the cite data intact. The score was 0.584. Then I randomized only the cite data and left the multi data intact. The score was 0.133.\n\nFrom this it is clear that the cite data dominates the evaluation!",
      "votes": null
    },
    {
      "id": "1929686",
      "postDate": "09/07/2022 09:22:53",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349583\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349583</a></p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349583",
      "votes": null
    },
    {
      "id": "1929805",
      "postDate": "09/07/2022 11:28:40",
      "content": "<p>here is stated that CITE vs MULTI weights are near 3:1<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349591\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349591</a></p>",
      "rawMarkdown": "here is stated that CITE vs MULTI weights are near 3:1\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/349591",
      "votes": null
    },
    {
      "id": "1986332",
      "postDate": "10/14/2022 03:02:00",
      "content": "<p>I was wondering the same question. The organizers could simply have removed the 70% rows of the test data which are never used. My guess was that historically, they first wanted to use the whole test data, but it was taking too long to score at each submission, or the submission file was too big to upload ?</p>",
      "rawMarkdown": "I was wondering the same question. The organizers could simply have removed the 70% rows of the test data which are never used. My guess was that historically, they first wanted to use the whole test data, but it was taking too long to score at each submission, or the submission file was too big to upload ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1928185,
      "author_name": "kseniyapetrova",
      "author_url": "",
      "post_date": "09/06/2022 10:02:54",
      "content": "<p>As I understand for the 1st question, if the 70% will be used for training, the model generalization (between donors and days) would not be measured as supposed by organizers.<br>\nFor the 2nd, for example, you can use all the data for correction for batch effect. It's the first idea, that comes to my mind.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1928519,
      "author_name": "ksuhre",
      "author_url": "",
      "post_date": "09/06/2022 13:59:44",
      "content": "<p>I don't think it is about defining the training and testing set - it is to reduce the number of predicted data points to a sampled subset that need to be submitted for scoring.</p>\n<p>On the other hand, what I don't quite see is how this scoring is balanced. There are only 140 predicted outcomes for cite (the RNA to PROTEIN part), while there are over 20,000 predictions for multi (DNA to RNA).</p>\n<p>As I understand it, these are actually two challenges merged into one score. It will be highly biased (by more than a factor of a hundred) toward the prediction of RNA from DNA, up to a point that one might want to focus only on this part. </p>\n<p>Or am I missing something?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1929151,
          "author_name": "thomasuriot",
          "author_url": "",
          "post_date": "09/06/2022 21:20:22",
          "content": "<p>you're missing that the pearson correlation is per row, so it does not matter the number of columns :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1929679,
          "author_name": "ksuhre",
          "author_url": "",
          "post_date": "09/07/2022 09:17:23",
          "content": "<p>No - that is not the case. </p>\n<p>I ran the below tests. </p>\n<p>It is clear that the predictions on cite weigh ten times more than the predictions on multi. </p>\n<p>Here is why:</p>\n<p>(1) Count the number of occurrences of of \"ENSG\" in evaluation_ids.csv (these are the gene ids).<br>\nThere are 58,931,360 gene ids out within the 65,744,180 (89.6%) evaluated data points, but only 6,812,820 protein ids.</p>\n<p>(2) To check this on the real evaluation of the kaggle site, I ran the simple test case of predicting all outcomes by their means. The score is 0.741. Then I randomized only the order of multi data predictions (means) and left the cite data intact. The score was 0.584. Then I randomized only the cite data and left the multi data intact. The score was 0.133.</p>\n<p>From this it is clear that the cite data dominates the evaluation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1929686,
          "author_name": "thomasuriot",
          "author_url": "",
          "post_date": "09/07/2022 09:22:53",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349583\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349583</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1929805,
          "author_name": "kaggledummie007",
          "author_url": "",
          "post_date": "09/07/2022 11:28:40",
          "content": "<p>here is stated that CITE vs MULTI weights are near 3:1<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349591\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349591</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1986332,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "10/14/2022 03:02:00",
      "content": "<p>I was wondering the same question. The organizers could simply have removed the 70% rows of the test data which are never used. My guess was that historically, they first wanted to use the whole test data, but it was taking too long to score at each submission, or the submission file was too big to upload ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1927329": "The subject to discuss is prediction for MULTIOME part of data.\n\n[It is explicitly stated](http://www.kaggle.com/competitions/open-problems-multimodal/data) that we should predict using only 30% of rows in test dataset:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9150719%2F018bfe51a19e73bd00940a9a024b34f0%2F1.jpg?generation=1662389609974489&alt=media)\n\nAnd it really is.\nFile “test_multi_inputs.h5” contains 55.935 rows each one for unique cell. But “sample_submission.csv” and “evaluation_ids.csv” links only to 16.780 cell_id’s for MULTIOME data.\n\nI am not experienced with Kaggle and DS but this looks strange for me.\nI have not seen something like that in any over competition or lesson.\nUsually we have to make predictions using all the test data.\nAt least all rows of it.\n\nSo I have 2 questions.\n1. Why not to limit test dataset to these 30% of rows and use other 70% for training?\n2. Is this a hint that the prediction should take into account all data in the test dataset, and not just use 30% of the rows?\nWhat algorithms can do this?",
    "1928185": "As I understand for the 1st question, if the 70% will be used for training, the model generalization (between donors and days) would not be measured as supposed by organizers.\nFor the 2nd, for example, you can use all the data for correction for batch effect. It's the first idea, that comes to my mind.",
    "1928519": "I don't think it is about defining the training and testing set - it is to reduce the number of predicted data points to a sampled subset that need to be submitted for scoring.\n\nOn the other hand, what I don't quite see is how this scoring is balanced. There are only 140 predicted outcomes for cite (the RNA to PROTEIN part), while there are over 20,000 predictions for multi (DNA to RNA).\n\nAs I understand it, these are actually two challenges merged into one score. It will be highly biased (by more than a factor of a hundred) toward the prediction of RNA from DNA, up to a point that one might want to focus only on this part. \n\nOr am I missing something?",
    "1929151": "you're missing that the pearson correlation is per row, so it does not matter the number of columns :)",
    "1929679": "No - that is not the case. \n\nI ran the below tests. \n\nIt is clear that the predictions on cite weigh ten times more than the predictions on multi. \n\nHere is why:\n\n(1) Count the number of occurrences of of \"ENSG\" in evaluation_ids.csv (these are the gene ids).\nThere are 58,931,360 gene ids out within the 65,744,180 (89.6%) evaluated data points, but only 6,812,820 protein ids.\n\n(2) To check this on the real evaluation of the kaggle site, I ran the simple test case of predicting all outcomes by their means. The score is 0.741. Then I randomized only the order of multi data predictions (means) and left the cite data intact. The score was 0.584. Then I randomized only the cite data and left the multi data intact. The score was 0.133.\n\nFrom this it is clear that the cite data dominates the evaluation!",
    "1929686": "https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349583",
    "1929805": "here is stated that CITE vs MULTI weights are near 3:1\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/349591",
    "1986332": "I was wondering the same question. The organizers could simply have removed the 70% rows of the test data which are never used. My guess was that historically, they first wanted to use the whole test data, but it was taking too long to score at each submission, or the submission file was too big to upload ?"
  },
  "source": "meta"
}