{
  "id": 345889,
  "title": "Some questions about generating the prediction data",
  "url": "/competitions/open-problems-multimodal/discussion/345889",
  "author_name": "TESUZI",
  "post_date": "2022-08-17T01:50:07.643000",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, I notice that in the evaluation table we have cell id and gene id, and I think if we intend to generate our prediction results, wo should search the training dataset for cell id and gene id then make the prediction. Is it correct? Thanks a lot.</p>",
  "messages": [
    {
      "id": 1902905,
      "postDate": "2022-08-17T01:50:07.643Z",
      "content": "<p>Hi, I notice that in the evaluation table we have cell id and gene id, and I think if we intend to generate our prediction results, wo should search the training dataset for cell id and gene id then make the prediction. Is it correct? Thanks a lot.</p>",
      "rawMarkdown": "Hi, I notice that in the evaluation table we have cell id and gene id, and I think if we intend to generate our prediction results, wo should search the training dataset for cell id and gene id then make the prediction. Is it correct? Thanks a lot.",
      "votes": 3
    },
    {
      "id": 1903400,
      "postDate": "2022-08-17T11:52:41.277Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/llttyy\" target=\"_blank\">@llttyy</a> ,</p>\n<p>The cell_id and gene_id will be in the test data (not the training data), but yes, you only need to make predictions for the values in the evaluation table.</p>",
      "rawMarkdown": "Hi @llttyy ,\n\nThe cell_id and gene_id will be in the test data (not the training data), but yes, you only need to make predictions for the values in the evaluation table.",
      "votes": 1,
      "replies": [
        {
          "id": 1903453,
          "postDate": "2022-08-17T12:51:53.010Z",
          "content": "<p>I see, thanks a lot!</p>",
          "rawMarkdown": "I see, thanks a lot!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1903869,
      "postDate": "2022-08-17T18:00:12.120Z",
      "content": "<p>Hi, and to be more specific, I think we also need to subsample the test dataset to generate the submission file, is it correct? I think the total cell*gene number of test datasets is large than the length of the evaluation table. Is it correct? Thanks a lot.</p>",
      "rawMarkdown": "Hi, and to be more specific, I think we also need to subsample the test dataset to generate the submission file, is it correct? I think the total cell*gene number of test datasets is large than the length of the evaluation table. Is it correct? Thanks a lot.",
      "replies": [
        {
          "id": 1903872,
          "postDate": "2022-08-17T18:02:22.350Z",
          "content": "<p>Yes. You can get the details of the subsampling in the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">Data Documentation</a>. It's near the bottom.</p>",
          "rawMarkdown": "Yes. You can get the details of the subsampling in the [Data Documentation](https://www.kaggle.com/competitions/open-problems-multimodal/data). It's near the bottom.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1903712,
      "postDate": "2022-08-17T15:39:59.627Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1903400,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2022-08-17T11:52:41.277000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/llttyy\" target=\"_blank\">@llttyy</a> ,</p>\n<p>The cell_id and gene_id will be in the test data (not the training data), but yes, you only need to make predictions for the values in the evaluation table.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1903453,
          "author_name": "TESUZI",
          "author_url": "",
          "post_date": "2022-08-17T12:51:53.010000",
          "content": "<p>I see, thanks a lot!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1903869,
      "author_name": "TESUZI",
      "author_url": "",
      "post_date": "2022-08-17T18:00:12.120000",
      "content": "<p>Hi, and to be more specific, I think we also need to subsample the test dataset to generate the submission file, is it correct? I think the total cell*gene number of test datasets is large than the length of the evaluation table. Is it correct? Thanks a lot.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1903872,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2022-08-17T18:02:22.350000",
          "content": "<p>Yes. You can get the details of the subsampling in the <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/data\" target=\"_blank\">Data Documentation</a>. It's near the bottom.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1903712,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-17T15:39:59.627000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1902905": "Hi, I notice that in the evaluation table we have cell id and gene id, and I think if we intend to generate our prediction results, wo should search the training dataset for cell id and gene id then make the prediction. Is it correct? Thanks a lot.",
    "1903400": "Hi @llttyy ,\n\nThe cell_id and gene_id will be in the test data (not the training data), but yes, you only need to make predictions for the values in the evaluation table.",
    "1903869": "Hi, and to be more specific, I think we also need to subsample the test dataset to generate the submission file, is it correct? I think the total cell*gene number of test datasets is large than the length of the evaluation table. Is it correct? Thanks a lot.",
    "1903712": ""
  }
}