{
  "id": 443330,
  "title": "Question about SMILES in the test dataset",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/443330",
  "author_name": "",
  "post_date": "2023-09-26T15:42:15.392520500Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was thinking about using SMILES for prediction, but there doesn't seem to be a lookup table for <code>sm_name</code> -&gt; <code>SMILES</code>, other than in the <code>de_train</code> file. Does this mean that for the unseen test dataset, all of the <code>sm_name</code> values will also come from <code>de_train</code>? That way we can do a lookup, apply some function, and add this as a feature to the original train/test features that our ML model is based on. Is this a correct understanding?</p>",
  "messages": [
    {
      "id": "2457076",
      "postDate": "09/26/2023 15:42:15",
      "content": "<p>I was thinking about using SMILES for prediction, but there doesn't seem to be a lookup table for <code>sm_name</code> -&gt; <code>SMILES</code>, other than in the <code>de_train</code> file. Does this mean that for the unseen test dataset, all of the <code>sm_name</code> values will also come from <code>de_train</code>? That way we can do a lookup, apply some function, and add this as a feature to the original train/test features that our ML model is based on. Is this a correct understanding?</p>",
      "rawMarkdown": "I was thinking about using SMILES for prediction, but there doesn't seem to be a lookup table for `sm_name` -> `SMILES`, other than in the `de_train` file. Does this mean that for the unseen test dataset, all of the `sm_name` values will also come from `de_train`? That way we can do a lookup, apply some function, and add this as a feature to the original train/test features that our ML model is based on. Is this a correct understanding?",
      "votes": null
    },
    {
      "id": "2457125",
      "postDate": "09/26/2023 16:27:53",
      "content": "<p>Or am I just horribly confused? I'm admittedly quite a noob when it comes to Kaggle competitions. Is the <code>id_map</code> the entirety of the test data, and the leaderboard only shows a % of it? Sorry in advance.</p>",
      "rawMarkdown": "Or am I just horribly confused? I'm admittedly quite a noob when it comes to Kaggle competitions. Is the `id_map` the entirety of the test data, and the leaderboard only shows a % of it? Sorry in advance.",
      "votes": null
    },
    {
      "id": "2457512",
      "postDate": "09/27/2023 02:16:19",
      "content": "<ol>\n<li>The <code>de_train.csv</code> contains all the SMILES, so we can do a lookup to apply the <code>sm_name -&gt; SMILES</code> for the <code>id_map.csv</code>.</li>\n<li>The <code>id_map.csv</code> is the entirety of the test data, but the public leaderboard score only shows a part of it.</li>\n<li>You can find these infos in the <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Data page</a> or my <a href=\"https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda#Data-split\" target=\"_blank\">notebook</a></li>\n</ol>",
      "rawMarkdown": "1. The `de_train.csv` contains all the SMILES, so we can do a lookup to apply the `sm_name -> SMILES` for the `id_map.csv`.\n2. The `id_map.csv` is the entirety of the test data, but the public leaderboard score only shows a part of it.\n3. You can find these infos in the [Data page](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data) or my [notebook](https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda#Data-split)",
      "votes": null
    },
    {
      "id": "2457553",
      "postDate": "09/27/2023 03:26:34",
      "content": "<p>Thanks for clarifying!</p>",
      "rawMarkdown": "Thanks for clarifying!",
      "votes": null
    },
    {
      "id": "2514199",
      "postDate": "11/06/2023 05:50:14",
      "content": "<p>Thanks for your reply, I tried to do so, and I found that only 129 sm_names in id_map are in the de_train.csv</p>",
      "rawMarkdown": "Thanks for your reply, I tried to do so, and I found that only 129 sm_names in id_map are in the de_train.csv",
      "votes": null
    },
    {
      "id": "2514209",
      "postDate": "11/06/2023 05:59:45",
      "content": "<p>I tried using de_train as a lookup table, but I found that id_map is not entirely lied in de_train, only 129 out of all are in de_train. Have you solve this issue?</p>",
      "rawMarkdown": "I tried using de_train as a lookup table, but I found that id_map is not entirely lied in de_train, only 129 out of all are in de_train. Have you solve this issue?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2457125,
      "author_name": "yuqizheng",
      "author_url": "",
      "post_date": "09/26/2023 16:27:53",
      "content": "<p>Or am I just horribly confused? I'm admittedly quite a noob when it comes to Kaggle competitions. Is the <code>id_map</code> the entirety of the test data, and the leaderboard only shows a % of it? Sorry in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2514209,
          "author_name": "ruiqizhuangricky",
          "author_url": "",
          "post_date": "11/06/2023 05:59:45",
          "content": "<p>I tried using de_train as a lookup table, but I found that id_map is not entirely lied in de_train, only 129 out of all are in de_train. Have you solve this issue?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2457512,
      "author_name": "awater1223",
      "author_url": "",
      "post_date": "09/27/2023 02:16:19",
      "content": "<ol>\n<li>The <code>de_train.csv</code> contains all the SMILES, so we can do a lookup to apply the <code>sm_name -&gt; SMILES</code> for the <code>id_map.csv</code>.</li>\n<li>The <code>id_map.csv</code> is the entirety of the test data, but the public leaderboard score only shows a part of it.</li>\n<li>You can find these infos in the <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Data page</a> or my <a href=\"https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda#Data-split\" target=\"_blank\">notebook</a></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2457553,
          "author_name": "yuqizheng",
          "author_url": "",
          "post_date": "09/27/2023 03:26:34",
          "content": "<p>Thanks for clarifying!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2514199,
          "author_name": "ruiqizhuangricky",
          "author_url": "",
          "post_date": "11/06/2023 05:50:14",
          "content": "<p>Thanks for your reply, I tried to do so, and I found that only 129 sm_names in id_map are in the de_train.csv</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2457076": "I was thinking about using SMILES for prediction, but there doesn't seem to be a lookup table for `sm_name` -> `SMILES`, other than in the `de_train` file. Does this mean that for the unseen test dataset, all of the `sm_name` values will also come from `de_train`? That way we can do a lookup, apply some function, and add this as a feature to the original train/test features that our ML model is based on. Is this a correct understanding?",
    "2457125": "Or am I just horribly confused? I'm admittedly quite a noob when it comes to Kaggle competitions. Is the `id_map` the entirety of the test data, and the leaderboard only shows a % of it? Sorry in advance.",
    "2457512": "1. The `de_train.csv` contains all the SMILES, so we can do a lookup to apply the `sm_name -> SMILES` for the `id_map.csv`.\n2. The `id_map.csv` is the entirety of the test data, but the public leaderboard score only shows a part of it.\n3. You can find these infos in the [Data page](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data) or my [notebook](https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda#Data-split)",
    "2457553": "Thanks for clarifying!",
    "2514199": "Thanks for your reply, I tried to do so, and I found that only 129 sm_names in id_map are in the de_train.csv",
    "2514209": "I tried using de_train as a lookup table, but I found that id_map is not entirely lied in de_train, only 129 out of all are in de_train. Have you solve this issue?"
  },
  "source": "meta"
}