{
  "id": 224100,
  "title": "Impact of Synthetic Data",
  "url": "/competitions/bms-molecular-translation/discussion/224100",
  "author_name": "",
  "post_date": "2021-03-06T23:23:39.696381200Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm curious how people feel about the role synthetic data will play in this competition. I made a quick <a href=\"https://www.kaggle.com/towardsentropy/create-additional-data\" target=\"_blank\">notebook</a> on how you can create Image/InChi pairs from SMILES strings with RDKit.</p>\n<p>While the synthetic data doesn't have the same \"scanned\" appearance as the competition data (at least until someone finds the right image processing), the similarity is close enough to be a major factor. </p>\n<p>Existing compound databases have billions of SMILES strings ready to do. Add in some combichem and there's really no limit to the amount of synthetic data you can create. Will this competition be decided by who invests the most time in creating synthetic data?</p>",
  "messages": [
    {
      "id": "1228968",
      "postDate": "03/06/2021 23:23:39",
      "content": "<p>I'm curious how people feel about the role synthetic data will play in this competition. I made a quick <a href=\"https://www.kaggle.com/towardsentropy/create-additional-data\" target=\"_blank\">notebook</a> on how you can create Image/InChi pairs from SMILES strings with RDKit.</p>\n<p>While the synthetic data doesn't have the same \"scanned\" appearance as the competition data (at least until someone finds the right image processing), the similarity is close enough to be a major factor. </p>\n<p>Existing compound databases have billions of SMILES strings ready to do. Add in some combichem and there's really no limit to the amount of synthetic data you can create. Will this competition be decided by who invests the most time in creating synthetic data?</p>",
      "rawMarkdown": "I'm curious how people feel about the role synthetic data will play in this competition. I made a quick [notebook](https://www.kaggle.com/towardsentropy/create-additional-data) on how you can create Image/InChi pairs from SMILES strings with RDKit.\n\nWhile the synthetic data doesn't have the same \"scanned\" appearance as the competition data (at least until someone finds the right image processing), the similarity is close enough to be a major factor. \n\nExisting compound databases have billions of SMILES strings ready to do. Add in some combichem and there's really no limit to the amount of synthetic data you can create. Will this competition be decided by who invests the most time in creating synthetic data?",
      "votes": null
    },
    {
      "id": "1229065",
      "postDate": "03/07/2021 03:49:03",
      "content": "<p>Since you mentioned virtually unlimited synthetic data, maybe it can be used as part of a pre-training strategy with the competition morphed like a downstream task. The different distribution might be an issue but I think there are approaches out there that work fine if fine-tuned properly. Thoughts?</p>",
      "rawMarkdown": "Since you mentioned virtually unlimited synthetic data, maybe it can be used as part of a pre-training strategy with the competition morphed like a downstream task. The different distribution might be an issue but I think there are approaches out there that work fine if fine-tuned properly. Thoughts?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1229065,
      "author_name": "arka47",
      "author_url": "",
      "post_date": "03/07/2021 03:49:03",
      "content": "<p>Since you mentioned virtually unlimited synthetic data, maybe it can be used as part of a pre-training strategy with the competition morphed like a downstream task. The different distribution might be an issue but I think there are approaches out there that work fine if fine-tuned properly. Thoughts?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1228968": "I'm curious how people feel about the role synthetic data will play in this competition. I made a quick [notebook](https://www.kaggle.com/towardsentropy/create-additional-data) on how you can create Image/InChi pairs from SMILES strings with RDKit.\n\nWhile the synthetic data doesn't have the same \"scanned\" appearance as the competition data (at least until someone finds the right image processing), the similarity is close enough to be a major factor. \n\nExisting compound databases have billions of SMILES strings ready to do. Add in some combichem and there's really no limit to the amount of synthetic data you can create. Will this competition be decided by who invests the most time in creating synthetic data?",
    "1229065": "Since you mentioned virtually unlimited synthetic data, maybe it can be used as part of a pre-training strategy with the competition morphed like a downstream task. The different distribution might be an issue but I think there are approaches out there that work fine if fine-tuned properly. Thoughts?"
  },
  "source": "meta"
}