{
  "id": 446397,
  "title": "Anyone tried cell data?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/446397",
  "author_name": "",
  "post_date": "2023-10-11T14:06:59.855532300Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello,  friends!</p>\n<p>Has anyone explored incorporating cell line data or other auxiliary data sources, alongside complex ensemble models and fine-tuning techniques, to improve model performance on the public test dataset? If you have ventured into this, I'd greatly appreciate learning from your insights and experiences.</p>\n<p>Thank you for sharing your knowledge.</p>\n<p>Best regards.</p>",
  "messages": [
    {
      "id": "2477709",
      "postDate": "10/11/2023 14:06:59",
      "content": "<p>Hello,  friends!</p>\n<p>Has anyone explored incorporating cell line data or other auxiliary data sources, alongside complex ensemble models and fine-tuning techniques, to improve model performance on the public test dataset? If you have ventured into this, I'd greatly appreciate learning from your insights and experiences.</p>\n<p>Thank you for sharing your knowledge.</p>\n<p>Best regards.</p>",
      "rawMarkdown": "Hello,  friends!\n\nHas anyone explored incorporating cell line data or other auxiliary data sources, alongside complex ensemble models and fine-tuning techniques, to improve model performance on the public test dataset? If you have ventured into this, I'd greatly appreciate learning from your insights and experiences.\n\nThank you for sharing your knowledge.\n\nBest regards.",
      "votes": null
    },
    {
      "id": "2478167",
      "postDate": "10/11/2023 19:35:53",
      "content": "<p>Not yet, I am trying various techniques to extract some features from the SMILES, but there are still no results.🤦‍♂️</p>",
      "rawMarkdown": "Not yet, I am trying various techniques to extract some features from the SMILES, but there are still no results.🤦‍♂️",
      "votes": null
    },
    {
      "id": "2478491",
      "postDate": "10/12/2023 03:42:31",
      "content": "<p>Awesome ! Have you considered using SMILES embeddings or leveraging pre-trained models for this?  Give it a try and let me know how it goes! Thanks for sharing your progress. 🙌</p>",
      "rawMarkdown": "Awesome ! Have you considered using SMILES embeddings or leveraging pre-trained models for this?  Give it a try and let me know how it goes! Thanks for sharing your progress. 🙌",
      "votes": null
    },
    {
      "id": "2478655",
      "postDate": "10/12/2023 06:18:58",
      "content": "<p>Yes, I tried with two strategies. The first one was Morgan fingerprints at 2048 bits. Initially, it seemed promising, but I had significant overfitting issues. Now, I'm experimenting with embeddings using Mol2Vec and have managed to reduce the RMSE a bit. Over the weekend, I'll play around with keras tuner and see how it goes </p>",
      "rawMarkdown": "Yes, I tried with two strategies. The first one was Morgan fingerprints at 2048 bits. Initially, it seemed promising, but I had significant overfitting issues. Now, I'm experimenting with embeddings using Mol2Vec and have managed to reduce the RMSE a bit. Over the weekend, I'll play around with keras tuner and see how it goes",
      "votes": null
    },
    {
      "id": "2479055",
      "postDate": "10/12/2023 11:30:12",
      "content": "<p>All the best!</p>",
      "rawMarkdown": "All the best!",
      "votes": null
    },
    {
      "id": "2479759",
      "postDate": "10/12/2023 20:51:28",
      "content": "<p>I am trying to build a two stage model, with one stage predicting the 978 landmark genes shared between the Kaggle data and the <a href=\"https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE70138\" target=\"_blank\">LINCS</a> across 107,404 samples, and then a second model predicting the full transcriptome using just the Kaggle data.<br>\nI haven't had much success, but maybe I haven't been using the right model.</p>\n<p>I made the basics of the <a href=\"https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation\" target=\"_blank\">dataset preparation</a> public and would be happy to talk more.</p>",
      "rawMarkdown": "I am trying to build a two stage model, with one stage predicting the 978 landmark genes shared between the Kaggle data and the [LINCS](https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE70138) across 107,404 samples, and then a second model predicting the full transcriptome using just the Kaggle data.\nI haven't had much success, but maybe I haven't been using the right model.\n\nI made the basics of the [dataset preparation](https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation) public and would be happy to talk more.",
      "votes": null
    },
    {
      "id": "2480011",
      "postDate": "10/13/2023 04:40:06",
      "content": "<p>Great progress! It's wonderful that you've shared your dataset preparation work. Building a two-stage model sounds like a promising approach.  Integrating predictions for the landmark genes from Kaggle and LINCS data across a significant number of samples is a complex task. I hope it ll go well.</p>\n<p>I tried pulling some chemical compound data from <a href=\"https://lincsportal.ccs.miami.edu/sigc-api/swagger-ui.html\" target=\"_blank\">lincs portal</a> and still processing it. </p>\n<p>Keep up the good work, and feel free to share any updates or challenges you face. Thank you for sharing your progress.</p>",
      "rawMarkdown": "Great progress! It's wonderful that you've shared your dataset preparation work. Building a two-stage model sounds like a promising approach.  Integrating predictions for the landmark genes from Kaggle and LINCS data across a significant number of samples is a complex task. I hope it ll go well.\n\nI tried pulling some chemical compound data from [lincs portal](https://lincsportal.ccs.miami.edu/sigc-api/swagger-ui.html) and still processing it. \n\nKeep up the good work, and feel free to share any updates or challenges you face. Thank you for sharing your progress.",
      "votes": null
    },
    {
      "id": "2486087",
      "postDate": "10/17/2023 17:09:00",
      "content": "<p>I have generated embeddings for the SMILE columns using DeepChem/ChemBERTa-77M-MTR (<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550#2485877\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550#2485877</a>) and I am now in the process of training a neural network based on this (<a href=\"https://www.kaggle.com/code/kishanvavdara/neural-network-regression)\" target=\"_blank\">https://www.kaggle.com/code/kishanvavdara/neural-network-regression)</a>. Then I would like to try XGBoost as well. Still quite a lot of things to figure out, but moving forward.</p>",
      "rawMarkdown": "I have generated embeddings for the SMILE columns using DeepChem/ChemBERTa-77M-MTR (https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550#2485877) and I am now in the process of training a neural network based on this (https://www.kaggle.com/code/kishanvavdara/neural-network-regression). Then I would like to try XGBoost as well. Still quite a lot of things to figure out, but moving forward.",
      "votes": null
    },
    {
      "id": "2486711",
      "postDate": "10/18/2023 05:56:46",
      "content": "<p>Great! I ll also try doing that. All the best <a href=\"https://www.kaggle.com/wguesdon\" target=\"_blank\">@wguesdon</a> Hope you get good results!</p>",
      "rawMarkdown": "Great! I ll also try doing that. All the best @wguesdon Hope you get good results!",
      "votes": null
    },
    {
      "id": "2487059",
      "postDate": "10/18/2023 10:44:25",
      "content": "<p>Thanks :) Does this improve model performance for you?<br>\nIt seems to decrease performance so far compared to not using embedding.  I am now trying to apply PCA or other feature engineering methods to those embeddings. </p>",
      "rawMarkdown": "Thanks :) Does this improve model performance for you?\nIt seems to decrease performance so far compared to not using embedding.  I am now trying to apply PCA or other feature engineering methods to those embeddings.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2478167,
      "author_name": "ermalossi",
      "author_url": "",
      "post_date": "10/11/2023 19:35:53",
      "content": "<p>Not yet, I am trying various techniques to extract some features from the SMILES, but there are still no results.🤦‍♂️</p>",
      "votes": null,
      "replies": [
        {
          "id": 2478491,
          "author_name": "kishanvavdara",
          "author_url": "",
          "post_date": "10/12/2023 03:42:31",
          "content": "<p>Awesome ! Have you considered using SMILES embeddings or leveraging pre-trained models for this?  Give it a try and let me know how it goes! Thanks for sharing your progress. 🙌</p>",
          "votes": null,
          "replies": [
            {
              "id": 2478655,
              "author_name": "ermalossi",
              "author_url": "",
              "post_date": "10/12/2023 06:18:58",
              "content": "<p>Yes, I tried with two strategies. The first one was Morgan fingerprints at 2048 bits. Initially, it seemed promising, but I had significant overfitting issues. Now, I'm experimenting with embeddings using Mol2Vec and have managed to reduce the RMSE a bit. Over the weekend, I'll play around with keras tuner and see how it goes </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2479055,
                  "author_name": "kishanvavdara",
                  "author_url": "",
                  "post_date": "10/12/2023 11:30:12",
                  "content": "<p>All the best!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2479759,
      "author_name": "laurasisson",
      "author_url": "",
      "post_date": "10/12/2023 20:51:28",
      "content": "<p>I am trying to build a two stage model, with one stage predicting the 978 landmark genes shared between the Kaggle data and the <a href=\"https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE70138\" target=\"_blank\">LINCS</a> across 107,404 samples, and then a second model predicting the full transcriptome using just the Kaggle data.<br>\nI haven't had much success, but maybe I haven't been using the right model.</p>\n<p>I made the basics of the <a href=\"https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation\" target=\"_blank\">dataset preparation</a> public and would be happy to talk more.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2480011,
          "author_name": "kishanvavdara",
          "author_url": "",
          "post_date": "10/13/2023 04:40:06",
          "content": "<p>Great progress! It's wonderful that you've shared your dataset preparation work. Building a two-stage model sounds like a promising approach.  Integrating predictions for the landmark genes from Kaggle and LINCS data across a significant number of samples is a complex task. I hope it ll go well.</p>\n<p>I tried pulling some chemical compound data from <a href=\"https://lincsportal.ccs.miami.edu/sigc-api/swagger-ui.html\" target=\"_blank\">lincs portal</a> and still processing it. </p>\n<p>Keep up the good work, and feel free to share any updates or challenges you face. Thank you for sharing your progress.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2486087,
      "author_name": "wguesdon",
      "author_url": "",
      "post_date": "10/17/2023 17:09:00",
      "content": "<p>I have generated embeddings for the SMILE columns using DeepChem/ChemBERTa-77M-MTR (<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550#2485877\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550#2485877</a>) and I am now in the process of training a neural network based on this (<a href=\"https://www.kaggle.com/code/kishanvavdara/neural-network-regression)\" target=\"_blank\">https://www.kaggle.com/code/kishanvavdara/neural-network-regression)</a>. Then I would like to try XGBoost as well. Still quite a lot of things to figure out, but moving forward.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2486711,
          "author_name": "kishanvavdara",
          "author_url": "",
          "post_date": "10/18/2023 05:56:46",
          "content": "<p>Great! I ll also try doing that. All the best <a href=\"https://www.kaggle.com/wguesdon\" target=\"_blank\">@wguesdon</a> Hope you get good results!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2487059,
              "author_name": "wguesdon",
              "author_url": "",
              "post_date": "10/18/2023 10:44:25",
              "content": "<p>Thanks :) Does this improve model performance for you?<br>\nIt seems to decrease performance so far compared to not using embedding.  I am now trying to apply PCA or other feature engineering methods to those embeddings. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2477709": "Hello,  friends!\n\nHas anyone explored incorporating cell line data or other auxiliary data sources, alongside complex ensemble models and fine-tuning techniques, to improve model performance on the public test dataset? If you have ventured into this, I'd greatly appreciate learning from your insights and experiences.\n\nThank you for sharing your knowledge.\n\nBest regards.",
    "2478167": "Not yet, I am trying various techniques to extract some features from the SMILES, but there are still no results.🤦‍♂️",
    "2478491": "Awesome ! Have you considered using SMILES embeddings or leveraging pre-trained models for this?  Give it a try and let me know how it goes! Thanks for sharing your progress. 🙌",
    "2478655": "Yes, I tried with two strategies. The first one was Morgan fingerprints at 2048 bits. Initially, it seemed promising, but I had significant overfitting issues. Now, I'm experimenting with embeddings using Mol2Vec and have managed to reduce the RMSE a bit. Over the weekend, I'll play around with keras tuner and see how it goes",
    "2479055": "All the best!",
    "2479759": "I am trying to build a two stage model, with one stage predicting the 978 landmark genes shared between the Kaggle data and the [LINCS](https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE70138) across 107,404 samples, and then a second model predicting the full transcriptome using just the Kaggle data.\nI haven't had much success, but maybe I haven't been using the right model.\n\nI made the basics of the [dataset preparation](https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation) public and would be happy to talk more.",
    "2480011": "Great progress! It's wonderful that you've shared your dataset preparation work. Building a two-stage model sounds like a promising approach.  Integrating predictions for the landmark genes from Kaggle and LINCS data across a significant number of samples is a complex task. I hope it ll go well.\n\nI tried pulling some chemical compound data from [lincs portal](https://lincsportal.ccs.miami.edu/sigc-api/swagger-ui.html) and still processing it. \n\nKeep up the good work, and feel free to share any updates or challenges you face. Thank you for sharing your progress.",
    "2486087": "I have generated embeddings for the SMILE columns using DeepChem/ChemBERTa-77M-MTR (https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/441550#2485877) and I am now in the process of training a neural network based on this (https://www.kaggle.com/code/kishanvavdara/neural-network-regression). Then I would like to try XGBoost as well. Still quite a lot of things to figure out, but moving forward.",
    "2486711": "Great! I ll also try doing that. All the best @wguesdon Hope you get good results!",
    "2487059": "Thanks :) Does this improve model performance for you?\nIt seems to decrease performance so far compared to not using embedding.  I am now trying to apply PCA or other feature engineering methods to those embeddings."
  },
  "source": "meta"
}