{
  "id": 458916,
  "title": "Lessons Learned",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/458916",
  "author_name": "",
  "post_date": "2023-12-02T09:10:52.083281200Z",
  "votes": 20,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my second competition, so I feel it's very helpful to make some conclusions for the future. Would be happy if you could share yours)</p>\n<p><strong>1) Start with exploratory data analysis, as deep as possible, it's worth the time! And come back to it from time to time.</strong></p>\n<p>I did a few:<br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-adata-analysis-with-seurat\" target=\"_blank\">Single-cell data analysis</a><br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda\" target=\"_blank\">Data overview</a><br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-drugbank-data-for-compounds-eda\" target=\"_blank\">DrugBank compound data</a></p>\n<p>But all I have learned from them is this:</p>\n<ul>\n<li>I cannot figure out the exact workflow of data preparation by the organizers (e.g., <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/453376****\" target=\"_blank\">did they account for the negative control effect? if so, at what exact stage of data processing?</a> Why did <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2487564\" target=\"_blank\">they select compounds with data for a certain number of cells from different donors for the test</a>, but not for the training data set?) </li>\n<li>There is no easy way to extract additional information for all 146 drugs (there are drugs I couldn't find in any database I know of). The fact that the drug names contain errors makes this even more difficult.</li>\n<li>There are outliers in the data and something is wrong with the adata_train, like missing counts after the negative control treatment. </li>\n</ul>\n<p>If I try even harder, I might succeed… happy to learn at least after the competition is over, that most answers were there, in the data - see <a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense\" target=\"_blank\">the smart EDA which makes sense</a> by Mr. <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p><strong>2) Find a suitable local validation and save all metrics for comparison.</strong></p>\n<p>After a few tries with random folds, I started to use validation on test drugs. This resulted in the training set of each fold containing all available data from the test cells (myeloid and B), while the test part contained only the drugs for which we needed to predict. This improved my LB results and I just stuck with it and didn't try any other schemes. With this validation we could not check simple tricks (noise, features, model parameters) locally, only big improvements were seen both locally and on the LB.</p>\n<p><strong>3) Don't stick to one model and think hard about post-processing strategies.</strong></p>\n<p>Correct me if I am wrong, but it seems that most winning solutions use multiple approaches rather than relying on a single model.  Also, post-processing can be very important to adjust for model biases.</p>\n<p>I stuck with just one model - <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-20th-place-solution\" target=\"_blank\">MLP</a> - and did hard work to improve it and make it more sophisticated. Well, based on the private scores I see that the last month I was doing nothing but overfilling (although locally there was some improvement).</p>\n<p>Luckily, I found a team and the final solution implies different boostings and NN.<br>\nAs for the post-processing, I left thinking about it until the end, and at the end I realized that my EDAs were unfinished, many questions did not allow me to find out what post-processing to use.</p>\n<p><strong>4) Build a team at the beginning, separate tasks and idea testing, agree on a validation that fits all models, and save the same Yoof to use for local blend validation.</strong></p>\n<p>I guess there is no need to explain… I'll just tell you that we didn't do any of this… yet our solutions improved each other significantly!! This means that we could have done a lot better if we had more time together 😉</p>\n<p>Did you follow any of this? Was it helpful? </p>",
  "messages": [
    {
      "id": "2546268",
      "postDate": "12/02/2023 09:10:52",
      "content": "<p>This is my second competition, so I feel it's very helpful to make some conclusions for the future. Would be happy if you could share yours)</p>\n<p><strong>1) Start with exploratory data analysis, as deep as possible, it's worth the time! And come back to it from time to time.</strong></p>\n<p>I did a few:<br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-adata-analysis-with-seurat\" target=\"_blank\">Single-cell data analysis</a><br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda\" target=\"_blank\">Data overview</a><br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-drugbank-data-for-compounds-eda\" target=\"_blank\">DrugBank compound data</a></p>\n<p>But all I have learned from them is this:</p>\n<ul>\n<li>I cannot figure out the exact workflow of data preparation by the organizers (e.g., <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/453376****\" target=\"_blank\">did they account for the negative control effect? if so, at what exact stage of data processing?</a> Why did <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2487564\" target=\"_blank\">they select compounds with data for a certain number of cells from different donors for the test</a>, but not for the training data set?) </li>\n<li>There is no easy way to extract additional information for all 146 drugs (there are drugs I couldn't find in any database I know of). The fact that the drug names contain errors makes this even more difficult.</li>\n<li>There are outliers in the data and something is wrong with the adata_train, like missing counts after the negative control treatment. </li>\n</ul>\n<p>If I try even harder, I might succeed… happy to learn at least after the competition is over, that most answers were there, in the data - see <a href=\"https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense\" target=\"_blank\">the smart EDA which makes sense</a> by Mr. <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>\n<p><strong>2) Find a suitable local validation and save all metrics for comparison.</strong></p>\n<p>After a few tries with random folds, I started to use validation on test drugs. This resulted in the training set of each fold containing all available data from the test cells (myeloid and B), while the test part contained only the drugs for which we needed to predict. This improved my LB results and I just stuck with it and didn't try any other schemes. With this validation we could not check simple tricks (noise, features, model parameters) locally, only big improvements were seen both locally and on the LB.</p>\n<p><strong>3) Don't stick to one model and think hard about post-processing strategies.</strong></p>\n<p>Correct me if I am wrong, but it seems that most winning solutions use multiple approaches rather than relying on a single model.  Also, post-processing can be very important to adjust for model biases.</p>\n<p>I stuck with just one model - <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-20th-place-solution\" target=\"_blank\">MLP</a> - and did hard work to improve it and make it more sophisticated. Well, based on the private scores I see that the last month I was doing nothing but overfilling (although locally there was some improvement).</p>\n<p>Luckily, I found a team and the final solution implies different boostings and NN.<br>\nAs for the post-processing, I left thinking about it until the end, and at the end I realized that my EDAs were unfinished, many questions did not allow me to find out what post-processing to use.</p>\n<p><strong>4) Build a team at the beginning, separate tasks and idea testing, agree on a validation that fits all models, and save the same Yoof to use for local blend validation.</strong></p>\n<p>I guess there is no need to explain… I'll just tell you that we didn't do any of this… yet our solutions improved each other significantly!! This means that we could have done a lot better if we had more time together 😉</p>\n<p>Did you follow any of this? Was it helpful? </p>",
      "rawMarkdown": "This is my second competition, so I feel it's very helpful to make some conclusions for the future. Would be happy if you could share yours)\n\n**1) Start with exploratory data analysis, as deep as possible, it's worth the time! And come back to it from time to time.**\n\nI did a few:\n[Single-cell data analysis](https://www.kaggle.com/code/antoninadolgorukova/op2-adata-analysis-with-seurat)\n[Data overview](https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda)\n[DrugBank compound data](https://www.kaggle.com/code/antoninadolgorukova/op2-drugbank-data-for-compounds-eda)\n\nBut all I have learned from them is this:\n- I cannot figure out the exact workflow of data preparation by the organizers (e.g., [did they account for the negative control effect? if so, at what exact stage of data processing?](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/453376****) Why did [they select compounds with data for a certain number of cells from different donors for the test](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2487564), but not for the training data set?) \n- There is no easy way to extract additional information for all 146 drugs (there are drugs I couldn't find in any database I know of). The fact that the drug names contain errors makes this even more difficult.\n- There are outliers in the data and something is wrong with the adata_train, like missing counts after the negative control treatment. \n\nIf I try even harder, I might succeed... happy to learn at least after the competition is over, that most answers were there, in the data - see [the smart EDA which makes sense](https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense) by Mr. @ambrosm \n\n**2) Find a suitable local validation and save all metrics for comparison.**\n\nAfter a few tries with random folds, I started to use validation on test drugs. This resulted in the training set of each fold containing all available data from the test cells (myeloid and B), while the test part contained only the drugs for which we needed to predict. This improved my LB results and I just stuck with it and didn't try any other schemes. With this validation we could not check simple tricks (noise, features, model parameters) locally, only big improvements were seen both locally and on the LB.\n\n**3) Don't stick to one model and think hard about post-processing strategies.**\n\nCorrect me if I am wrong, but it seems that most winning solutions use multiple approaches rather than relying on a single model.  Also, post-processing can be very important to adjust for model biases.\n\nI stuck with just one model - [MLP](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-20th-place-solution) - and did hard work to improve it and make it more sophisticated. Well, based on the private scores I see that the last month I was doing nothing but overfilling (although locally there was some improvement).\n\nLuckily, I found a team and the final solution implies different boostings and NN.\nAs for the post-processing, I left thinking about it until the end, and at the end I realized that my EDAs were unfinished, many questions did not allow me to find out what post-processing to use.\n\n**4) Build a team at the beginning, separate tasks and idea testing, agree on a validation that fits all models, and save the same Yoof to use for local blend validation.**\n\nI guess there is no need to explain... I'll just tell you that we didn't do any of this... yet our solutions improved each other significantly!! This means that we could have done a lot better if we had more time together 😉\n\nDid you follow any of this? Was it helpful?",
      "votes": null
    },
    {
      "id": "2546465",
      "postDate": "12/02/2023 13:16:56",
      "content": "<p>They were looking for PubChem, weren't they? They have a python API and can be searched by SMILES.</p>",
      "rawMarkdown": "They were looking for PubChem, weren't they? They have a python API and can be searched by SMILES.",
      "votes": null
    },
    {
      "id": "2546508",
      "postDate": "12/02/2023 13:48:35",
      "content": "<p>It's just my personal experience, didn't mean it is not possible)</p>\n<p>I did try pubchem and other databases, using drug names and SMILES. As for SMILES, it appears there are many different variations (canonical, non-canonical, isomeric, with aromatic atoms represented differenty, etc.), and for someone who is not a specialist, it can be quite challenging to determine which SMILES are provided. In my non-specialist opinion, the SMILES in 'de_train' are canonical, whereas in DrugBank, for example, only isomeric SMILES are given. I don't remember want kind of in PubChem, but I tried indeed, spent a lot of time on this and couldn't extract all the drugs, only part of them.</p>",
      "rawMarkdown": "It's just my personal experience, didn't mean it is not possible)\n\nI did try pubchem and other databases, using drug names and SMILES. As for SMILES, it appears there are many different variations (canonical, non-canonical, isomeric, with aromatic atoms represented differenty, etc.), and for someone who is not a specialist, it can be quite challenging to determine which SMILES are provided. In my non-specialist opinion, the SMILES in 'de_train' are canonical, whereas in DrugBank, for example, only isomeric SMILES are given. I don't remember want kind of in PubChem, but I tried indeed, spent a lot of time on this and couldn't extract all the drugs, only part of them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2546465,
      "author_name": "alekseytrepetsky",
      "author_url": "",
      "post_date": "12/02/2023 13:16:56",
      "content": "<p>They were looking for PubChem, weren't they? They have a python API and can be searched by SMILES.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2546508,
          "author_name": "antoninadolgorukova",
          "author_url": "",
          "post_date": "12/02/2023 13:48:35",
          "content": "<p>It's just my personal experience, didn't mean it is not possible)</p>\n<p>I did try pubchem and other databases, using drug names and SMILES. As for SMILES, it appears there are many different variations (canonical, non-canonical, isomeric, with aromatic atoms represented differenty, etc.), and for someone who is not a specialist, it can be quite challenging to determine which SMILES are provided. In my non-specialist opinion, the SMILES in 'de_train' are canonical, whereas in DrugBank, for example, only isomeric SMILES are given. I don't remember want kind of in PubChem, but I tried indeed, spent a lot of time on this and couldn't extract all the drugs, only part of them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2546268": "This is my second competition, so I feel it's very helpful to make some conclusions for the future. Would be happy if you could share yours)\n\n**1) Start with exploratory data analysis, as deep as possible, it's worth the time! And come back to it from time to time.**\n\nI did a few:\n[Single-cell data analysis](https://www.kaggle.com/code/antoninadolgorukova/op2-adata-analysis-with-seurat)\n[Data overview](https://www.kaggle.com/code/antoninadolgorukova/op2-stats-sign-de-fingerprints-eda)\n[DrugBank compound data](https://www.kaggle.com/code/antoninadolgorukova/op2-drugbank-data-for-compounds-eda)\n\nBut all I have learned from them is this:\n- I cannot figure out the exact workflow of data preparation by the organizers (e.g., [did they account for the negative control effect? if so, at what exact stage of data processing?](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/453376****) Why did [they select compounds with data for a certain number of cells from different donors for the test](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/445883#2487564), but not for the training data set?) \n- There is no easy way to extract additional information for all 146 drugs (there are drugs I couldn't find in any database I know of). The fact that the drug names contain errors makes this even more difficult.\n- There are outliers in the data and something is wrong with the adata_train, like missing counts after the negative control treatment. \n\nIf I try even harder, I might succeed... happy to learn at least after the competition is over, that most answers were there, in the data - see [the smart EDA which makes sense](https://www.kaggle.com/code/ambrosm/scp-eda-which-makes-sense) by Mr. @ambrosm \n\n**2) Find a suitable local validation and save all metrics for comparison.**\n\nAfter a few tries with random folds, I started to use validation on test drugs. This resulted in the training set of each fold containing all available data from the test cells (myeloid and B), while the test part contained only the drugs for which we needed to predict. This improved my LB results and I just stuck with it and didn't try any other schemes. With this validation we could not check simple tricks (noise, features, model parameters) locally, only big improvements were seen both locally and on the LB.\n\n**3) Don't stick to one model and think hard about post-processing strategies.**\n\nCorrect me if I am wrong, but it seems that most winning solutions use multiple approaches rather than relying on a single model.  Also, post-processing can be very important to adjust for model biases.\n\nI stuck with just one model - [MLP](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-20th-place-solution) - and did hard work to improve it and make it more sophisticated. Well, based on the private scores I see that the last month I was doing nothing but overfilling (although locally there was some improvement).\n\nLuckily, I found a team and the final solution implies different boostings and NN.\nAs for the post-processing, I left thinking about it until the end, and at the end I realized that my EDAs were unfinished, many questions did not allow me to find out what post-processing to use.\n\n**4) Build a team at the beginning, separate tasks and idea testing, agree on a validation that fits all models, and save the same Yoof to use for local blend validation.**\n\nI guess there is no need to explain... I'll just tell you that we didn't do any of this... yet our solutions improved each other significantly!! This means that we could have done a lot better if we had more time together 😉\n\nDid you follow any of this? Was it helpful?",
    "2546465": "They were looking for PubChem, weren't they? They have a python API and can be searched by SMILES.",
    "2546508": "It's just my personal experience, didn't mean it is not possible)\n\nI did try pubchem and other databases, using drug names and SMILES. As for SMILES, it appears there are many different variations (canonical, non-canonical, isomeric, with aromatic atoms represented differenty, etc.), and for someone who is not a specialist, it can be quite challenging to determine which SMILES are provided. In my non-specialist opinion, the SMILES in 'de_train' are canonical, whereas in DrugBank, for example, only isomeric SMILES are given. I don't remember want kind of in PubChem, but I tried indeed, spent a lot of time on this and couldn't extract all the drugs, only part of them."
  },
  "source": "meta"
}