{
  "id": 457782,
  "title": "Mistakes you made ?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/457782",
  "author_name": "",
  "post_date": "2023-11-26T20:38:22.664987Z",
  "votes": 16,
  "comment_count": 6,
  "views": 0,
  "content": "<p>If start from scratch what would you do in the other way or not do at all ?</p>\n<p>From my side: </p>\n<p>1) low number of samples made me think that we need simple models , while in reality NN, pyboost seems work better. Probably the issue is not in  sample number but samples * targets - which is quite big.</p>\n<p>2)  Cross-validation vs LB: most of the time it seemed to me that it is hopeless to find good correspondence again mainly due to small sample size, however near to the end we found that situation is not that much pessimistic - and if start from scratch I would spend more time exploring various schemes and more thinking…</p>\n<p>3) Analysis of  outliers - that seems to be crucial - however still not clear for me how to treat that problem efficiently…  </p>\n<p>And many more, may be add later</p>\n<p>PS<br>\nFor those have not seen before here is \"Kaggle wisdom from Grandmaster senkin13\" quite useful Kaggle advices:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351021\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351021</a></p>",
  "messages": [
    {
      "id": "2539212",
      "postDate": "11/26/2023 20:38:22",
      "content": "<p>If start from scratch what would you do in the other way or not do at all ?</p>\n<p>From my side: </p>\n<p>1) low number of samples made me think that we need simple models , while in reality NN, pyboost seems work better. Probably the issue is not in  sample number but samples * targets - which is quite big.</p>\n<p>2)  Cross-validation vs LB: most of the time it seemed to me that it is hopeless to find good correspondence again mainly due to small sample size, however near to the end we found that situation is not that much pessimistic - and if start from scratch I would spend more time exploring various schemes and more thinking…</p>\n<p>3) Analysis of  outliers - that seems to be crucial - however still not clear for me how to treat that problem efficiently…  </p>\n<p>And many more, may be add later</p>\n<p>PS<br>\nFor those have not seen before here is \"Kaggle wisdom from Grandmaster senkin13\" quite useful Kaggle advices:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351021\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351021</a></p>",
      "rawMarkdown": "If start from scratch what would you do in the other way or not do at all ?\n\nFrom my side: \n\n1) low number of samples made me think that we need simple models , while in reality NN, pyboost seems work better. Probably the issue is not in  sample number but samples * targets - which is quite big.\n\n2)  Cross-validation vs LB: most of the time it seemed to me that it is hopeless to find good correspondence again mainly due to small sample size, however near to the end we found that situation is not that much pessimistic - and if start from scratch I would spend more time exploring various schemes and more thinking...\n\n3) Analysis of  outliers - that seems to be crucial - however still not clear for me how to treat that problem efficiently...  \n\nAnd many more, may be add later\n\nPS\nFor those have not seen before here is \"Kaggle wisdom from Grandmaster senkin13\" quite useful Kaggle advices:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/351021",
      "votes": null
    },
    {
      "id": "2539401",
      "postDate": "11/27/2023 03:36:20",
      "content": "<p>Thank you, Alexander, for the invitation to reflect on and learn from mistakes made!</p>\n<p>Definitely +1 to your 2) and I wish I had realized earlier that XGBoost has GPU support – I spent quite some time slowly experimenting on 32 CPUs. For large datasets I plan to use XGBoost even for plain random forests in the future to leverage its GPU acceleration (see documentation <a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/rf.html\" target=\"_blank\">here</a>).</p>",
      "rawMarkdown": "Thank you, Alexander, for the invitation to reflect on and learn from mistakes made!\n\nDefinitely +1 to your 2) and I wish I had realized earlier that XGBoost has GPU support – I spent quite some time slowly experimenting on 32 CPUs. For large datasets I plan to use XGBoost even for plain random forests in the future to leverage its GPU acceleration (see documentation [here](https://xgboost.readthedocs.io/en/stable/tutorials/rf.html)).",
      "votes": null
    },
    {
      "id": "2539608",
      "postDate": "11/27/2023 07:40:57",
      "content": "<p>From my side:</p>\n<ol>\n<li><p>Not having a clean pre/post processing pipeline. Being lazy about it makes developement quick at the begining, but a diaster towards the later stage.</p></li>\n<li><p>Multi-tasking multiple kaggle competitions can be fun and spark ideas, but also very very very very challenging. </p></li>\n</ol>",
      "rawMarkdown": "From my side:\n\n1. Not having a clean pre/post processing pipeline. Being lazy about it makes developement quick at the begining, but a diaster towards the later stage.\n\n2. Multi-tasking multiple kaggle competitions can be fun and spark ideas, but also very very very very challenging.",
      "votes": null
    },
    {
      "id": "2540001",
      "postDate": "11/27/2023 12:52:45",
      "content": "<p>No idea until after the competition. Perhaps either \"took too much notice of public blends\" or \"took too little notice of public blends\". I know I'm making assumptions as to whether single models are competitive or whether ensembles are essential.</p>",
      "rawMarkdown": "No idea until after the competition. Perhaps either \"took too much notice of public blends\" or \"took too little notice of public blends\". I know I'm making assumptions as to whether single models are competitive or whether ensembles are essential.",
      "votes": null
    },
    {
      "id": "2544037",
      "postDate": "11/30/2023 14:58:05",
      "content": "<p>What? I don't know XGBoost has GPU support, thanks Frenio🙏</p>",
      "rawMarkdown": "What? I don't know XGBoost has GPU support, thanks Frenio🙏",
      "votes": null
    },
    {
      "id": "2544505",
      "postDate": "11/30/2023 21:35:06",
      "content": "<p>Just 614 rows in train data, which is chemical reactions in multiple kind of cells. Almost everyone may be mistaken.</p>",
      "rawMarkdown": "Just 614 rows in train data, which is chemical reactions in multiple kind of cells. Almost everyone may be mistaken.",
      "votes": null
    },
    {
      "id": "2544635",
      "postDate": "12/01/2023 02:01:12",
      "content": "<p>This comment relates to the first point of Alexander's reflection. I've been thinking about it differently; pretty early on I started converting the training data into long format which resulted in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (cell_type, sm_name, gene), and one target (signed -log(p-value) which I just called value). Therefore, all my models had to predict only one target for each cell type/chemical/gene-combination. Inference yielded a list of 4,643,805 prediction values which I just reshaped back into the submission format of 255x18,211.</p>",
      "rawMarkdown": "This comment relates to the first point of Alexander's reflection. I've been thinking about it differently; pretty early on I started converting the training data into long format which resulted in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (cell_type, sm_name, gene), and one target (signed -log(p-value) which I just called value). Therefore, all my models had to predict only one target for each cell type/chemical/gene-combination. Inference yielded a list of 4,643,805 prediction values which I just reshaped back into the submission format of 255x18,211.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2539401,
      "author_name": "frenio",
      "author_url": "",
      "post_date": "11/27/2023 03:36:20",
      "content": "<p>Thank you, Alexander, for the invitation to reflect on and learn from mistakes made!</p>\n<p>Definitely +1 to your 2) and I wish I had realized earlier that XGBoost has GPU support – I spent quite some time slowly experimenting on 32 CPUs. For large datasets I plan to use XGBoost even for plain random forests in the future to leverage its GPU acceleration (see documentation <a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/rf.html\" target=\"_blank\">here</a>).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2544037,
          "author_name": "ermalossi",
          "author_url": "",
          "post_date": "11/30/2023 14:58:05",
          "content": "<p>What? I don't know XGBoost has GPU support, thanks Frenio🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2539608,
      "author_name": "qihuaz",
      "author_url": "",
      "post_date": "11/27/2023 07:40:57",
      "content": "<p>From my side:</p>\n<ol>\n<li><p>Not having a clean pre/post processing pipeline. Being lazy about it makes developement quick at the begining, but a diaster towards the later stage.</p></li>\n<li><p>Multi-tasking multiple kaggle competitions can be fun and spark ideas, but also very very very very challenging. </p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2540001,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "11/27/2023 12:52:45",
      "content": "<p>No idea until after the competition. Perhaps either \"took too much notice of public blends\" or \"took too little notice of public blends\". I know I'm making assumptions as to whether single models are competitive or whether ensembles are essential.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2544505,
      "author_name": "hdynamics",
      "author_url": "",
      "post_date": "11/30/2023 21:35:06",
      "content": "<p>Just 614 rows in train data, which is chemical reactions in multiple kind of cells. Almost everyone may be mistaken.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2544635,
          "author_name": "frenio",
          "author_url": "",
          "post_date": "12/01/2023 02:01:12",
          "content": "<p>This comment relates to the first point of Alexander's reflection. I've been thinking about it differently; pretty early on I started converting the training data into long format which resulted in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (cell_type, sm_name, gene), and one target (signed -log(p-value) which I just called value). Therefore, all my models had to predict only one target for each cell type/chemical/gene-combination. Inference yielded a list of 4,643,805 prediction values which I just reshaped back into the submission format of 255x18,211.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2539212": "If start from scratch what would you do in the other way or not do at all ?\n\nFrom my side: \n\n1) low number of samples made me think that we need simple models , while in reality NN, pyboost seems work better. Probably the issue is not in  sample number but samples * targets - which is quite big.\n\n2)  Cross-validation vs LB: most of the time it seemed to me that it is hopeless to find good correspondence again mainly due to small sample size, however near to the end we found that situation is not that much pessimistic - and if start from scratch I would spend more time exploring various schemes and more thinking...\n\n3) Analysis of  outliers - that seems to be crucial - however still not clear for me how to treat that problem efficiently...  \n\nAnd many more, may be add later\n\nPS\nFor those have not seen before here is \"Kaggle wisdom from Grandmaster senkin13\" quite useful Kaggle advices:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/351021",
    "2539401": "Thank you, Alexander, for the invitation to reflect on and learn from mistakes made!\n\nDefinitely +1 to your 2) and I wish I had realized earlier that XGBoost has GPU support – I spent quite some time slowly experimenting on 32 CPUs. For large datasets I plan to use XGBoost even for plain random forests in the future to leverage its GPU acceleration (see documentation [here](https://xgboost.readthedocs.io/en/stable/tutorials/rf.html)).",
    "2539608": "From my side:\n\n1. Not having a clean pre/post processing pipeline. Being lazy about it makes developement quick at the begining, but a diaster towards the later stage.\n\n2. Multi-tasking multiple kaggle competitions can be fun and spark ideas, but also very very very very challenging.",
    "2540001": "No idea until after the competition. Perhaps either \"took too much notice of public blends\" or \"took too little notice of public blends\". I know I'm making assumptions as to whether single models are competitive or whether ensembles are essential.",
    "2544037": "What? I don't know XGBoost has GPU support, thanks Frenio🙏",
    "2544505": "Just 614 rows in train data, which is chemical reactions in multiple kind of cells. Almost everyone may be mistaken.",
    "2544635": "This comment relates to the first point of Alexander's reflection. I've been thinking about it differently; pretty early on I started converting the training data into long format which resulted in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (cell_type, sm_name, gene), and one target (signed -log(p-value) which I just called value). Therefore, all my models had to predict only one target for each cell type/chemical/gene-combination. Inference yielded a list of 4,643,805 prediction values which I just reshaped back into the submission format of 255x18,211."
  },
  "source": "meta"
}