{
  "id": 366455,
  "title": "12th-Place Solution",
  "url": "/competitions/open-problems-multimodal/writeups/silogram-12th-place-solution",
  "author_name": "",
  "post_date": "2022-11-16T08:26:55.390Z",
  "votes": 53,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First, congratulations to the winners, especially <a href=\"https://www.kaggle.com/senkin\" target=\"_blank\">@senkin</a> and <a href=\"https://www.kaggle.com/tmp\" target=\"_blank\">@tmp</a> for leading throughout and <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> for his big jump from public to private LB. I'm eager to hear from both teams about their techniques. Overall, it was an interesting competition that gave me a greater appreciation for the challenges that bioinformaticians face. </p>\n<p>I don't know if others feel the same, but it seemed like a very long competition to me. I ran out of gas about midway and didn't really work on it much the last 4 weeks or so, which means I didn't use the raw counts at all. My solution, therefore, is fairly simple. </p>\n<p><strong>CV setup:</strong> I assigned each batch (unique user/day) to a separate fold, so I had 9 folds for citeseq and 12 for multiome. This was expensive, but had the virtue that it wasn't optimized towards either new donors (public LB) or new days (private LB). LB scores tracked local CV scores very closely. Even small gains in local CV almost always led to similar gains on the LB. This CV scheme was probably the reason that I fared well on the private LB (46-&gt;12).</p>\n<p><strong>Data transformations:</strong> I tried a lot of ways to denoise and transform the data, but most of them failed. In the end, I just used PCA and tSVD of the original data. </p>\n<p><strong>Feature engineering:</strong> None for Multiome. For Citeseq, I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.</p>\n<p><strong>Modeling:</strong> For Citeseq, I trained both single-target (140 separate models) and multi-target NN models (using Fastai). I also trained LGB and CatBoost models for each target. Altogether, I trained over 20 sets of models using different variations of the PCA data combined with selected features from the feature engineering. Individually, the models had local CV scores in the range 0.8995 - 0.9017. For multiome, I trained 3 multi-target NNs on the tSVD-reduced targets, and one CatBoost model.</p>\n<p><strong>Ensembling:</strong> I blended the models together using a very simple optimized weighting scheme where the only possible weights were 0,1,2, or 3. I tried other ensembling techniques that had higher CV scores, but they performed worse on the LB. I was afraid this might be due to some hidden leakage between folds, so I stuck to the simpler weighting scheme. This led to local CV scores of 0.9039 for Citseq and 0.669 for Multiome.</p>\n<p><strong>Thoughts about trends in data:</strong> I was intrigued by <a href=\"https://www.kaggle.com/AmbroseM\" target=\"_blank\">@AmbroseM</a>'s posts arguing that the data is a time series. There are undoubtedly trends over the 7 days of training data, but I was concerned about whether these trends would continue to day 10. I don't know enough about the biology, but it seems likely that cell behavior has both long-term trends (aging) and short-term trends based on things like diet, exercise, illness, etc. I decided that the trends visible in the 7 days of training data could very easily be short-term trends that would reverse themselves after 7 days, or could be just coincidental due to technical aspects of the data collection. Based on <a href=\"https://www.kaggle.com/AmbroseM\" target=\"_blank\">@AmbroseM</a>'s post here (<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395)\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395)</a>, it seems like the trends did, in fact, continue to the test set. I would be very interested to hear from some cell scientists about what these trends might signify, whether they're cyclical in nature, and if yes, how long each cycle is typically.</p>",
  "messages": [
    {
      "id": "2031734",
      "postDate": "11/16/2022 08:24:11",
      "content": "<p>First, congratulations to the winners, especially <a href=\"https://www.kaggle.com/senkin\" target=\"_blank\">@senkin</a> and <a href=\"https://www.kaggle.com/tmp\" target=\"_blank\">@tmp</a> for leading throughout and <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> for his big jump from public to private LB. I'm eager to hear from both teams about their techniques. Overall, it was an interesting competition that gave me a greater appreciation for the challenges that bioinformaticians face. </p>\n<p>I don't know if others feel the same, but it seemed like a very long competition to me. I ran out of gas about midway and didn't really work on it much the last 4 weeks or so, which means I didn't use the raw counts at all. My solution, therefore, is fairly simple. </p>\n<p><strong>CV setup:</strong> I assigned each batch (unique user/day) to a separate fold, so I had 9 folds for citeseq and 12 for multiome. This was expensive, but had the virtue that it wasn't optimized towards either new donors (public LB) or new days (private LB). LB scores tracked local CV scores very closely. Even small gains in local CV almost always led to similar gains on the LB. This CV scheme was probably the reason that I fared well on the private LB (46-&gt;12).</p>\n<p><strong>Data transformations:</strong> I tried a lot of ways to denoise and transform the data, but most of them failed. In the end, I just used PCA and tSVD of the original data. </p>\n<p><strong>Feature engineering:</strong> None for Multiome. For Citeseq, I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.</p>\n<p><strong>Modeling:</strong> For Citeseq, I trained both single-target (140 separate models) and multi-target NN models (using Fastai). I also trained LGB and CatBoost models for each target. Altogether, I trained over 20 sets of models using different variations of the PCA data combined with selected features from the feature engineering. Individually, the models had local CV scores in the range 0.8995 - 0.9017. For multiome, I trained 3 multi-target NNs on the tSVD-reduced targets, and one CatBoost model.</p>\n<p><strong>Ensembling:</strong> I blended the models together using a very simple optimized weighting scheme where the only possible weights were 0,1,2, or 3. I tried other ensembling techniques that had higher CV scores, but they performed worse on the LB. I was afraid this might be due to some hidden leakage between folds, so I stuck to the simpler weighting scheme. This led to local CV scores of 0.9039 for Citseq and 0.669 for Multiome.</p>\n<p><strong>Thoughts about trends in data:</strong> I was intrigued by <a href=\"https://www.kaggle.com/AmbroseM\" target=\"_blank\">@AmbroseM</a>'s posts arguing that the data is a time series. There are undoubtedly trends over the 7 days of training data, but I was concerned about whether these trends would continue to day 10. I don't know enough about the biology, but it seems likely that cell behavior has both long-term trends (aging) and short-term trends based on things like diet, exercise, illness, etc. I decided that the trends visible in the 7 days of training data could very easily be short-term trends that would reverse themselves after 7 days, or could be just coincidental due to technical aspects of the data collection. Based on <a href=\"https://www.kaggle.com/AmbroseM\" target=\"_blank\">@AmbroseM</a>'s post here (<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395)\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395)</a>, it seems like the trends did, in fact, continue to the test set. I would be very interested to hear from some cell scientists about what these trends might signify, whether they're cyclical in nature, and if yes, how long each cycle is typically.</p>",
      "rawMarkdown": "First, congratulations to the winners, especially @senkin and @tmp for leading throughout and @shujisuzuki65 for his big jump from public to private LB. I'm eager to hear from both teams about their techniques. Overall, it was an interesting competition that gave me a greater appreciation for the challenges that bioinformaticians face. \n\nI don't know if others feel the same, but it seemed like a very long competition to me. I ran out of gas about midway and didn't really work on it much the last 4 weeks or so, which means I didn't use the raw counts at all. My solution, therefore, is fairly simple. \n\n**CV setup:** I assigned each batch (unique user/day) to a separate fold, so I had 9 folds for citeseq and 12 for multiome. This was expensive, but had the virtue that it wasn't optimized towards either new donors (public LB) or new days (private LB). LB scores tracked local CV scores very closely. Even small gains in local CV almost always led to similar gains on the LB. This CV scheme was probably the reason that I fared well on the private LB (46->12).\n\n**Data transformations:** I tried a lot of ways to denoise and transform the data, but most of them failed. In the end, I just used PCA and tSVD of the original data. \n\n**Feature engineering:** None for Multiome. For Citeseq, I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\n\n**Modeling:** For Citeseq, I trained both single-target (140 separate models) and multi-target NN models (using Fastai). I also trained LGB and CatBoost models for each target. Altogether, I trained over 20 sets of models using different variations of the PCA data combined with selected features from the feature engineering. Individually, the models had local CV scores in the range 0.8995 - 0.9017. For multiome, I trained 3 multi-target NNs on the tSVD-reduced targets, and one CatBoost model.\n\n**Ensembling:** I blended the models together using a very simple optimized weighting scheme where the only possible weights were 0,1,2, or 3. I tried other ensembling techniques that had higher CV scores, but they performed worse on the LB. I was afraid this might be due to some hidden leakage between folds, so I stuck to the simpler weighting scheme. This led to local CV scores of 0.9039 for Citseq and 0.669 for Multiome.\n\n**Thoughts about trends in data:** I was intrigued by @AmbroseM's posts arguing that the data is a time series. There are undoubtedly trends over the 7 days of training data, but I was concerned about whether these trends would continue to day 10. I don't know enough about the biology, but it seems likely that cell behavior has both long-term trends (aging) and short-term trends based on things like diet, exercise, illness, etc. I decided that the trends visible in the 7 days of training data could very easily be short-term trends that would reverse themselves after 7 days, or could be just coincidental due to technical aspects of the data collection. Based on @AmbroseM's post here (https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395), it seems like the trends did, in fact, continue to the test set. I would be very interested to hear from some cell scientists about what these trends might signify, whether they're cyclical in nature, and if yes, how long each cycle is typically.",
      "votes": null
    },
    {
      "id": "2031775",
      "postDate": "11/16/2022 08:37:05",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> </p>",
      "rawMarkdown": "Thanks for sharing @psilogram",
      "votes": null
    },
    {
      "id": "2031783",
      "postDate": "11/16/2022 08:40:19",
      "content": "<blockquote>\n  <p>The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.</p>\n</blockquote>\n<p>I think this is really cool!</p>",
      "rawMarkdown": "> The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\n\nI think this is really cool!",
      "votes": null
    },
    {
      "id": "2032872",
      "postDate": "11/16/2022 21:01:14",
      "content": "<p>Hearty congratulations <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>, thanks for the approach too. All the best!!</p>",
      "rawMarkdown": "Hearty congratulations @psilogram, thanks for the approach too. All the best!!",
      "votes": null
    },
    {
      "id": "2034071",
      "postDate": "11/17/2022 19:56:08",
      "content": "<p>Thanks for sharing ! And congratulations with the gold !</p>\n<p>May I ask a question about the validation scheme:<br>\nfrom the naive point of view it is not that much strong -  validation fold and train folds have the same day and the donor - so from the naive point of view it is not easy to expect that it would be similar to LB which contains completely new day and donor.</p>\n<p>What do you think ? </p>\n<p>Senkin13 writes that even random validation worked well for him, but they carefully checked the features.</p>\n<p>Validation scheme by KSS (and similar our scheme) <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366471\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366471</a><br>\nmight be stronger, but for after private LB was opened - I do not see full correspondence between private LB and our  scheme.</p>\n<p>So for you, Senkin some more simple validation scheme works, but for me even more sophisticated does not fully work,<br>\nso it seems the validation scheme is not the guarantee for the full correspondence between CV and LB.</p>\n<p>May be there are some additional trick to control features to include ?</p>\n<p>What are your opinion about that ? </p>",
      "rawMarkdown": "Thanks for sharing ! And congratulations with the gold !\n\nMay I ask a question about the validation scheme:\nfrom the naive point of view it is not that much strong -  validation fold and train folds have the same day and the donor - so from the naive point of view it is not easy to expect that it would be similar to LB which contains completely new day and donor.\n\nWhat do you think ? \n\nSenkin13 writes that even random validation worked well for him, but they carefully checked the features.\n\nValidation scheme by KSS (and similar our scheme) https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366471\nmight be stronger, but for after private LB was opened - I do not see full correspondence between private LB and our  scheme.\n\nSo for you, Senkin some more simple validation scheme works, but for me even more sophisticated does not fully work,\nso it seems the validation scheme is not the guarantee for the full correspondence between CV and LB.\n\nMay be there are some additional trick to control features to include ?\n\nWhat are your opinion about that ?",
      "votes": null
    },
    {
      "id": "2034524",
      "postDate": "11/18/2022 07:52:20",
      "content": "<p>I looked at it from a slightly different perspective. The public LB data has the same days but a different donor, so to optimize CV for the public score, we would want 3 folds, one for each donor in the training set. But this set-up could potentially penalize the private LB since it would not optimize for unseen days. </p>\n<p>Conversely, to optimize for the private LB, we would want separate folds by day (or even possibly past/future days if we believe it is a time series problem), but such a set-up might give us misleading feedback on the public LB if it turns out that the data is highly specific to donors. </p>\n<p>My system is a compromise between the donor/day approaches. The fact that it aligned so closely with the public LB gave me confidence that it was generalizing well for new donors, and the fact that each training fold contained data from all days meant that it wasn't biased too much by any particular day. </p>\n<p>The one thing my CV set-up doesn't do is optimize for future days, and this turned out to be its biggest failing.  Despite some clear trends in the data, I didn't know enough about the underlying biology to be persuaded that the trends would continue into the private test days. In retrospect, this is my biggest regret. Also, that I didn't bite the bullet and re-run my entire pipeline with the new raw data when it was released.</p>",
      "rawMarkdown": "I looked at it from a slightly different perspective. The public LB data has the same days but a different donor, so to optimize CV for the public score, we would want 3 folds, one for each donor in the training set. But this set-up could potentially penalize the private LB since it would not optimize for unseen days. \n\nConversely, to optimize for the private LB, we would want separate folds by day (or even possibly past/future days if we believe it is a time series problem), but such a set-up might give us misleading feedback on the public LB if it turns out that the data is highly specific to donors. \n\nMy system is a compromise between the donor/day approaches. The fact that it aligned so closely with the public LB gave me confidence that it was generalizing well for new donors, and the fact that each training fold contained data from all days meant that it wasn't biased too much by any particular day. \n\nThe one thing my CV set-up doesn't do is optimize for future days, and this turned out to be its biggest failing.  Despite some clear trends in the data, I didn't know enough about the underlying biology to be persuaded that the trends would continue into the private test days. In retrospect, this is my biggest regret. Also, that I didn't bite the bullet and re-run my entire pipeline with the new raw data when it was released.",
      "votes": null
    },
    {
      "id": "2035030",
      "postDate": "11/18/2022 14:47:56",
      "content": "<p>I agree with this!!! I thought about the same thing!! I guess that I was a bit skeptical that in fact the trend was going to continue, so I decided to make a general solution considering both public and private LB. </p>",
      "rawMarkdown": "I agree with this!!! I thought about the same thing!! I guess that I was a bit skeptical that in fact the trend was going to continue, so I decided to make a general solution considering both public and private LB.",
      "votes": null
    },
    {
      "id": "2035506",
      "postDate": "11/18/2022 23:55:54",
      "content": "<p>Congratulations Silogram. Great CV scheme and jump upward from public to private LB!</p>",
      "rawMarkdown": "Congratulations Silogram. Great CV scheme and jump upward from public to private LB!",
      "votes": null
    },
    {
      "id": "2041641",
      "postDate": "11/24/2022 05:19:04",
      "content": "<p>Congratulations and thanks for posting your great solution Silogram!</p>",
      "rawMarkdown": "Congratulations and thanks for posting your great solution Silogram!",
      "votes": null
    },
    {
      "id": "2044905",
      "postDate": "11/26/2022 19:52:18",
      "content": "<p>Thanks again for sharing and comments ! </p>\n<p>May I ask you about the following, would you be so kind to share the results you mention:<br>\n\"I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\"</p>\n<p>It would be great - if you can share these importances (and/or those 100-200 selected features) ! <br>\nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us. <br>\nPS<br>\nTechnically:  may be create some like \"Kaggle dataset\" and save csv files with importances  (and/or list of those 100-200 selected features)<br>\nSomething like that: <a href=\"https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\" target=\"_blank\">https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances</a></p>",
      "rawMarkdown": "Thanks again for sharing and comments ! \n\nMay I ask you about the following, would you be so kind to share the results you mention:\n\"I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\"\n\nIt would be great - if you can share these importances (and/or those 100-200 selected features) ! \nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us. \nPS\nTechnically:  may be create some like \"Kaggle dataset\" and save csv files with importances  (and/or list of those 100-200 selected features)\nSomething like that: https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances",
      "votes": null
    },
    {
      "id": "2057610",
      "postDate": "12/07/2022 08:52:30",
      "content": "<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> Sorry for the delay. I've uploaded the feature importance tables here: <a href=\"https://www.kaggle.com/datasets/psilogram/citeq-feature-importance\" target=\"_blank\">https://www.kaggle.com/datasets/psilogram/citeq-feature-importance</a></p>\n<p>there are a couple different versions from different runs. Let me know if you have any questions.</p>",
      "rawMarkdown": "alexandervc Sorry for the delay. I've uploaded the feature importance tables here: https://www.kaggle.com/datasets/psilogram/citeq-feature-importance\n\nthere are a couple different versions from different runs. Let me know if you have any questions.",
      "votes": null
    },
    {
      "id": "2059822",
      "postDate": "12/09/2022 08:59:30",
      "content": "<p>Thank you very much for sharing ! <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances</a><br>\nHere is some brief look on importances for CD36. </p>\n<p>We can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.</p>\n<p>Also look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.</p>\n<p>That is a first look - we need some time to get more insights from the data.</p>",
      "rawMarkdown": "Thank you very much for sharing ! \nhttps://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\nHere is some brief look on importances for CD36. \n\nWe can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.\n\nAlso look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.\n\nThat is a first look - we need some time to get more insights from the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031775,
      "author_name": "zvr842",
      "author_url": "",
      "post_date": "11/16/2022 08:37:05",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2031783,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "11/16/2022 08:40:19",
      "content": "<blockquote>\n  <p>The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.</p>\n</blockquote>\n<p>I think this is really cool!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032872,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/16/2022 21:01:14",
      "content": "<p>Hearty congratulations <a href=\"https://www.kaggle.com/psilogram\" target=\"_blank\">@psilogram</a>, thanks for the approach too. All the best!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2034071,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/17/2022 19:56:08",
      "content": "<p>Thanks for sharing ! And congratulations with the gold !</p>\n<p>May I ask a question about the validation scheme:<br>\nfrom the naive point of view it is not that much strong -  validation fold and train folds have the same day and the donor - so from the naive point of view it is not easy to expect that it would be similar to LB which contains completely new day and donor.</p>\n<p>What do you think ? </p>\n<p>Senkin13 writes that even random validation worked well for him, but they carefully checked the features.</p>\n<p>Validation scheme by KSS (and similar our scheme) <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366471\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366471</a><br>\nmight be stronger, but for after private LB was opened - I do not see full correspondence between private LB and our  scheme.</p>\n<p>So for you, Senkin some more simple validation scheme works, but for me even more sophisticated does not fully work,<br>\nso it seems the validation scheme is not the guarantee for the full correspondence between CV and LB.</p>\n<p>May be there are some additional trick to control features to include ?</p>\n<p>What are your opinion about that ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2034524,
          "author_name": "psilogram",
          "author_url": "",
          "post_date": "11/18/2022 07:52:20",
          "content": "<p>I looked at it from a slightly different perspective. The public LB data has the same days but a different donor, so to optimize CV for the public score, we would want 3 folds, one for each donor in the training set. But this set-up could potentially penalize the private LB since it would not optimize for unseen days. </p>\n<p>Conversely, to optimize for the private LB, we would want separate folds by day (or even possibly past/future days if we believe it is a time series problem), but such a set-up might give us misleading feedback on the public LB if it turns out that the data is highly specific to donors. </p>\n<p>My system is a compromise between the donor/day approaches. The fact that it aligned so closely with the public LB gave me confidence that it was generalizing well for new donors, and the fact that each training fold contained data from all days meant that it wasn't biased too much by any particular day. </p>\n<p>The one thing my CV set-up doesn't do is optimize for future days, and this turned out to be its biggest failing.  Despite some clear trends in the data, I didn't know enough about the underlying biology to be persuaded that the trends would continue into the private test days. In retrospect, this is my biggest regret. Also, that I didn't bite the bullet and re-run my entire pipeline with the new raw data when it was released.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2035030,
          "author_name": "sergiomiguelm",
          "author_url": "",
          "post_date": "11/18/2022 14:47:56",
          "content": "<p>I agree with this!!! I thought about the same thing!! I guess that I was a bit skeptical that in fact the trend was going to continue, so I decided to make a general solution considering both public and private LB. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2035506,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/18/2022 23:55:54",
      "content": "<p>Congratulations Silogram. Great CV scheme and jump upward from public to private LB!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2041641,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "11/24/2022 05:19:04",
      "content": "<p>Congratulations and thanks for posting your great solution Silogram!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2044905,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/26/2022 19:52:18",
      "content": "<p>Thanks again for sharing and comments ! </p>\n<p>May I ask you about the following, would you be so kind to share the results you mention:<br>\n\"I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\"</p>\n<p>It would be great - if you can share these importances (and/or those 100-200 selected features) ! <br>\nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us. <br>\nPS<br>\nTechnically:  may be create some like \"Kaggle dataset\" and save csv files with importances  (and/or list of those 100-200 selected features)<br>\nSomething like that: <a href=\"https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\" target=\"_blank\">https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2057610,
          "author_name": "psilogram",
          "author_url": "",
          "post_date": "12/07/2022 08:52:30",
          "content": "<p><a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> Sorry for the delay. I've uploaded the feature importance tables here: <a href=\"https://www.kaggle.com/datasets/psilogram/citeq-feature-importance\" target=\"_blank\">https://www.kaggle.com/datasets/psilogram/citeq-feature-importance</a></p>\n<p>there are a couple different versions from different runs. Let me know if you have any questions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2059822,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "12/09/2022 08:59:30",
              "content": "<p>Thank you very much for sharing ! <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances</a><br>\nHere is some brief look on importances for CD36. </p>\n<p>We can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.</p>\n<p>Also look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.</p>\n<p>That is a first look - we need some time to get more insights from the data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2031734": "First, congratulations to the winners, especially @senkin and @tmp for leading throughout and @shujisuzuki65 for his big jump from public to private LB. I'm eager to hear from both teams about their techniques. Overall, it was an interesting competition that gave me a greater appreciation for the challenges that bioinformaticians face. \n\nI don't know if others feel the same, but it seemed like a very long competition to me. I ran out of gas about midway and didn't really work on it much the last 4 weeks or so, which means I didn't use the raw counts at all. My solution, therefore, is fairly simple. \n\n**CV setup:** I assigned each batch (unique user/day) to a separate fold, so I had 9 folds for citeseq and 12 for multiome. This was expensive, but had the virtue that it wasn't optimized towards either new donors (public LB) or new days (private LB). LB scores tracked local CV scores very closely. Even small gains in local CV almost always led to similar gains on the LB. This CV scheme was probably the reason that I fared well on the private LB (46->12).\n\n**Data transformations:** I tried a lot of ways to denoise and transform the data, but most of them failed. In the end, I just used PCA and tSVD of the original data. \n\n**Feature engineering:** None for Multiome. For Citeseq, I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\n\n**Modeling:** For Citeseq, I trained both single-target (140 separate models) and multi-target NN models (using Fastai). I also trained LGB and CatBoost models for each target. Altogether, I trained over 20 sets of models using different variations of the PCA data combined with selected features from the feature engineering. Individually, the models had local CV scores in the range 0.8995 - 0.9017. For multiome, I trained 3 multi-target NNs on the tSVD-reduced targets, and one CatBoost model.\n\n**Ensembling:** I blended the models together using a very simple optimized weighting scheme where the only possible weights were 0,1,2, or 3. I tried other ensembling techniques that had higher CV scores, but they performed worse on the LB. I was afraid this might be due to some hidden leakage between folds, so I stuck to the simpler weighting scheme. This led to local CV scores of 0.9039 for Citseq and 0.669 for Multiome.\n\n**Thoughts about trends in data:** I was intrigued by @AmbroseM's posts arguing that the data is a time series. There are undoubtedly trends over the 7 days of training data, but I was concerned about whether these trends would continue to day 10. I don't know enough about the biology, but it seems likely that cell behavior has both long-term trends (aging) and short-term trends based on things like diet, exercise, illness, etc. I decided that the trends visible in the 7 days of training data could very easily be short-term trends that would reverse themselves after 7 days, or could be just coincidental due to technical aspects of the data collection. Based on @AmbroseM's post here (https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395), it seems like the trends did, in fact, continue to the test set. I would be very interested to hear from some cell scientists about what these trends might signify, whether they're cyclical in nature, and if yes, how long each cycle is typically.",
    "2031775": "Thanks for sharing @psilogram",
    "2031783": "> The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\n\nI think this is really cool!",
    "2032872": "Hearty congratulations @psilogram, thanks for the approach too. All the best!!",
    "2034071": "Thanks for sharing ! And congratulations with the gold !\n\nMay I ask a question about the validation scheme:\nfrom the naive point of view it is not that much strong -  validation fold and train folds have the same day and the donor - so from the naive point of view it is not easy to expect that it would be similar to LB which contains completely new day and donor.\n\nWhat do you think ? \n\nSenkin13 writes that even random validation worked well for him, but they carefully checked the features.\n\nValidation scheme by KSS (and similar our scheme) https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366471\nmight be stronger, but for after private LB was opened - I do not see full correspondence between private LB and our  scheme.\n\nSo for you, Senkin some more simple validation scheme works, but for me even more sophisticated does not fully work,\nso it seems the validation scheme is not the guarantee for the full correspondence between CV and LB.\n\nMay be there are some additional trick to control features to include ?\n\nWhat are your opinion about that ?",
    "2034524": "I looked at it from a slightly different perspective. The public LB data has the same days but a different donor, so to optimize CV for the public score, we would want 3 folds, one for each donor in the training set. But this set-up could potentially penalize the private LB since it would not optimize for unseen days. \n\nConversely, to optimize for the private LB, we would want separate folds by day (or even possibly past/future days if we believe it is a time series problem), but such a set-up might give us misleading feedback on the public LB if it turns out that the data is highly specific to donors. \n\nMy system is a compromise between the donor/day approaches. The fact that it aligned so closely with the public LB gave me confidence that it was generalizing well for new donors, and the fact that each training fold contained data from all days meant that it wasn't biased too much by any particular day. \n\nThe one thing my CV set-up doesn't do is optimize for future days, and this turned out to be its biggest failing.  Despite some clear trends in the data, I didn't know enough about the underlying biology to be persuaded that the trends would continue into the private test days. In retrospect, this is my biggest regret. Also, that I didn't bite the bullet and re-run my entire pipeline with the new raw data when it was released.",
    "2035030": "I agree with this!!! I thought about the same thing!! I guess that I was a bit skeptical that in fact the trend was going to continue, so I decided to make a general solution considering both public and private LB.",
    "2035506": "Congratulations Silogram. Great CV scheme and jump upward from public to private LB!",
    "2041641": "Congratulations and thanks for posting your great solution Silogram!",
    "2044905": "Thanks again for sharing and comments ! \n\nMay I ask you about the following, would you be so kind to share the results you mention:\n\"I trained 140 shallow LGB models (one for each target) using the full set of data. The goal here was not to use the models themselves, but to see which features were important for each target. I used the top 100-200 features per target in the later, deeper modeling.\"\n\nIt would be great - if you can share these importances (and/or those 100-200 selected features) ! \nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us. \nPS\nTechnically:  may be create some like \"Kaggle dataset\" and save csv files with importances  (and/or list of those 100-200 selected features)\nSomething like that: https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances",
    "2057610": "alexandervc Sorry for the delay. I've uploaded the feature importance tables here: https://www.kaggle.com/datasets/psilogram/citeq-feature-importance\n\nthere are a couple different versions from different runs. Let me know if you have any questions.",
    "2059822": "Thank you very much for sharing ! \nhttps://www.kaggle.com/code/alexandervc/silogram-cd36-feature-importances\nHere is some brief look on importances for CD36. \n\nWe can see quite some biology from them - e.g. - CD36 protein is higly activated in erythroid like cells - similar to red blood cells - so you can see genes like HBD, HBB, HBA1 as important features - that various forms of the hemoglobin - so quite as expected from biology.\n\nAlso look on the so-called genes enrichment analysis with KEGG pathways - again biologically reasonable results, and compared with importances by other methods - they are quite consistent.\n\nThat is a first look - we need some time to get more insights from the data."
  },
  "source": "meta"
}