{
  "id": 366471,
  "title": "7th Place Solution Summary",
  "url": "/competitions/open-problems-multimodal/writeups/chromosom-7th-place-solution-summary",
  "author_name": "",
  "post_date": "2022-11-16T10:24:14.600Z",
  "votes": 55,
  "comment_count": 12,
  "views": 0,
  "content": "<p>First of all, we would like to thank the organizers for an interesting challenge, as well as for the opportunity to use Saturn cloud.</p>\n<p><strong>CV Scheme</strong></p>\n<p>Since the organizer initially provided the information how public and private parts are splitted, we decided to utilize this info and design our local validation as similar as possible to the private part. So, we can say that the public part of the leaderboard was of little interest to us.</p>\n<p>The validation scheme is shown in the picture below (for citeseq part). Since we have three days in train dataset (citeseq task), the number of splits equals 6, for the multiome part the scheme was the same. However, there were 9 splits due to the larger number of days (at the very end we slightly modified this scheme so that validation fold always contained only one (nearest) day).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1243561%2F761304b2456999a0a31225be67858207%2Fcv_citeseq.png?generation=1668591942826347&amp;alt=media\" alt=\"\"></p>\n<p><strong>How to submit</strong></p>\n<p>Since the last available day (in train dataset) was always in validation folds, we could not use this validation scheme for the traditional folds blending. So, when we wanted to make submit, we used the usual KFold validation with shuffle.<br>\nThe approach was the following: all hypotheses and experiments, including hyperparameters tuning, were tested on true cv and afterwards these models were retrained on KFold.</p>\n<p><strong>Features and dimensionality reduction</strong></p>\n<p>We used PCA for dimensionality reduction for both tasks. Autoencoder was tested on citeseq dataset too, but performed slightly worse than PCA.<br>\nWe also used \"important\" features in their raw form for citeseq task.<br>\nMoreover we created some additional features based on aggregations of \"important\" columns over metadata (for example, \"mean_feature_1_by_donor\"). Such features gave us + 0.0003 for GBDT model on local CV.</p>\n<p><strong>Models (ensemble)</strong></p>\n<p><em>Citeseq task</em> : 3x multilayer perceptrons, 1x Conv1D, 1x pyBoost (best single model).<br>\n<em>Multiome task</em> : 1x multilayer perceptron, 1x TabNet, 1x pyBoost (best single model).</p>\n<p>pyBoost seems to be the new SOTA on multioutput tasks (at least among GBDT models).<br>\nIt's extremely fast to train as it uses GPU only and super easy to customize.<br>\n<a href=\"https://openreview.net/forum?id=WSxarC8t-T\" target=\"_blank\">Paper</a><br>\n<a href=\"https://github.com/sb-ai-lab/Py-Boost\" target=\"_blank\">Code</a></p>\n<p><strong>Some remarks about models</strong></p>\n<ul>\n<li><p><em>Multiome task</em>: all neural nets had 23418 output neurons. For pyBoost we reduced targets' dimension to 64 components using PCA.</p></li>\n<li><p>pyBoost was the best single model on True CV and KFold validation on citeseq data.</p></li>\n<li><p>For multiome task, it was the best model according to True CV and the<br>\nworst by KFold.</p></li>\n<li><p>We noticed that splitting targets  into groups and building a separate pyBoost model for each group improved our local CV a lot. By default, pyBoost can split targets into groups randomly, so we decided to try to improve it by splitting targets into groups based on their clusters, however, in the end it worked nearly the same as random splitting.</p></li>\n</ul>\n<p><strong>Data used to train</strong></p>\n<p><em>Citeseq task</em> : all available data<br>\n<em>Multiome task</em> : day 7 data only.</p>\n<p>Solving multiome task, we noticed that there is a significant performance drop on day 7 (on True CV).<br>\nThere was an idea that the reason is data drift in time, so we tried to train the model not on all available days, but only on the last available one. Locally, this improved our score by + 0.02. However, the problem was that there were no unseen days on the public leaderboard,<br>\nso training the model only on the last available day seriously dropped our public score (from 0.814 to 0.808). Nevertheless, we decided to follow the mantra \"trust your cv\" and as a result, this particular submission became our best on private.</p>\n<p>We also conducted a study on the similarity of days and found out that among available training days, day 3 is the most similar to private day 10 (day 7 is the second most similar).<br>\nNevertheless, since for other days the most similar day was always the closest in time, we decided to train our models on day 7.</p>",
  "messages": [
    {
      "id": "2031889",
      "postDate": "11/16/2022 09:36:53",
      "content": "<p>First of all, we would like to thank the organizers for an interesting challenge, as well as for the opportunity to use Saturn cloud.</p>\n<p><strong>CV Scheme</strong></p>\n<p>Since the organizer initially provided the information how public and private parts are splitted, we decided to utilize this info and design our local validation as similar as possible to the private part. So, we can say that the public part of the leaderboard was of little interest to us.</p>\n<p>The validation scheme is shown in the picture below (for citeseq part). Since we have three days in train dataset (citeseq task), the number of splits equals 6, for the multiome part the scheme was the same. However, there were 9 splits due to the larger number of days (at the very end we slightly modified this scheme so that validation fold always contained only one (nearest) day).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1243561%2F761304b2456999a0a31225be67858207%2Fcv_citeseq.png?generation=1668591942826347&amp;alt=media\" alt=\"\"></p>\n<p><strong>How to submit</strong></p>\n<p>Since the last available day (in train dataset) was always in validation folds, we could not use this validation scheme for the traditional folds blending. So, when we wanted to make submit, we used the usual KFold validation with shuffle.<br>\nThe approach was the following: all hypotheses and experiments, including hyperparameters tuning, were tested on true cv and afterwards these models were retrained on KFold.</p>\n<p><strong>Features and dimensionality reduction</strong></p>\n<p>We used PCA for dimensionality reduction for both tasks. Autoencoder was tested on citeseq dataset too, but performed slightly worse than PCA.<br>\nWe also used \"important\" features in their raw form for citeseq task.<br>\nMoreover we created some additional features based on aggregations of \"important\" columns over metadata (for example, \"mean_feature_1_by_donor\"). Such features gave us + 0.0003 for GBDT model on local CV.</p>\n<p><strong>Models (ensemble)</strong></p>\n<p><em>Citeseq task</em> : 3x multilayer perceptrons, 1x Conv1D, 1x pyBoost (best single model).<br>\n<em>Multiome task</em> : 1x multilayer perceptron, 1x TabNet, 1x pyBoost (best single model).</p>\n<p>pyBoost seems to be the new SOTA on multioutput tasks (at least among GBDT models).<br>\nIt's extremely fast to train as it uses GPU only and super easy to customize.<br>\n<a href=\"https://openreview.net/forum?id=WSxarC8t-T\" target=\"_blank\">Paper</a><br>\n<a href=\"https://github.com/sb-ai-lab/Py-Boost\" target=\"_blank\">Code</a></p>\n<p><strong>Some remarks about models</strong></p>\n<ul>\n<li><p><em>Multiome task</em>: all neural nets had 23418 output neurons. For pyBoost we reduced targets' dimension to 64 components using PCA.</p></li>\n<li><p>pyBoost was the best single model on True CV and KFold validation on citeseq data.</p></li>\n<li><p>For multiome task, it was the best model according to True CV and the<br>\nworst by KFold.</p></li>\n<li><p>We noticed that splitting targets  into groups and building a separate pyBoost model for each group improved our local CV a lot. By default, pyBoost can split targets into groups randomly, so we decided to try to improve it by splitting targets into groups based on their clusters, however, in the end it worked nearly the same as random splitting.</p></li>\n</ul>\n<p><strong>Data used to train</strong></p>\n<p><em>Citeseq task</em> : all available data<br>\n<em>Multiome task</em> : day 7 data only.</p>\n<p>Solving multiome task, we noticed that there is a significant performance drop on day 7 (on True CV).<br>\nThere was an idea that the reason is data drift in time, so we tried to train the model not on all available days, but only on the last available one. Locally, this improved our score by + 0.02. However, the problem was that there were no unseen days on the public leaderboard,<br>\nso training the model only on the last available day seriously dropped our public score (from 0.814 to 0.808). Nevertheless, we decided to follow the mantra \"trust your cv\" and as a result, this particular submission became our best on private.</p>\n<p>We also conducted a study on the similarity of days and found out that among available training days, day 3 is the most similar to private day 10 (day 7 is the second most similar).<br>\nNevertheless, since for other days the most similar day was always the closest in time, we decided to train our models on day 7.</p>",
      "rawMarkdown": "First of all, we would like to thank the organizers for an interesting challenge, as well as for the opportunity to use Saturn cloud.\n\n**CV Scheme**\n\nSince the organizer initially provided the information how public and private parts are splitted, we decided to utilize this info and design our local validation as similar as possible to the private part. So, we can say that the public part of the leaderboard was of little interest to us.\n\nThe validation scheme is shown in the picture below (for citeseq part). Since we have three days in train dataset (citeseq task), the number of splits equals 6, for the multiome part the scheme was the same. However, there were 9 splits due to the larger number of days (at the very end we slightly modified this scheme so that validation fold always contained only one (nearest) day).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1243561%2F761304b2456999a0a31225be67858207%2Fcv_citeseq.png?generation=1668591942826347&alt=media)\n\n\n**How to submit**\n\nSince the last available day (in train dataset) was always in validation folds, we could not use this validation scheme for the traditional folds blending. So, when we wanted to make submit, we used the usual KFold validation with shuffle.\nThe approach was the following: all hypotheses and experiments, including hyperparameters tuning, were tested on true cv and afterwards these models were retrained on KFold.\n\n\n**Features and dimensionality reduction**\n\nWe used PCA for dimensionality reduction for both tasks. Autoencoder was tested on citeseq dataset too, but performed slightly worse than PCA.\nWe also used \"important\" features in their raw form for citeseq task.\nMoreover we created some additional features based on aggregations of \"important\" columns over metadata (for example, \"mean_feature_1_by_donor\"). Such features gave us + 0.0003 for GBDT model on local CV.\n\n\n**Models (ensemble)**\n\n*Citeseq task* : 3x multilayer perceptrons, 1x Conv1D, 1x pyBoost (best single model).\n*Multiome task* : 1x multilayer perceptron, 1x TabNet, 1x pyBoost (best single model).\n\npyBoost seems to be the new SOTA on multioutput tasks (at least among GBDT models).\nIt's extremely fast to train as it uses GPU only and super easy to customize.\n[Paper](https://openreview.net/forum?id=WSxarC8t-T)\n[Code](https://github.com/sb-ai-lab/Py-Boost)\n\n\n**Some remarks about models**\n\n- *Multiome task*: all neural nets had 23418 output neurons. For pyBoost we reduced targets' dimension to 64 components using PCA.\n\n- pyBoost was the best single model on True CV and KFold validation on citeseq data.\n\n- For multiome task, it was the best model according to True CV and the\nworst by KFold.\n\n- We noticed that splitting targets  into groups and building a separate pyBoost model for each group improved our local CV a lot. By default, pyBoost can split targets into groups randomly, so we decided to try to improve it by splitting targets into groups based on their clusters, however, in the end it worked nearly the same as random splitting.\n\n\n**Data used to train**\n\n*Citeseq task* : all available data\n*Multiome task* : day 7 data only.\n\nSolving multiome task, we noticed that there is a significant performance drop on day 7 (on True CV).\nThere was an idea that the reason is data drift in time, so we tried to train the model not on all available days, but only on the last available one. Locally, this improved our score by + 0.02. However, the problem was that there were no unseen days on the public leaderboard,\nso training the model only on the last available day seriously dropped our public score (from 0.814 to 0.808). Nevertheless, we decided to follow the mantra \"trust your cv\" and as a result, this particular submission became our best on private.\n\nWe also conducted a study on the similarity of days and found out that among available training days, day 3 is the most similar to private day 10 (day 7 is the second most similar).\nNevertheless, since for other days the most similar day was always the closest in time, we decided to train our models on day 7.",
      "votes": null
    },
    {
      "id": "2031943",
      "postDate": "11/16/2022 10:22:08",
      "content": "<p>Great ! Thanks for the summary and congratulations with the Gold ! You fairly deserved it! </p>\n<p>Question 1: </p>\n<p>how did you measure the quality of the model - AVERAGE scores over folds, is it correct ? Or some weights for folds - since sizes are different ?  </p>\n<p>In classical cv we have two options - average over the folds VS create Out-Of-Fold prediction and score(Y_true, Y_oof) (without fold averaging).<br>\nIn that scheme you do not have OOF, is it correct ?</p>\n<p>Question 2: <br>\nHave you considered blend and tuning/optimizing coefficients for blend ? <br>\nIf yes - how do approach it ? <br>\nIs it - blending over each validation fold and again taking the average ? </p>\n<p>Question 3: <br>\nDo you think about having something like a additional  \"holdout\"  - because if you do something like extensive hyperparameter search or feature selection over CV you might get better results due to randomness ( similar  problem in statistics is resolved by<br>\nBonferroni correction -  which should depend on  the number of parameter trials <a href=\"https://en.wikipedia.org/wiki/Bonferroni_correction\" target=\"_blank\">https://en.wikipedia.org/wiki/Bonferroni_correction</a> ).<br>\nSo to control it - one may introduce \"holdout\" - which should NOT be seen during hyperparam optiomization/feature selection. </p>\n<p>===============</p>\n<p>Remarks: <br>\nthe proposed CV scheme <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a> <br>\nis quite similar to yours. <br>\nIn particular you folds 3,4 are exactly the same as in our scheme</p>\n<ul>\n<li>you excluded from the train day 4 and donors 31800 for one fold and 32606 for other.<br>\nSo our idea was just to be symmetric - exclude day 2, exclude 3 - and do the same as you for excluded day 4  - and thus we  get 6 folds.</li>\n</ul>\n<p>With our CV we can get two interesting things:</p>\n<p>1) We can create \"public like score\" and \"private like score\" - one corresponds to predict other  donor , another to predict the other day</p>\n<p>2) We can make  Y_{oof} of the FULL SIZE - so we can score , not by averaging over folds -  but like in classical scheme,<br>\nto do that - we need a trick:<br>\nwe average predictions from two subfolds, e.g. OOF_{day4} =  1/2 ( Y_{pred from days 2,3 exclude 31800} +   Y_{pred from days 2,3 exclude 32606} )  - so difference with the classical scheme that we need to average predictions to form the Y_{oof} !</p>\n<p>Not sure I am very clear, may be the code is more clear:</p>\n<p>Create folds:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=21</a></p>\n<p>Create predictions: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=26\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=26</a> </p>\n<p>====</p>\n<p>To conclude in some sense your CV scheme is more \"stronger\" - you train one just single day and required the model to perform good on TWO days (your folds 1-3),   but our CV in some sense is more general - you can use it for any similar task when  you have two groups and \"two leaderbords\" - public and private - we have scores for the both of them. </p>",
      "rawMarkdown": "Great ! Thanks for the summary and congratulations with the Gold ! You fairly deserved it! \n\nQuestion 1: \n\nhow did you measure the quality of the model - AVERAGE scores over folds, is it correct ? Or some weights for folds - since sizes are different ?  \n\nIn classical cv we have two options - average over the folds VS create Out-Of-Fold prediction and score(Y_true, Y_oof) (without fold averaging).\nIn that scheme you do not have OOF, is it correct ?\n\nQuestion 2: \nHave you considered blend and tuning/optimizing coefficients for blend ? \nIf yes - how do approach it ? \nIs it - blending over each validation fold and again taking the average ? \n\nQuestion 3: \nDo you think about having something like a additional  \"holdout\"  - because if you do something like extensive hyperparameter search or feature selection over CV you might get better results due to randomness ( similar  problem in statistics is resolved by\nBonferroni correction -  which should depend on  the number of parameter trials https://en.wikipedia.org/wiki/Bonferroni_correction ).\nSo to control it - one may introduce \"holdout\" - which should NOT be seen during hyperparam optiomization/feature selection. \n\n\n===============\n\nRemarks: \nthe proposed CV scheme https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860 \nis quite similar to yours. \nIn particular you folds 3,4 are exactly the same as in our scheme\n- you excluded from the train day 4 and donors 31800 for one fold and 32606 for other.\nSo our idea was just to be symmetric - exclude day 2, exclude 3 - and do the same as you for excluded day 4  - and thus we  get 6 folds.\n\nWith our CV we can get two interesting things:\n\n1) We can create \"public like score\" and \"private like score\" - one corresponds to predict other  donor , another to predict the other day\n\n2) We can make  Y_{oof} of the FULL SIZE - so we can score , not by averaging over folds -  but like in classical scheme,\nto do that - we need a trick:\nwe average predictions from two subfolds, e.g. OOF_{day4} =  1/2 ( Y_{pred from days 2,3 exclude 31800} +   Y_{pred from days 2,3 exclude 32606} )  - so difference with the classical scheme that we need to average predictions to form the Y_{oof} !\n\nNot sure I am very clear, may be the code is more clear:\n\nCreate folds:\nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&cellId=21\n\nCreate predictions: \nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&cellId=26 \n\n ====\n\nTo conclude in some sense your CV scheme is more \"stronger\" - you train one just single day and required the model to perform good on TWO days (your folds 1-3),   but our CV in some sense is more general - you can use it for any similar task when  you have two groups and \"two leaderbords\" - public and private - we have scores for the both of them.",
      "votes": null
    },
    {
      "id": "2031973",
      "postDate": "11/16/2022 10:38:44",
      "content": "<p>Answer_1: yep, you're right, we measured model performance by taking simple average.</p>\n<p>Answer_2: blend weights were tuned with optuna (we've tried out stacking with pyBoost but in the end simple linear combination of predictions worked better for us).</p>\n<p>Since the dataset with predictions was quite large, the weights were selected not for all folds, but only for a part of them.</p>",
      "rawMarkdown": "Answer_1: yep, you're right, we measured model performance by taking simple average.\n\nAnswer_2: blend weights were tuned with optuna (we've tried out stacking with pyBoost but in the end simple linear combination of predictions worked better for us).\n\nSince the dataset with predictions was quite large, the weights were selected not for all folds, but only for a part of them.",
      "votes": null
    },
    {
      "id": "2032002",
      "postDate": "11/16/2022 11:01:15",
      "content": "<p>Thanks for answers ! <br>\nFor the answer 2 - so you blend on each fold, and average the scores from each fold - that what you need to optimize with Optuna - is it correct ? <br>\nPS<br>\nAlso added the question 3 - how to control possible overfit to validation scheme - if extensive hyperparam search / feature selection was done - please take a look when you have free time.  </p>",
      "rawMarkdown": "Thanks for answers ! \nFor the answer 2 - so you blend on each fold, and average the scores from each fold - that what you need to optimize with Optuna - is it correct ? \nPS\nAlso added the question 3 - how to control possible overfit to validation scheme - if extensive hyperparam search / feature selection was done - please take a look when you have free time.",
      "votes": null
    },
    {
      "id": "2032317",
      "postDate": "11/16/2022 14:53:06",
      "content": "<p>Congrats! I got a question that how do u compute the similarity of days? Would u introduce more details. Thanks for sharing, u did a great job.</p>",
      "rawMarkdown": "Congrats! I got a question that how do u compute the similarity of days? Would u introduce more details. Thanks for sharing, u did a great job.",
      "votes": null
    },
    {
      "id": "2032344",
      "postDate": "11/16/2022 15:12:05",
      "content": "<p>Thank you very much for the kind words!</p>\n<p>In fact, the most similar day to private day 10 was determined in two ways:</p>\n<ol>\n<li><p>Visual analysis of the distribution of individual features + comparing the number of non-zero elements for different days</p></li>\n<li><p>We trained CatBoost classifier that tried to distinguish the day from the training dataset from day 10. The day with lowest ROC AUC we considered as the most similar to day 10.</p></li>\n</ol>\n<p>According to both of these approaches, day 3 looked as the most similar day (the next one was day 7)</p>",
      "rawMarkdown": "Thank you very much for the kind words!\n\nIn fact, the most similar day to private day 10 was determined in two ways:\n\n1. Visual analysis of the distribution of individual features + comparing the number of non-zero elements for different days\n\n2. We trained CatBoost classifier that tried to distinguish the day from the training dataset from day 10. The day with lowest ROC AUC we considered as the most similar to day 10.\n\nAccording to both of these approaches, day 3 looked as the most similar day (the next one was day 7)",
      "votes": null
    },
    {
      "id": "2033589",
      "postDate": "11/17/2022 11:30:46",
      "content": "<p>nice catch at Multiome CV. congratulations for your results</p>",
      "rawMarkdown": "nice catch at Multiome CV. congratulations for your results",
      "votes": null
    },
    {
      "id": "2034316",
      "postDate": "11/18/2022 03:26:43",
      "content": "<p>Thank you for publishing this.<br>\nI have one question. <br>\nHow did you decide to use Kfold for the final submission? I mean, how were you convinced that using the features and parameters which scored highest in the CV mimicking Private would also score well in Kfold?</p>",
      "rawMarkdown": "Thank you for publishing this.\nI have one question. \nHow did you decide to use Kfold for the final submission? I mean, how were you convinced that using the features and parameters which scored highest in the CV mimicking Private would also score well in Kfold?",
      "votes": null
    },
    {
      "id": "2034354",
      "postDate": "11/18/2022 04:42:47",
      "content": "<p>Congratulations on your finish.  Thanks for sharing the links on PyBoost.</p>\n<p>Wondered about the days and differences or similarities.  Day 7 consistently had the lowest number of cell ids for all Donors, 6960-7466 and seemed to score poorly for all, best for the donor with the highest number. Whereas Day 4 which had more cell ids for the Train Donors (no day 4 in Public though), 9407-10933 always seemed to score the best in folds.  If trying to balance the number of cell ids or generate some artificial data or from other sources might have helped models. </p>\n<p>Did you find any donor for day 10 that was more similar to the Test donor?   </p>",
      "rawMarkdown": "Congratulations on your finish.  Thanks for sharing the links on PyBoost.\n\nWondered about the days and differences or similarities.  Day 7 consistently had the lowest number of cell ids for all Donors, 6960-7466 and seemed to score poorly for all, best for the donor with the highest number. Whereas Day 4 which had more cell ids for the Train Donors (no day 4 in Public though), 9407-10933 always seemed to score the best in folds.  If trying to balance the number of cell ids or generate some artificial data or from other sources might have helped models. \n\nDid you find any donor for day 10 that was more similar to the Test donor?",
      "votes": null
    },
    {
      "id": "2034604",
      "postDate": "11/18/2022 09:23:07",
      "content": "<p>Hello!</p>\n<p>We were not interested in how our model performs on KFold cv. The only purpose of its usage was the possibility to train multiple models and blend them (instead of training one model on the whole train dataset). As an alternative for Kfold validation, a common bootstrap could also be used here.</p>",
      "rawMarkdown": "Hello!\n\nWe were not interested in how our model performs on KFold cv. The only purpose of its usage was the possibility to train multiple models and blend them (instead of training one model on the whole train dataset). As an alternative for Kfold validation, a common bootstrap could also be used here.",
      "votes": null
    },
    {
      "id": "2034608",
      "postDate": "11/18/2022 09:26:30",
      "content": "<p>Good question! Frankly speaking, no. As we observed model performance across all folds, it seemed that new day was always much more problematic than new donor. So we concentrated on time component of this problem only.</p>",
      "rawMarkdown": "Good question! Frankly speaking, no. As we observed model performance across all folds, it seemed that new day was always much more problematic than new donor. So we concentrated on time component of this problem only.",
      "votes": null
    },
    {
      "id": "2035546",
      "postDate": "11/19/2022 02:22:26",
      "content": "<p>Great! thanks for sharing.</p>",
      "rawMarkdown": "Great! thanks for sharing.",
      "votes": null
    },
    {
      "id": "2042890",
      "postDate": "11/25/2022 05:25:01",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031943,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/16/2022 10:22:08",
      "content": "<p>Great ! Thanks for the summary and congratulations with the Gold ! You fairly deserved it! </p>\n<p>Question 1: </p>\n<p>how did you measure the quality of the model - AVERAGE scores over folds, is it correct ? Or some weights for folds - since sizes are different ?  </p>\n<p>In classical cv we have two options - average over the folds VS create Out-Of-Fold prediction and score(Y_true, Y_oof) (without fold averaging).<br>\nIn that scheme you do not have OOF, is it correct ?</p>\n<p>Question 2: <br>\nHave you considered blend and tuning/optimizing coefficients for blend ? <br>\nIf yes - how do approach it ? <br>\nIs it - blending over each validation fold and again taking the average ? </p>\n<p>Question 3: <br>\nDo you think about having something like a additional  \"holdout\"  - because if you do something like extensive hyperparameter search or feature selection over CV you might get better results due to randomness ( similar  problem in statistics is resolved by<br>\nBonferroni correction -  which should depend on  the number of parameter trials <a href=\"https://en.wikipedia.org/wiki/Bonferroni_correction\" target=\"_blank\">https://en.wikipedia.org/wiki/Bonferroni_correction</a> ).<br>\nSo to control it - one may introduce \"holdout\" - which should NOT be seen during hyperparam optiomization/feature selection. </p>\n<p>===============</p>\n<p>Remarks: <br>\nthe proposed CV scheme <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860</a> <br>\nis quite similar to yours. <br>\nIn particular you folds 3,4 are exactly the same as in our scheme</p>\n<ul>\n<li>you excluded from the train day 4 and donors 31800 for one fold and 32606 for other.<br>\nSo our idea was just to be symmetric - exclude day 2, exclude 3 - and do the same as you for excluded day 4  - and thus we  get 6 folds.</li>\n</ul>\n<p>With our CV we can get two interesting things:</p>\n<p>1) We can create \"public like score\" and \"private like score\" - one corresponds to predict other  donor , another to predict the other day</p>\n<p>2) We can make  Y_{oof} of the FULL SIZE - so we can score , not by averaging over folds -  but like in classical scheme,<br>\nto do that - we need a trick:<br>\nwe average predictions from two subfolds, e.g. OOF_{day4} =  1/2 ( Y_{pred from days 2,3 exclude 31800} +   Y_{pred from days 2,3 exclude 32606} )  - so difference with the classical scheme that we need to average predictions to form the Y_{oof} !</p>\n<p>Not sure I am very clear, may be the code is more clear:</p>\n<p>Create folds:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=21</a></p>\n<p>Create predictions: <br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=26\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&amp;cellId=26</a> </p>\n<p>====</p>\n<p>To conclude in some sense your CV scheme is more \"stronger\" - you train one just single day and required the model to perform good on TWO days (your folds 1-3),   but our CV in some sense is more general - you can use it for any similar task when  you have two groups and \"two leaderbords\" - public and private - we have scores for the both of them. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2031973,
          "author_name": "l0glikelihood",
          "author_url": "",
          "post_date": "11/16/2022 10:38:44",
          "content": "<p>Answer_1: yep, you're right, we measured model performance by taking simple average.</p>\n<p>Answer_2: blend weights were tuned with optuna (we've tried out stacking with pyBoost but in the end simple linear combination of predictions worked better for us).</p>\n<p>Since the dataset with predictions was quite large, the weights were selected not for all folds, but only for a part of them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2032002,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "11/16/2022 11:01:15",
          "content": "<p>Thanks for answers ! <br>\nFor the answer 2 - so you blend on each fold, and average the scores from each fold - that what you need to optimize with Optuna - is it correct ? <br>\nPS<br>\nAlso added the question 3 - how to control possible overfit to validation scheme - if extensive hyperparam search / feature selection was done - please take a look when you have free time.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2032317,
      "author_name": "alvinai9603",
      "author_url": "",
      "post_date": "11/16/2022 14:53:06",
      "content": "<p>Congrats! I got a question that how do u compute the similarity of days? Would u introduce more details. Thanks for sharing, u did a great job.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2032344,
          "author_name": "l0glikelihood",
          "author_url": "",
          "post_date": "11/16/2022 15:12:05",
          "content": "<p>Thank you very much for the kind words!</p>\n<p>In fact, the most similar day to private day 10 was determined in two ways:</p>\n<ol>\n<li><p>Visual analysis of the distribution of individual features + comparing the number of non-zero elements for different days</p></li>\n<li><p>We trained CatBoost classifier that tried to distinguish the day from the training dataset from day 10. The day with lowest ROC AUC we considered as the most similar to day 10.</p></li>\n</ol>\n<p>According to both of these approaches, day 3 looked as the most similar day (the next one was day 7)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2035546,
          "author_name": "alvinai9603",
          "author_url": "",
          "post_date": "11/19/2022 02:22:26",
          "content": "<p>Great! thanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2033589,
      "author_name": "leandrodestefani",
      "author_url": "",
      "post_date": "11/17/2022 11:30:46",
      "content": "<p>nice catch at Multiome CV. congratulations for your results</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2034316,
      "author_name": "aesoptacit",
      "author_url": "",
      "post_date": "11/18/2022 03:26:43",
      "content": "<p>Thank you for publishing this.<br>\nI have one question. <br>\nHow did you decide to use Kfold for the final submission? I mean, how were you convinced that using the features and parameters which scored highest in the CV mimicking Private would also score well in Kfold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2034604,
          "author_name": "l0glikelihood",
          "author_url": "",
          "post_date": "11/18/2022 09:23:07",
          "content": "<p>Hello!</p>\n<p>We were not interested in how our model performs on KFold cv. The only purpose of its usage was the possibility to train multiple models and blend them (instead of training one model on the whole train dataset). As an alternative for Kfold validation, a common bootstrap could also be used here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2042890,
          "author_name": "aesoptacit",
          "author_url": "",
          "post_date": "11/25/2022 05:25:01",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2034354,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "11/18/2022 04:42:47",
      "content": "<p>Congratulations on your finish.  Thanks for sharing the links on PyBoost.</p>\n<p>Wondered about the days and differences or similarities.  Day 7 consistently had the lowest number of cell ids for all Donors, 6960-7466 and seemed to score poorly for all, best for the donor with the highest number. Whereas Day 4 which had more cell ids for the Train Donors (no day 4 in Public though), 9407-10933 always seemed to score the best in folds.  If trying to balance the number of cell ids or generate some artificial data or from other sources might have helped models. </p>\n<p>Did you find any donor for day 10 that was more similar to the Test donor?   </p>",
      "votes": null,
      "replies": [
        {
          "id": 2034608,
          "author_name": "l0glikelihood",
          "author_url": "",
          "post_date": "11/18/2022 09:26:30",
          "content": "<p>Good question! Frankly speaking, no. As we observed model performance across all folds, it seemed that new day was always much more problematic than new donor. So we concentrated on time component of this problem only.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2031889": "First of all, we would like to thank the organizers for an interesting challenge, as well as for the opportunity to use Saturn cloud.\n\n**CV Scheme**\n\nSince the organizer initially provided the information how public and private parts are splitted, we decided to utilize this info and design our local validation as similar as possible to the private part. So, we can say that the public part of the leaderboard was of little interest to us.\n\nThe validation scheme is shown in the picture below (for citeseq part). Since we have three days in train dataset (citeseq task), the number of splits equals 6, for the multiome part the scheme was the same. However, there were 9 splits due to the larger number of days (at the very end we slightly modified this scheme so that validation fold always contained only one (nearest) day).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1243561%2F761304b2456999a0a31225be67858207%2Fcv_citeseq.png?generation=1668591942826347&alt=media)\n\n\n**How to submit**\n\nSince the last available day (in train dataset) was always in validation folds, we could not use this validation scheme for the traditional folds blending. So, when we wanted to make submit, we used the usual KFold validation with shuffle.\nThe approach was the following: all hypotheses and experiments, including hyperparameters tuning, were tested on true cv and afterwards these models were retrained on KFold.\n\n\n**Features and dimensionality reduction**\n\nWe used PCA for dimensionality reduction for both tasks. Autoencoder was tested on citeseq dataset too, but performed slightly worse than PCA.\nWe also used \"important\" features in their raw form for citeseq task.\nMoreover we created some additional features based on aggregations of \"important\" columns over metadata (for example, \"mean_feature_1_by_donor\"). Such features gave us + 0.0003 for GBDT model on local CV.\n\n\n**Models (ensemble)**\n\n*Citeseq task* : 3x multilayer perceptrons, 1x Conv1D, 1x pyBoost (best single model).\n*Multiome task* : 1x multilayer perceptron, 1x TabNet, 1x pyBoost (best single model).\n\npyBoost seems to be the new SOTA on multioutput tasks (at least among GBDT models).\nIt's extremely fast to train as it uses GPU only and super easy to customize.\n[Paper](https://openreview.net/forum?id=WSxarC8t-T)\n[Code](https://github.com/sb-ai-lab/Py-Boost)\n\n\n**Some remarks about models**\n\n- *Multiome task*: all neural nets had 23418 output neurons. For pyBoost we reduced targets' dimension to 64 components using PCA.\n\n- pyBoost was the best single model on True CV and KFold validation on citeseq data.\n\n- For multiome task, it was the best model according to True CV and the\nworst by KFold.\n\n- We noticed that splitting targets  into groups and building a separate pyBoost model for each group improved our local CV a lot. By default, pyBoost can split targets into groups randomly, so we decided to try to improve it by splitting targets into groups based on their clusters, however, in the end it worked nearly the same as random splitting.\n\n\n**Data used to train**\n\n*Citeseq task* : all available data\n*Multiome task* : day 7 data only.\n\nSolving multiome task, we noticed that there is a significant performance drop on day 7 (on True CV).\nThere was an idea that the reason is data drift in time, so we tried to train the model not on all available days, but only on the last available one. Locally, this improved our score by + 0.02. However, the problem was that there were no unseen days on the public leaderboard,\nso training the model only on the last available day seriously dropped our public score (from 0.814 to 0.808). Nevertheless, we decided to follow the mantra \"trust your cv\" and as a result, this particular submission became our best on private.\n\nWe also conducted a study on the similarity of days and found out that among available training days, day 3 is the most similar to private day 10 (day 7 is the second most similar).\nNevertheless, since for other days the most similar day was always the closest in time, we decided to train our models on day 7.",
    "2031943": "Great ! Thanks for the summary and congratulations with the Gold ! You fairly deserved it! \n\nQuestion 1: \n\nhow did you measure the quality of the model - AVERAGE scores over folds, is it correct ? Or some weights for folds - since sizes are different ?  \n\nIn classical cv we have two options - average over the folds VS create Out-Of-Fold prediction and score(Y_true, Y_oof) (without fold averaging).\nIn that scheme you do not have OOF, is it correct ?\n\nQuestion 2: \nHave you considered blend and tuning/optimizing coefficients for blend ? \nIf yes - how do approach it ? \nIs it - blending over each validation fold and again taking the average ? \n\nQuestion 3: \nDo you think about having something like a additional  \"holdout\"  - because if you do something like extensive hyperparameter search or feature selection over CV you might get better results due to randomness ( similar  problem in statistics is resolved by\nBonferroni correction -  which should depend on  the number of parameter trials https://en.wikipedia.org/wiki/Bonferroni_correction ).\nSo to control it - one may introduce \"holdout\" - which should NOT be seen during hyperparam optiomization/feature selection. \n\n\n===============\n\nRemarks: \nthe proposed CV scheme https://www.kaggle.com/competitions/open-problems-multimodal/discussion/358860 \nis quite similar to yours. \nIn particular you folds 3,4 are exactly the same as in our scheme\n- you excluded from the train day 4 and donors 31800 for one fold and 32606 for other.\nSo our idea was just to be symmetric - exclude day 2, exclude 3 - and do the same as you for excluded day 4  - and thus we  get 6 folds.\n\nWith our CV we can get two interesting things:\n\n1) We can create \"public like score\" and \"private like score\" - one corresponds to predict other  donor , another to predict the other day\n\n2) We can make  Y_{oof} of the FULL SIZE - so we can score , not by averaging over folds -  but like in classical scheme,\nto do that - we need a trick:\nwe average predictions from two subfolds, e.g. OOF_{day4} =  1/2 ( Y_{pred from days 2,3 exclude 31800} +   Y_{pred from days 2,3 exclude 32606} )  - so difference with the classical scheme that we need to average predictions to form the Y_{oof} !\n\nNot sure I am very clear, may be the code is more clear:\n\nCreate folds:\nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&cellId=21\n\nCreate predictions: \nhttps://www.kaggle.com/code/alexandervc/mmscel-cv-modeling-advanced?scriptVersionId=111042574&cellId=26 \n\n ====\n\nTo conclude in some sense your CV scheme is more \"stronger\" - you train one just single day and required the model to perform good on TWO days (your folds 1-3),   but our CV in some sense is more general - you can use it for any similar task when  you have two groups and \"two leaderbords\" - public and private - we have scores for the both of them.",
    "2031973": "Answer_1: yep, you're right, we measured model performance by taking simple average.\n\nAnswer_2: blend weights were tuned with optuna (we've tried out stacking with pyBoost but in the end simple linear combination of predictions worked better for us).\n\nSince the dataset with predictions was quite large, the weights were selected not for all folds, but only for a part of them.",
    "2032002": "Thanks for answers ! \nFor the answer 2 - so you blend on each fold, and average the scores from each fold - that what you need to optimize with Optuna - is it correct ? \nPS\nAlso added the question 3 - how to control possible overfit to validation scheme - if extensive hyperparam search / feature selection was done - please take a look when you have free time.",
    "2032317": "Congrats! I got a question that how do u compute the similarity of days? Would u introduce more details. Thanks for sharing, u did a great job.",
    "2032344": "Thank you very much for the kind words!\n\nIn fact, the most similar day to private day 10 was determined in two ways:\n\n1. Visual analysis of the distribution of individual features + comparing the number of non-zero elements for different days\n\n2. We trained CatBoost classifier that tried to distinguish the day from the training dataset from day 10. The day with lowest ROC AUC we considered as the most similar to day 10.\n\nAccording to both of these approaches, day 3 looked as the most similar day (the next one was day 7)",
    "2033589": "nice catch at Multiome CV. congratulations for your results",
    "2034316": "Thank you for publishing this.\nI have one question. \nHow did you decide to use Kfold for the final submission? I mean, how were you convinced that using the features and parameters which scored highest in the CV mimicking Private would also score well in Kfold?",
    "2034354": "Congratulations on your finish.  Thanks for sharing the links on PyBoost.\n\nWondered about the days and differences or similarities.  Day 7 consistently had the lowest number of cell ids for all Donors, 6960-7466 and seemed to score poorly for all, best for the donor with the highest number. Whereas Day 4 which had more cell ids for the Train Donors (no day 4 in Public though), 9407-10933 always seemed to score the best in folds.  If trying to balance the number of cell ids or generate some artificial data or from other sources might have helped models. \n\nDid you find any donor for day 10 that was more similar to the Test donor?",
    "2034604": "Hello!\n\nWe were not interested in how our model performs on KFold cv. The only purpose of its usage was the possibility to train multiple models and blend them (instead of training one model on the whole train dataset). As an alternative for Kfold validation, a common bootstrap could also be used here.",
    "2034608": "Good question! Frankly speaking, no. As we observed model performance across all folds, it seemed that new day was always much more problematic than new donor. So we concentrated on time component of this problem only.",
    "2035546": "Great! thanks for sharing.",
    "2042890": "Thanks a lot!"
  },
  "source": "meta"
}