{
  "id": 366392,
  "title": "Public 6th Private 14 Solution",
  "url": "/competitions/open-problems-multimodal/writeups/protein-shake-public-6th-private-14-solution",
  "author_name": "",
  "post_date": "2022-11-16T08:00:06.890Z",
  "votes": 23,
  "comment_count": 12,
  "views": 0,
  "content": "<h2>Intro</h2>\n<ul>\n<li>I have been mainly working on the Cite part. I tried many things, </li>\n<li>multi part by <a href=\"https://www.kaggle.com/paragkale\" target=\"_blank\">@paragkale</a> <a href=\"https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\" target=\"_blank\">https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion</a></li>\n<li>The following tricks gave me the most gain.</li>\n<li>I will update in this post about the code and what didn't work.</li>\n</ul>\n<h2>Extra Data</h2>\n<ul>\n<li>The raw count data released by the host.</li>\n</ul>\n<h2>Dimensionality Reduction</h2>\n<p>I think the most helpful one are:</p>\n<ul>\n<li>sklearn.decomposition.TruncatedSVD (128 comps)</li>\n<li>Self-made denosing auto encoder (128 hidden nodes)</li>\n</ul>\n<h2>Direct Features</h2>\n<ul>\n<li>Direct features based on matching the names</li>\n<li>Direct features based on absolute correlation to targets</li>\n<li>Direct features based on the list shared by the hosts in this thread <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366392\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366392</a></li>\n</ul>\n<h2>Use base models outcomes as NN inputs feature for ensembling</h2>\n<p>It is known that MSE is not a really good loss function for the competition metric. Therefore within each fold, we trained 4 base models and used their features for the NN input.</p>\n<ul>\n<li>sklearn.linear_model.Ridge</li>\n<li>sklearn.linear_model.MultiTaskElasticNet</li>\n<li>sklearn.kernel_ridge.KernelRidge</li>\n<li>sklearn.ensemble.HistGradientBoostingRegressor</li>\n</ul>\n<p>We added heavy noise to their predictions to make sure the NN can learn from other features as well</p>\n<pre><code>        self.blender = torch.nn.Sequential(\n            GaussianNoise(self.blend_noise),\n            torch.nn.Linear(out_dim * 4, 128),\n            torch.nn.LayerNorm(128),\n            activation(),\n            torch.nn.Dropout(self.blend_dropout),\n        )\n</code></pre>\n<h2>GroupK Cross Validation on Target Clusters</h2>\n<p>The tricks in this section increased both public and private LB, but we cannot compare the CV because it is a CV scheme change. Luckily it is (relatively, I guess?) performing well on both public and private.</p>\n<p>It is known that there are some subtle domain shifts between train, private and public test sets. However, the difficulty is that the shift is happening in at least 3 directions (donor, day, cell types). To create a hard but not too hard CV scheme, we find that clustering the target values performed very well on both of the public and private leaderboard.</p>\n<p>Let's consider the CV scheme selection as a spectrum:</p>\n<ul>\n<li>The easiest CV scheme: Random K fold (Downside: not representative of the test set)</li>\n<li>The hard CV scheme: GroupKfold by day/donor (Downside: too few fold to train)</li>\n<li>The hardest CV scheme 1: Time series split (Downside: wasting the last day data)</li>\n<li>The hardest CV scheme 2: Excluding the 1 day or 1 donor completely from the training set (Downside: too hard/defensive)</li>\n</ul>\n<p>Another reason of doing the clustering is that the day here is categorical, however in real life, time is continuous. GroupK CV by day is not that satisfying.</p>\n<p>The first image shows the target kmeans result (colors)  visualized with the tsvd targets (points):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fb5e02c4c3d9edc2029a1458d02976f34%2F1.png?generation=1668564701067041&amp;alt=media\" alt=\"\"></p>\n<p>Next, you can see the target clusters capture the cell type differences:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fa76444828dbe22e970e9af2a5909005a%2F2.png?generation=1668564719243229&amp;alt=media\" alt=\"\"></p>\n<p>And the shifts of day and donor are not that significant compared to the cell types in the context of target clustering:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fd061e807ed3799640c34629a225bc5cc%2F3.png?generation=1668564731817794&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F53706c2ecd0ffc45cdf32d2517b71a8a%2F4.png?generation=1668564744081492&amp;alt=media\" alt=\"\"></p>\n<h2>Regularization/Augmentation</h2>\n<p>The tricks in this section increased both CV and LB.</p>\n<h3>Seed-bagging</h3>\n<p>I think most people have done this, we trained the same model a few more times with the different seeds for blending.</p>\n<h3>Mixup Augmentation and Stochastic Weight Averaging</h3>\n<p>The training is done on roughly 3 stages</p>\n<h4>1. Mix up augmentation stage</h4>\n<p>Since all features are numerical values, mixup worked well.</p>\n<pre><code>def mixup_augmentation(x: torch.Tensor, y: torch.Tensor, alpha: float = 5):\n    lam = np.random.beta(alpha, alpha)\n    rand_idx = torch.randperm(x.shape[0])\n    mixed_x = lam * x + (1 - lam) * x[rand_idx, :]\n    target_a, target_b = y, y[rand_idx]\n    return mixed_x, target_a, target_b, lam\n</code></pre>\n<h4>2. Normal training stage</h4>\n<h4>3. SWA stage</h4>\n<p><a href=\"https://pytorch.org/docs/stable/optim.html#putting-it-all-together\" target=\"_blank\">https://pytorch.org/docs/stable/optim.html#putting-it-all-together</a><br>\nThis is similar to seed-bagging, I am not sure if they are overlapping or if they have their benefits here. </p>",
  "messages": [
    {
      "id": "2031325",
      "postDate": "11/16/2022 02:18:52",
      "content": "<h2>Intro</h2>\n<ul>\n<li>I have been mainly working on the Cite part. I tried many things, </li>\n<li>multi part by <a href=\"https://www.kaggle.com/paragkale\" target=\"_blank\">@paragkale</a> <a href=\"https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\" target=\"_blank\">https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion</a></li>\n<li>The following tricks gave me the most gain.</li>\n<li>I will update in this post about the code and what didn't work.</li>\n</ul>\n<h2>Extra Data</h2>\n<ul>\n<li>The raw count data released by the host.</li>\n</ul>\n<h2>Dimensionality Reduction</h2>\n<p>I think the most helpful one are:</p>\n<ul>\n<li>sklearn.decomposition.TruncatedSVD (128 comps)</li>\n<li>Self-made denosing auto encoder (128 hidden nodes)</li>\n</ul>\n<h2>Direct Features</h2>\n<ul>\n<li>Direct features based on matching the names</li>\n<li>Direct features based on absolute correlation to targets</li>\n<li>Direct features based on the list shared by the hosts in this thread <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366392\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366392</a></li>\n</ul>\n<h2>Use base models outcomes as NN inputs feature for ensembling</h2>\n<p>It is known that MSE is not a really good loss function for the competition metric. Therefore within each fold, we trained 4 base models and used their features for the NN input.</p>\n<ul>\n<li>sklearn.linear_model.Ridge</li>\n<li>sklearn.linear_model.MultiTaskElasticNet</li>\n<li>sklearn.kernel_ridge.KernelRidge</li>\n<li>sklearn.ensemble.HistGradientBoostingRegressor</li>\n</ul>\n<p>We added heavy noise to their predictions to make sure the NN can learn from other features as well</p>\n<pre><code>        self.blender = torch.nn.Sequential(\n            GaussianNoise(self.blend_noise),\n            torch.nn.Linear(out_dim * 4, 128),\n            torch.nn.LayerNorm(128),\n            activation(),\n            torch.nn.Dropout(self.blend_dropout),\n        )\n</code></pre>\n<h2>GroupK Cross Validation on Target Clusters</h2>\n<p>The tricks in this section increased both public and private LB, but we cannot compare the CV because it is a CV scheme change. Luckily it is (relatively, I guess?) performing well on both public and private.</p>\n<p>It is known that there are some subtle domain shifts between train, private and public test sets. However, the difficulty is that the shift is happening in at least 3 directions (donor, day, cell types). To create a hard but not too hard CV scheme, we find that clustering the target values performed very well on both of the public and private leaderboard.</p>\n<p>Let's consider the CV scheme selection as a spectrum:</p>\n<ul>\n<li>The easiest CV scheme: Random K fold (Downside: not representative of the test set)</li>\n<li>The hard CV scheme: GroupKfold by day/donor (Downside: too few fold to train)</li>\n<li>The hardest CV scheme 1: Time series split (Downside: wasting the last day data)</li>\n<li>The hardest CV scheme 2: Excluding the 1 day or 1 donor completely from the training set (Downside: too hard/defensive)</li>\n</ul>\n<p>Another reason of doing the clustering is that the day here is categorical, however in real life, time is continuous. GroupK CV by day is not that satisfying.</p>\n<p>The first image shows the target kmeans result (colors)  visualized with the tsvd targets (points):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fb5e02c4c3d9edc2029a1458d02976f34%2F1.png?generation=1668564701067041&amp;alt=media\" alt=\"\"></p>\n<p>Next, you can see the target clusters capture the cell type differences:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fa76444828dbe22e970e9af2a5909005a%2F2.png?generation=1668564719243229&amp;alt=media\" alt=\"\"></p>\n<p>And the shifts of day and donor are not that significant compared to the cell types in the context of target clustering:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fd061e807ed3799640c34629a225bc5cc%2F3.png?generation=1668564731817794&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F53706c2ecd0ffc45cdf32d2517b71a8a%2F4.png?generation=1668564744081492&amp;alt=media\" alt=\"\"></p>\n<h2>Regularization/Augmentation</h2>\n<p>The tricks in this section increased both CV and LB.</p>\n<h3>Seed-bagging</h3>\n<p>I think most people have done this, we trained the same model a few more times with the different seeds for blending.</p>\n<h3>Mixup Augmentation and Stochastic Weight Averaging</h3>\n<p>The training is done on roughly 3 stages</p>\n<h4>1. Mix up augmentation stage</h4>\n<p>Since all features are numerical values, mixup worked well.</p>\n<pre><code>def mixup_augmentation(x: torch.Tensor, y: torch.Tensor, alpha: float = 5):\n    lam = np.random.beta(alpha, alpha)\n    rand_idx = torch.randperm(x.shape[0])\n    mixed_x = lam * x + (1 - lam) * x[rand_idx, :]\n    target_a, target_b = y, y[rand_idx]\n    return mixed_x, target_a, target_b, lam\n</code></pre>\n<h4>2. Normal training stage</h4>\n<h4>3. SWA stage</h4>\n<p><a href=\"https://pytorch.org/docs/stable/optim.html#putting-it-all-together\" target=\"_blank\">https://pytorch.org/docs/stable/optim.html#putting-it-all-together</a><br>\nThis is similar to seed-bagging, I am not sure if they are overlapping or if they have their benefits here. </p>",
      "rawMarkdown": "## Intro\n- I have been mainly working on the Cite part. I tried many things, \n- multi part by @paragkale https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\n- The following tricks gave me the most gain.\n- I will update in this post about the code and what didn't work.\n\n## Extra Data\n- The raw count data released by the host.\n\n##Dimensionality Reduction\n\nI think the most helpful one are:\n- sklearn.decomposition.TruncatedSVD (128 comps)\n- Self-made denosing auto encoder (128 hidden nodes)\n\n## Direct Features\n- Direct features based on matching the names\n- Direct features based on absolute correlation to targets\n- Direct features based on the list shared by the hosts in this thread https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366392\n\n## Use base models outcomes as NN inputs feature for ensembling\n\nIt is known that MSE is not a really good loss function for the competition metric. Therefore within each fold, we trained 4 base models and used their features for the NN input.\n- sklearn.linear_model.Ridge\n- sklearn.linear_model.MultiTaskElasticNet\n- sklearn.kernel_ridge.KernelRidge\n- sklearn.ensemble.HistGradientBoostingRegressor\n\nWe added heavy noise to their predictions to make sure the NN can learn from other features as well\n```\n        self.blender = torch.nn.Sequential(\n            GaussianNoise(self.blend_noise),\n            torch.nn.Linear(out_dim * 4, 128),\n            torch.nn.LayerNorm(128),\n            activation(),\n            torch.nn.Dropout(self.blend_dropout),\n        )\n```\n\n## GroupK Cross Validation on Target Clusters\nThe tricks in this section increased both public and private LB, but we cannot compare the CV because it is a CV scheme change. Luckily it is (relatively, I guess?) performing well on both public and private.\n\nIt is known that there are some subtle domain shifts between train, private and public test sets. However, the difficulty is that the shift is happening in at least 3 directions (donor, day, cell types). To create a hard but not too hard CV scheme, we find that clustering the target values performed very well on both of the public and private leaderboard.\n\nLet's consider the CV scheme selection as a spectrum:\n- The easiest CV scheme: Random K fold (Downside: not representative of the test set)\n- The hard CV scheme: GroupKfold by day/donor (Downside: too few fold to train)\n- The hardest CV scheme 1: Time series split (Downside: wasting the last day data)\n- The hardest CV scheme 2: Excluding the 1 day or 1 donor completely from the training set (Downside: too hard/defensive)\n\nAnother reason of doing the clustering is that the day here is categorical, however in real life, time is continuous. GroupK CV by day is not that satisfying.\n\nThe first image shows the target kmeans result (colors)  visualized with the tsvd targets (points):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fb5e02c4c3d9edc2029a1458d02976f34%2F1.png?generation=1668564701067041&alt=media)\n\nNext, you can see the target clusters capture the cell type differences:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fa76444828dbe22e970e9af2a5909005a%2F2.png?generation=1668564719243229&alt=media)\n\nAnd the shifts of day and donor are not that significant compared to the cell types in the context of target clustering:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fd061e807ed3799640c34629a225bc5cc%2F3.png?generation=1668564731817794&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F53706c2ecd0ffc45cdf32d2517b71a8a%2F4.png?generation=1668564744081492&alt=media)\n\n## Regularization/Augmentation\nThe tricks in this section increased both CV and LB.\n\n### Seed-bagging\nI think most people have done this, we trained the same model a few more times with the different seeds for blending.\n\n### Mixup Augmentation and Stochastic Weight Averaging\nThe training is done on roughly 3 stages\n#### 1. Mix up augmentation stage\nSince all features are numerical values, mixup worked well.\n```\ndef mixup_augmentation(x: torch.Tensor, y: torch.Tensor, alpha: float = 5):\n    lam = np.random.beta(alpha, alpha)\n    rand_idx = torch.randperm(x.shape[0])\n    mixed_x = lam * x + (1 - lam) * x[rand_idx, :]\n    target_a, target_b = y, y[rand_idx]\n    return mixed_x, target_a, target_b, lam\n```\n#### 2. Normal training stage\n#### 3. SWA stage\nhttps://pytorch.org/docs/stable/optim.html#putting-it-all-together\nThis is similar to seed-bagging, I am not sure if they are overlapping or if they have their benefits here.",
      "votes": null
    },
    {
      "id": "2031343",
      "postDate": "11/16/2022 02:38:31",
      "content": "<p>great solution! congrats!</p>\n<blockquote>\n  <p>Seed-bagging</p>\n</blockquote>\n<p>Can I ask how much improvement you got with seed-bagging?</p>",
      "rawMarkdown": "great solution! congrats!\n\n>Seed-bagging\n\nCan I ask how much improvement you got with seed-bagging?",
      "votes": null
    },
    {
      "id": "2031350",
      "postDate": "11/16/2022 02:50:05",
      "content": "<p>For example, this is 4-seed groupk cv by donor result cv</p>\n<pre><code>cite_tsvd_50_torch_nn_oof_0.894400.npz 0.8943999611714102\ncite_tsvd_50_torch_nn_oof_0.894471.npz 0.8944711303264004\ncite_tsvd_50_torch_nn_oof_0.894500.npz 0.8945004071945007\ncite_tsvd_50_torch_nn_oof_0.894614.npz 0.8946143443923297\n0.8956681046650627\n</code></pre>\n<p>I haven't compare the lb of seed bagging for so long time, so cannot give you a number now,.</p>",
      "rawMarkdown": "For example, this is 4-seed groupk cv by donor result cv\n```\ncite_tsvd_50_torch_nn_oof_0.894400.npz 0.8943999611714102\ncite_tsvd_50_torch_nn_oof_0.894471.npz 0.8944711303264004\ncite_tsvd_50_torch_nn_oof_0.894500.npz 0.8945004071945007\ncite_tsvd_50_torch_nn_oof_0.894614.npz 0.8946143443923297\n0.8956681046650627\n```\n\nI haven't compare the lb of seed bagging for so long time, so cannot give you a number now,.",
      "votes": null
    },
    {
      "id": "2031363",
      "postDate": "11/16/2022 03:10:57",
      "content": "<p>Congrats! I would like to clarify the sentence \"We added heavy noise to their predictions to make sure the NN can learn from other features as well\".. Do you mean adding a big dropout value to the base model predictions?</p>",
      "rawMarkdown": "Congrats! I would like to clarify the sentence \"We added heavy noise to their predictions to make sure the NN can learn from other features as well\".. Do you mean adding a big dropout value to the base model predictions?",
      "votes": null
    },
    {
      "id": "2031369",
      "postDate": "11/16/2022 03:15:48",
      "content": "<p>We have both noise and dropout. </p>\n<p>We have added this layer copied from the intenet  with stddev ~ 0.8</p>\n<pre><code>class GaussianNoise(torch.nn.Module):\n    def __init__(self, stddev):\n        super().__init__()\n        self.stddev = stddev\n\n    def forward(self, din):\n        if self.training:\n            return din + torch.autograd.Variable(\n                torch.randn(din.size(), device=din.device) * self.stddev\n            )\n        return din\n</code></pre>\n<p>also 0.8 dropout as well</p>",
      "rawMarkdown": "We have both noise and dropout. \n\nWe have added this layer copied from the intenet  with stddev ~ 0.8\n```\nclass GaussianNoise(torch.nn.Module):\n    def __init__(self, stddev):\n        super().__init__()\n        self.stddev = stddev\n\n    def forward(self, din):\n        if self.training:\n            return din + torch.autograd.Variable(\n                torch.randn(din.size(), device=din.device) * self.stddev\n            )\n        return din\n```\n\nalso 0.8 dropout as well",
      "votes": null
    },
    {
      "id": "2031382",
      "postDate": "11/16/2022 03:35:25",
      "content": "<p>Thanks for the clarification nlgn… Good work.</p>",
      "rawMarkdown": "Thanks for the clarification nlgn... Good work.",
      "votes": null
    },
    {
      "id": "2031428",
      "postDate": "11/16/2022 04:40:33",
      "content": "<p>Congrats on results and condolences to gold <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> and <a href=\"https://www.kaggle.com/paragkale\" target=\"_blank\">@paragkale</a> </p>",
      "rawMarkdown": "Congrats on results and condolences to gold @kingychiu and @paragkale",
      "votes": null
    },
    {
      "id": "2031443",
      "postDate": "11/16/2022 04:52:30",
      "content": "<p>so for the cv score, 4 seed bagging got <code>0.001</code> improvement over 1 seed? Impressive!</p>",
      "rawMarkdown": "so for the cv score, 4 seed bagging got `0.001` improvement over 1 seed? Impressive!",
      "votes": null
    },
    {
      "id": "2031675",
      "postDate": "11/16/2022 07:58:28",
      "content": "<p>Here is the link to the multiome portion of our solution <a href=\"https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\" target=\"_blank\">https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion</a></p>\n<p>Very simple script, uses 12 folds for donor x day.</p>",
      "rawMarkdown": "Here is the link to the multiome portion of our solution https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\n\nVery simple script, uses 12 folds for donor x day.",
      "votes": null
    },
    {
      "id": "2043289",
      "postDate": "11/25/2022 14:24:53",
      "content": "<p>Congrats! The idea of GroupK Cross Validation on the Target Clusters is so novel and interesting for me! I learned a lot from it.</p>",
      "rawMarkdown": "Congrats! The idea of GroupK Cross Validation on the Target Clusters is so novel and interesting for me! I learned a lot from it.",
      "votes": null
    },
    {
      "id": "2044402",
      "postDate": "11/26/2022 14:23:07",
      "content": "<p>Thanks for you kind words.</p>\n<p>As always, I can't conclude it is really helpful based on only the final result; I think we need more experiments for most tricks reported in this comp to \"conclude\" what tricks are actually helpful to both unseen donors and unseen days.</p>\n<p>If you look at our sub sorted by private score in the below image, groupk by donor got the gold range private scores, but the public score is very low… target clustering seems to be ok-ish…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F86c6579d569842d93719098a799d146b%2FScreenshot%202022-11-26%20at%2010.56.25%20PM.png?generation=1669478203034830&amp;alt=media\" alt=\"\"></p>\n<p>If I could redo the entire competition again, I think I would do a simple K-fold, but validate/early-stop on the last day of the validation fold data. This trick has been used on kaggle many times to allow us to train on full data but still fit towards the latest data in time.</p>\n<ul>\n<li>You can see <a href=\"https://www.kaggle.com/AmbrosM\" target=\"_blank\">@AmbrosM</a> mentioned here as well <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395#2031471\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395#2031471</a> </li>\n<li>My old team has used this trick before as well (3 years ago): <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119395\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119395</a></li>\n</ul>\n<p>I think I was overthinking about the \"domain shift\" here. There are always some domain shifts in kaggle data, if you compare the shakeup this time to other historical kaggle competitions, this time is not huge… And looking at the gold solutions, nothing crazy/fancy domain adaptation skills have been performed. Mostly is about careful feature generation/selection if I hasn't missed anything.</p>",
      "rawMarkdown": "Thanks for you kind words.\n\nAs always, I can't conclude it is really helpful based on only the final result; I think we need more experiments for most tricks reported in this comp to \"conclude\" what tricks are actually helpful to both unseen donors and unseen days.\n\nIf you look at our sub sorted by private score in the below image, groupk by donor got the gold range private scores, but the public score is very low... target clustering seems to be ok-ish...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F86c6579d569842d93719098a799d146b%2FScreenshot%202022-11-26%20at%2010.56.25%20PM.png?generation=1669478203034830&alt=media)\n\nIf I could redo the entire competition again, I think I would do a simple K-fold, but validate/early-stop on the last day of the validation fold data. This trick has been used on kaggle many times to allow us to train on full data but still fit towards the latest data in time.\n\n- You can see @AmbrosM mentioned here as well https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395#2031471 \n- My old team has used this trick before as well (3 years ago): https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119395\n\nI think I was overthinking about the \"domain shift\" here. There are always some domain shifts in kaggle data, if you compare the shakeup this time to other historical kaggle competitions, this time is not huge... And looking at the gold solutions, nothing crazy/fancy domain adaptation skills have been performed. Mostly is about careful feature generation/selection if I hasn't missed anything.",
      "votes": null
    },
    {
      "id": "2044502",
      "postDate": "11/26/2022 15:15:53",
      "content": "<p>To be clear, I am not saying we should ignore shifts in data.</p>\n<p>In the context of competition, the highest chance is that all participants cannot make a significantly closer public-private leaderboard score gap, which means the gap is more like a hidden difference between datasets, which is really hard to solve in a 3-month competition (or even years of research). </p>\n<p>Btw, I learned the concept and the difficulties of dataset shift in this paper: <a href=\"https://arxiv.org/abs/2007.00644\" target=\"_blank\">https://arxiv.org/abs/2007.00644</a></p>\n<blockquote>\n  <p>Most research on robustness focuses on synthetic image perturbations (noise, simulated weather artifacts, adversarial examples, etc.), which leaves open how robustness on synthetic distribution shift relates to distribution shift arising in real data. …. most current techniques provide no robustness to the natural distribution shifts in our testbed. The main exception is training on larger and more diverse datasets</p>\n</blockquote>\n<p>For research, of course, we want to make the public-private leaderboard gap as close as possible.  So, my reflection on this is we should have the leaderboard ranking based on the public-private leaderboard score gap in this type of competition. Then the competition ranking aligns with the host's objective of studying domain shift and domain adaptation. I think this is a feature request to <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> , because it seems like people now care more about robustness then performance.</p>\n<p>Lastly, out of curiosity, I want to ask for the host's comment <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> on the current finalized public-private leaderboard gap (0.81x ~ 0.77x). Is this gap  \"good enough\", \"can be improved\" or \"totally unacceptable\"? Without domain knowledge, I think the gap is not crazily huge, is it expected?</p>",
      "rawMarkdown": "To be clear, I am not saying we should ignore shifts in data.\n\nIn the context of competition, the highest chance is that all participants cannot make a significantly closer public-private leaderboard score gap, which means the gap is more like a hidden difference between datasets, which is really hard to solve in a 3-month competition (or even years of research). \n\nBtw, I learned the concept and the difficulties of dataset shift in this paper: https://arxiv.org/abs/2007.00644\n\n> Most research on robustness focuses on synthetic image perturbations (noise, simulated weather artifacts, adversarial examples, etc.), which leaves open how robustness on synthetic distribution shift relates to distribution shift arising in real data. .... most current techniques provide no robustness to the natural distribution shifts in our testbed. The main exception is training on larger and more diverse datasets\n\nFor research, of course, we want to make the public-private leaderboard gap as close as possible.  So, my reflection on this is we should have the leaderboard ranking based on the public-private leaderboard score gap in this type of competition. Then the competition ranking aligns with the host's objective of studying domain shift and domain adaptation. I think this is a feature request to @ryanholbrook , because it seems like people now care more about robustness then performance.\n\nLastly, out of curiosity, I want to ask for the host's comment @danielburkhardt on the current finalized public-private leaderboard gap (0.81x ~ 0.77x). Is this gap  \"good enough\", \"can be improved\" or \"totally unacceptable\"? Without domain knowledge, I think the gap is not crazily huge, is it expected?",
      "votes": null
    },
    {
      "id": "2048756",
      "postDate": "11/29/2022 16:49:22",
      "content": "<p>I tried to rank the teams by 2 different \"robustness measures\" in this notebook<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/robustness-on-open-problems-multimodal?scriptVersionId=112466254\" target=\"_blank\">https://www.kaggle.com/code/kingychiu/robustness-on-open-problems-multimodal?scriptVersionId=112466254</a></p>\n<p>haha, I realized it is tricky to use robustness for a leaderboard because, generally, underfit models / poorly performed teams got quite robust scores… </p>",
      "rawMarkdown": "I tried to rank the teams by 2 different \"robustness measures\" in this notebook\nhttps://www.kaggle.com/code/kingychiu/robustness-on-open-problems-multimodal?scriptVersionId=112466254\n\nhaha, I realized it is tricky to use robustness for a leaderboard because, generally, underfit models / poorly performed teams got quite robust scores...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031343,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/16/2022 02:38:31",
      "content": "<p>great solution! congrats!</p>\n<blockquote>\n  <p>Seed-bagging</p>\n</blockquote>\n<p>Can I ask how much improvement you got with seed-bagging?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031350,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "11/16/2022 02:50:05",
          "content": "<p>For example, this is 4-seed groupk cv by donor result cv</p>\n<pre><code>cite_tsvd_50_torch_nn_oof_0.894400.npz 0.8943999611714102\ncite_tsvd_50_torch_nn_oof_0.894471.npz 0.8944711303264004\ncite_tsvd_50_torch_nn_oof_0.894500.npz 0.8945004071945007\ncite_tsvd_50_torch_nn_oof_0.894614.npz 0.8946143443923297\n0.8956681046650627\n</code></pre>\n<p>I haven't compare the lb of seed bagging for so long time, so cannot give you a number now,.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2031443,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "11/16/2022 04:52:30",
          "content": "<p>so for the cv score, 4 seed bagging got <code>0.001</code> improvement over 1 seed? Impressive!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2031363,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "11/16/2022 03:10:57",
      "content": "<p>Congrats! I would like to clarify the sentence \"We added heavy noise to their predictions to make sure the NN can learn from other features as well\".. Do you mean adding a big dropout value to the base model predictions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031369,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "11/16/2022 03:15:48",
          "content": "<p>We have both noise and dropout. </p>\n<p>We have added this layer copied from the intenet  with stddev ~ 0.8</p>\n<pre><code>class GaussianNoise(torch.nn.Module):\n    def __init__(self, stddev):\n        super().__init__()\n        self.stddev = stddev\n\n    def forward(self, din):\n        if self.training:\n            return din + torch.autograd.Variable(\n                torch.randn(din.size(), device=din.device) * self.stddev\n            )\n        return din\n</code></pre>\n<p>also 0.8 dropout as well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2031382,
          "author_name": "drpatrickchan",
          "author_url": "",
          "post_date": "11/16/2022 03:35:25",
          "content": "<p>Thanks for the clarification nlgn… Good work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2031428,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "11/16/2022 04:40:33",
      "content": "<p>Congrats on results and condolences to gold <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> and <a href=\"https://www.kaggle.com/paragkale\" target=\"_blank\">@paragkale</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2031675,
      "author_name": "paragkale",
      "author_url": "",
      "post_date": "11/16/2022 07:58:28",
      "content": "<p>Here is the link to the multiome portion of our solution <a href=\"https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\" target=\"_blank\">https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion</a></p>\n<p>Very simple script, uses 12 folds for donor x day.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2043289,
      "author_name": "kmkm1234",
      "author_url": "",
      "post_date": "11/25/2022 14:24:53",
      "content": "<p>Congrats! The idea of GroupK Cross Validation on the Target Clusters is so novel and interesting for me! I learned a lot from it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2044402,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "11/26/2022 14:23:07",
          "content": "<p>Thanks for you kind words.</p>\n<p>As always, I can't conclude it is really helpful based on only the final result; I think we need more experiments for most tricks reported in this comp to \"conclude\" what tricks are actually helpful to both unseen donors and unseen days.</p>\n<p>If you look at our sub sorted by private score in the below image, groupk by donor got the gold range private scores, but the public score is very low… target clustering seems to be ok-ish…</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F86c6579d569842d93719098a799d146b%2FScreenshot%202022-11-26%20at%2010.56.25%20PM.png?generation=1669478203034830&amp;alt=media\" alt=\"\"></p>\n<p>If I could redo the entire competition again, I think I would do a simple K-fold, but validate/early-stop on the last day of the validation fold data. This trick has been used on kaggle many times to allow us to train on full data but still fit towards the latest data in time.</p>\n<ul>\n<li>You can see <a href=\"https://www.kaggle.com/AmbrosM\" target=\"_blank\">@AmbrosM</a> mentioned here as well <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395#2031471\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395#2031471</a> </li>\n<li>My old team has used this trick before as well (3 years ago): <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119395\" target=\"_blank\">https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119395</a></li>\n</ul>\n<p>I think I was overthinking about the \"domain shift\" here. There are always some domain shifts in kaggle data, if you compare the shakeup this time to other historical kaggle competitions, this time is not huge… And looking at the gold solutions, nothing crazy/fancy domain adaptation skills have been performed. Mostly is about careful feature generation/selection if I hasn't missed anything.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2044502,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "11/26/2022 15:15:53",
          "content": "<p>To be clear, I am not saying we should ignore shifts in data.</p>\n<p>In the context of competition, the highest chance is that all participants cannot make a significantly closer public-private leaderboard score gap, which means the gap is more like a hidden difference between datasets, which is really hard to solve in a 3-month competition (or even years of research). </p>\n<p>Btw, I learned the concept and the difficulties of dataset shift in this paper: <a href=\"https://arxiv.org/abs/2007.00644\" target=\"_blank\">https://arxiv.org/abs/2007.00644</a></p>\n<blockquote>\n  <p>Most research on robustness focuses on synthetic image perturbations (noise, simulated weather artifacts, adversarial examples, etc.), which leaves open how robustness on synthetic distribution shift relates to distribution shift arising in real data. …. most current techniques provide no robustness to the natural distribution shifts in our testbed. The main exception is training on larger and more diverse datasets</p>\n</blockquote>\n<p>For research, of course, we want to make the public-private leaderboard gap as close as possible.  So, my reflection on this is we should have the leaderboard ranking based on the public-private leaderboard score gap in this type of competition. Then the competition ranking aligns with the host's objective of studying domain shift and domain adaptation. I think this is a feature request to <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> , because it seems like people now care more about robustness then performance.</p>\n<p>Lastly, out of curiosity, I want to ask for the host's comment <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a> on the current finalized public-private leaderboard gap (0.81x ~ 0.77x). Is this gap  \"good enough\", \"can be improved\" or \"totally unacceptable\"? Without domain knowledge, I think the gap is not crazily huge, is it expected?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2048756,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "11/29/2022 16:49:22",
          "content": "<p>I tried to rank the teams by 2 different \"robustness measures\" in this notebook<br>\n<a href=\"https://www.kaggle.com/code/kingychiu/robustness-on-open-problems-multimodal?scriptVersionId=112466254\" target=\"_blank\">https://www.kaggle.com/code/kingychiu/robustness-on-open-problems-multimodal?scriptVersionId=112466254</a></p>\n<p>haha, I realized it is tricky to use robustness for a leaderboard because, generally, underfit models / poorly performed teams got quite robust scores… </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2031325": "## Intro\n- I have been mainly working on the Cite part. I tried many things, \n- multi part by @paragkale https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\n- The following tricks gave me the most gain.\n- I will update in this post about the code and what didn't work.\n\n## Extra Data\n- The raw count data released by the host.\n\n##Dimensionality Reduction\n\nI think the most helpful one are:\n- sklearn.decomposition.TruncatedSVD (128 comps)\n- Self-made denosing auto encoder (128 hidden nodes)\n\n## Direct Features\n- Direct features based on matching the names\n- Direct features based on absolute correlation to targets\n- Direct features based on the list shared by the hosts in this thread https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366392\n\n## Use base models outcomes as NN inputs feature for ensembling\n\nIt is known that MSE is not a really good loss function for the competition metric. Therefore within each fold, we trained 4 base models and used their features for the NN input.\n- sklearn.linear_model.Ridge\n- sklearn.linear_model.MultiTaskElasticNet\n- sklearn.kernel_ridge.KernelRidge\n- sklearn.ensemble.HistGradientBoostingRegressor\n\nWe added heavy noise to their predictions to make sure the NN can learn from other features as well\n```\n        self.blender = torch.nn.Sequential(\n            GaussianNoise(self.blend_noise),\n            torch.nn.Linear(out_dim * 4, 128),\n            torch.nn.LayerNorm(128),\n            activation(),\n            torch.nn.Dropout(self.blend_dropout),\n        )\n```\n\n## GroupK Cross Validation on Target Clusters\nThe tricks in this section increased both public and private LB, but we cannot compare the CV because it is a CV scheme change. Luckily it is (relatively, I guess?) performing well on both public and private.\n\nIt is known that there are some subtle domain shifts between train, private and public test sets. However, the difficulty is that the shift is happening in at least 3 directions (donor, day, cell types). To create a hard but not too hard CV scheme, we find that clustering the target values performed very well on both of the public and private leaderboard.\n\nLet's consider the CV scheme selection as a spectrum:\n- The easiest CV scheme: Random K fold (Downside: not representative of the test set)\n- The hard CV scheme: GroupKfold by day/donor (Downside: too few fold to train)\n- The hardest CV scheme 1: Time series split (Downside: wasting the last day data)\n- The hardest CV scheme 2: Excluding the 1 day or 1 donor completely from the training set (Downside: too hard/defensive)\n\nAnother reason of doing the clustering is that the day here is categorical, however in real life, time is continuous. GroupK CV by day is not that satisfying.\n\nThe first image shows the target kmeans result (colors)  visualized with the tsvd targets (points):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fb5e02c4c3d9edc2029a1458d02976f34%2F1.png?generation=1668564701067041&alt=media)\n\nNext, you can see the target clusters capture the cell type differences:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fa76444828dbe22e970e9af2a5909005a%2F2.png?generation=1668564719243229&alt=media)\n\nAnd the shifts of day and donor are not that significant compared to the cell types in the context of target clustering:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fd061e807ed3799640c34629a225bc5cc%2F3.png?generation=1668564731817794&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F53706c2ecd0ffc45cdf32d2517b71a8a%2F4.png?generation=1668564744081492&alt=media)\n\n## Regularization/Augmentation\nThe tricks in this section increased both CV and LB.\n\n### Seed-bagging\nI think most people have done this, we trained the same model a few more times with the different seeds for blending.\n\n### Mixup Augmentation and Stochastic Weight Averaging\nThe training is done on roughly 3 stages\n#### 1. Mix up augmentation stage\nSince all features are numerical values, mixup worked well.\n```\ndef mixup_augmentation(x: torch.Tensor, y: torch.Tensor, alpha: float = 5):\n    lam = np.random.beta(alpha, alpha)\n    rand_idx = torch.randperm(x.shape[0])\n    mixed_x = lam * x + (1 - lam) * x[rand_idx, :]\n    target_a, target_b = y, y[rand_idx]\n    return mixed_x, target_a, target_b, lam\n```\n#### 2. Normal training stage\n#### 3. SWA stage\nhttps://pytorch.org/docs/stable/optim.html#putting-it-all-together\nThis is similar to seed-bagging, I am not sure if they are overlapping or if they have their benefits here.",
    "2031343": "great solution! congrats!\n\n>Seed-bagging\n\nCan I ask how much improvement you got with seed-bagging?",
    "2031350": "For example, this is 4-seed groupk cv by donor result cv\n```\ncite_tsvd_50_torch_nn_oof_0.894400.npz 0.8943999611714102\ncite_tsvd_50_torch_nn_oof_0.894471.npz 0.8944711303264004\ncite_tsvd_50_torch_nn_oof_0.894500.npz 0.8945004071945007\ncite_tsvd_50_torch_nn_oof_0.894614.npz 0.8946143443923297\n0.8956681046650627\n```\n\nI haven't compare the lb of seed bagging for so long time, so cannot give you a number now,.",
    "2031363": "Congrats! I would like to clarify the sentence \"We added heavy noise to their predictions to make sure the NN can learn from other features as well\".. Do you mean adding a big dropout value to the base model predictions?",
    "2031369": "We have both noise and dropout. \n\nWe have added this layer copied from the intenet  with stddev ~ 0.8\n```\nclass GaussianNoise(torch.nn.Module):\n    def __init__(self, stddev):\n        super().__init__()\n        self.stddev = stddev\n\n    def forward(self, din):\n        if self.training:\n            return din + torch.autograd.Variable(\n                torch.randn(din.size(), device=din.device) * self.stddev\n            )\n        return din\n```\n\nalso 0.8 dropout as well",
    "2031382": "Thanks for the clarification nlgn... Good work.",
    "2031428": "Congrats on results and condolences to gold @kingychiu and @paragkale",
    "2031443": "so for the cv score, 4 seed bagging got `0.001` improvement over 1 seed? Impressive!",
    "2031675": "Here is the link to the multiome portion of our solution https://www.kaggle.com/paragkale/private-14th-public-6th-multiome-portion\n\nVery simple script, uses 12 folds for donor x day.",
    "2043289": "Congrats! The idea of GroupK Cross Validation on the Target Clusters is so novel and interesting for me! I learned a lot from it.",
    "2044402": "Thanks for you kind words.\n\nAs always, I can't conclude it is really helpful based on only the final result; I think we need more experiments for most tricks reported in this comp to \"conclude\" what tricks are actually helpful to both unseen donors and unseen days.\n\nIf you look at our sub sorted by private score in the below image, groupk by donor got the gold range private scores, but the public score is very low... target clustering seems to be ok-ish...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2F86c6579d569842d93719098a799d146b%2FScreenshot%202022-11-26%20at%2010.56.25%20PM.png?generation=1669478203034830&alt=media)\n\nIf I could redo the entire competition again, I think I would do a simple K-fold, but validate/early-stop on the last day of the validation fold data. This trick has been used on kaggle many times to allow us to train on full data but still fit towards the latest data in time.\n\n- You can see @AmbrosM mentioned here as well https://www.kaggle.com/competitions/open-problems-multimodal/discussion/366395#2031471 \n- My old team has used this trick before as well (3 years ago): https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119395\n\nI think I was overthinking about the \"domain shift\" here. There are always some domain shifts in kaggle data, if you compare the shakeup this time to other historical kaggle competitions, this time is not huge... And looking at the gold solutions, nothing crazy/fancy domain adaptation skills have been performed. Mostly is about careful feature generation/selection if I hasn't missed anything.",
    "2044502": "To be clear, I am not saying we should ignore shifts in data.\n\nIn the context of competition, the highest chance is that all participants cannot make a significantly closer public-private leaderboard score gap, which means the gap is more like a hidden difference between datasets, which is really hard to solve in a 3-month competition (or even years of research). \n\nBtw, I learned the concept and the difficulties of dataset shift in this paper: https://arxiv.org/abs/2007.00644\n\n> Most research on robustness focuses on synthetic image perturbations (noise, simulated weather artifacts, adversarial examples, etc.), which leaves open how robustness on synthetic distribution shift relates to distribution shift arising in real data. .... most current techniques provide no robustness to the natural distribution shifts in our testbed. The main exception is training on larger and more diverse datasets\n\nFor research, of course, we want to make the public-private leaderboard gap as close as possible.  So, my reflection on this is we should have the leaderboard ranking based on the public-private leaderboard score gap in this type of competition. Then the competition ranking aligns with the host's objective of studying domain shift and domain adaptation. I think this is a feature request to @ryanholbrook , because it seems like people now care more about robustness then performance.\n\nLastly, out of curiosity, I want to ask for the host's comment @danielburkhardt on the current finalized public-private leaderboard gap (0.81x ~ 0.77x). Is this gap  \"good enough\", \"can be improved\" or \"totally unacceptable\"? Without domain knowledge, I think the gap is not crazily huge, is it expected?",
    "2048756": "I tried to rank the teams by 2 different \"robustness measures\" in this notebook\nhttps://www.kaggle.com/code/kingychiu/robustness-on-open-problems-multimodal?scriptVersionId=112466254\n\nhaha, I realized it is tricky to use robustness for a leaderboard because, generally, underfit models / poorly performed teams got quite robust scores..."
  },
  "source": "meta"
}