{
  "id": 498983,
  "title": "CV/LB thread",
  "url": "/competitions/leash-BELKA/discussion/498983",
  "author_name": "",
  "post_date": "2024-04-30T07:56:34.489856Z",
  "votes": 28,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Please share your CV/LB results here. </p>\n<p>I'm using my <a href=\"https://www.kaggle.com/datasets/thedrcat/belka-cv-split\" target=\"_blank\">CV split</a>, here are most recent results:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV-random</th>\n<th>CV-noshare</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Molformer</td>\n<td>0.64</td>\n<td>0.008</td>\n<td>0.564</td>\n</tr>\n<tr>\n<td>Molformer</td>\n<td>0.62</td>\n<td>0.017</td>\n<td>0.554</td>\n</tr>\n</tbody>\n</table>\n<p>From other discussions and my initial results it's likely that public LB seems correlated with random split, but private results will most likely depend on how well you can model the no-share / test distribution. </p>\n<p>Below more details by protein target (if you're not using <a href=\"https://wandb.ai\" target=\"_blank\">W&amp;B</a> for tracking, you're missing out):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1328185%2Fa57a0fd2ce86c30dd3b0345958eb2c4f%2FScreenshot%202024-04-30%20at%2009.46.49.png?generation=1714463680621641&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2784340",
      "postDate": "04/30/2024 07:56:34",
      "content": "<p>Please share your CV/LB results here. </p>\n<p>I'm using my <a href=\"https://www.kaggle.com/datasets/thedrcat/belka-cv-split\" target=\"_blank\">CV split</a>, here are most recent results:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV-random</th>\n<th>CV-noshare</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Molformer</td>\n<td>0.64</td>\n<td>0.008</td>\n<td>0.564</td>\n</tr>\n<tr>\n<td>Molformer</td>\n<td>0.62</td>\n<td>0.017</td>\n<td>0.554</td>\n</tr>\n</tbody>\n</table>\n<p>From other discussions and my initial results it's likely that public LB seems correlated with random split, but private results will most likely depend on how well you can model the no-share / test distribution. </p>\n<p>Below more details by protein target (if you're not using <a href=\"https://wandb.ai\" target=\"_blank\">W&amp;B</a> for tracking, you're missing out):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1328185%2Fa57a0fd2ce86c30dd3b0345958eb2c4f%2FScreenshot%202024-04-30%20at%2009.46.49.png?generation=1714463680621641&amp;alt=media\"></p>",
      "rawMarkdown": "Please share your CV/LB results here. \n\nI'm using my [CV split](https://www.kaggle.com/datasets/thedrcat/belka-cv-split), here are most recent results:\n\n| Model | CV-random | CV-noshare | LB |\n| --- | --- | --- | --- |\n| Molformer | 0.64 | 0.008 | 0.564 |\n| Molformer | 0.62 | 0.017 | 0.554 |\n\n\nFrom other discussions and my initial results it's likely that public LB seems correlated with random split, but private results will most likely depend on how well you can model the no-share / test distribution. \n\nBelow more details by protein target (if you're not using [W&B](https://wandb.ai) for tracking, you're missing out):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1328185%2Fa57a0fd2ce86c30dd3b0345958eb2c4f%2FScreenshot%202024-04-30%20at%2009.46.49.png?generation=1714463680621641&alt=media)",
      "votes": null
    },
    {
      "id": "2785587",
      "postDate": "04/30/2024 22:56:03",
      "content": "<p>My CV LB split is very split dependent. The more samples I add to my experiments the lower my CV (using scaffold splitting) but higher LB. I'd put some examples here, but I am still experimenting.</p>",
      "rawMarkdown": "My CV LB split is very split dependent. The more samples I add to my experiments the lower my CV (using scaffold splitting) but higher LB. I'd put some examples here, but I am still experimenting.",
      "votes": null
    },
    {
      "id": "2785685",
      "postDate": "05/01/2024 01:45:03",
      "content": "<p>the best split is probably random, see also paper.<br>\ni suggest 2 separate splits and experiments</p>\n<p>my strategy:<br>\n[1] cv score versus num of share blocks</p>\n<ul>\n<li>train samples = share blocks</li>\n<li>valid samples with 0,1,2,3 share blocks<br>\nmake a plot of cv score versus num of share blocks<br>\nmake tsnet map or knn results for the blocks<br>\ntry some external dataset<br>\nmy conclusion here so far is that for non share blocks, results are really bad (need to solve later)</li>\n</ul>\n<p>[2] cv score for share blocks</p>\n<ul>\n<li>best results is probably using random split and all samples</li>\n</ul>\n<p>paper on cv split, datasize, etc:<br>\nUncovering Neural Scaling Laws in Molecular Representation Learning<br>\n<a href=\"https://arxiv.org/abs/2309.15123\" target=\"_blank\">https://arxiv.org/abs/2309.15123</a></p>",
      "rawMarkdown": "the best split is probably random, see also paper.\ni suggest 2 separate splits and experiments\n\nmy strategy:\n[1] cv score versus num of share blocks\n- train samples = share blocks\n- valid samples with 0,1,2,3 share blocks\nmake a plot of cv score versus num of share blocks\nmake tsnet map or knn results for the blocks\ntry some external dataset\nmy conclusion here so far is that for non share blocks, results are really bad (need to solve later)\n\n[2] cv score for share blocks\n- best results is probably using random split and all samples\n\n\n\npaper on cv split, datasize, etc:\nUncovering Neural Scaling Laws in Molecular Representation Learning\nhttps://arxiv.org/abs/2309.15123",
      "votes": null
    },
    {
      "id": "2785690",
      "postDate": "05/01/2024 01:50:00",
      "content": "<p>\"I'm using my CV split, here are most recent results:\"</p>\n<p>both molformer and chemberta should be about to give you lb 0.575. (CV 0.670), i am using random splits<br>\nI think this is the limit of SMILES string transformer?<br>\n(i think some kaggler can get up to lb0.605(CV 0.700))?</p>\n<p>molformer = 12 layers<br>\nchemberta = 3 layers</p>\n<p>different pretrain chemberta 5,10,77M MLM,MTR has slight different results</p>\n<p>Note there a 2 different molformer in papers (one is based on SMILES string , another is a GNN)</p>",
      "rawMarkdown": "\"I'm using my CV split, here are most recent results:\"\n\nboth molformer and chemberta should be about to give you lb 0.575. (CV 0.670), i am using random splits\nI think this is the limit of SMILES string transformer?\n(i think some kaggler can get up to lb0.605(CV 0.700))?\n\nmolformer = 12 layers\nchemberta = 3 layers\n\ndifferent pretrain chemberta 5,10,77M MLM,MTR has slight different results\n\nNote there a 2 different molformer in papers (one is based on SMILES string , another is a GNN)",
      "votes": null
    },
    {
      "id": "2786617",
      "postDate": "05/01/2024 11:27:41",
      "content": "<p>I agree there will likely need to be ways to handle compounds close to the train set and those that aren't. I have 64gb of RAM at home and it seems that even with np.packbits I can't train models with half of the train rows. So, now I'm focused on finding a subsampling strategy. I'm learning a lot about how to manage large datasets in this competition.</p>",
      "rawMarkdown": "I agree there will likely need to be ways to handle compounds close to the train set and those that aren't. I have 64gb of RAM at home and it seems that even with np.packbits I can't train models with half of the train rows. So, now I'm focused on finding a subsampling strategy. I'm learning a lot about how to manage large datasets in this competition.",
      "votes": null
    },
    {
      "id": "2786912",
      "postDate": "05/01/2024 13:47:12",
      "content": "<p>actually i am think that for those kagglers that shown good performance, espeically those with good domain knowledge, google should get them better compute power, e.g. discount price at google cloud for better gpu, etc</p>\n<p>or kaggler ourselves can setup some fund.<br>\nif the solution is good, then we can all learn something valuable, rather than trial and error like a headless fly</p>",
      "rawMarkdown": "actually i am think that for those kagglers that shown good performance, espeically those with good domain knowledge, google should get them better compute power, e.g. discount price at google cloud for better gpu, etc\n\nor kaggler ourselves can setup some fund.\nif the solution is good, then we can all learn something valuable, rather than trial and error like a headless fly",
      "votes": null
    },
    {
      "id": "2787584",
      "postDate": "05/01/2024 20:36:32",
      "content": "<p>CV 0.69, LB 0.585<br>\nmodel : 1dcnn trained on all data<br>\nI made this model public : <a href=\"https://www.kaggle.com/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">https://www.kaggle.com/ahmedelfazouan/belka-1dcnn-starter-with-all-data</a></p>",
      "rawMarkdown": "CV 0.69, LB 0.585\nmodel : 1dcnn trained on all data\nI made this model public : https://www.kaggle.com/ahmedelfazouan/belka-1dcnn-starter-with-all-data",
      "votes": null
    },
    {
      "id": "2787594",
      "postDate": "05/01/2024 20:46:46",
      "content": "<p>This is pretty cool <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> ! Looking forward to reading through it.</p>",
      "rawMarkdown": "This is pretty cool @ahmedelfazouan ! Looking forward to reading through it.",
      "votes": null
    },
    {
      "id": "2787630",
      "postDate": "05/01/2024 21:22:27",
      "content": "<p>I've considered renting more compute power, but I do think exploring a subsampling strategy is a good learning experience for me. </p>",
      "rawMarkdown": "I've considered renting more compute power, but I do think exploring a subsampling strategy is a good learning experience for me.",
      "votes": null
    },
    {
      "id": "2788308",
      "postDate": "05/02/2024 07:17:13",
      "content": "<p>CV: 0.597, 0.600, 0.608, 0.613<br>\nLB: 0.575, 0.586, 0.589, 0.585</p>\n<p>Model: small nn on top of frozen molformer<br>\nVal Split: positive ratio 1% (random, len 1e5)</p>",
      "rawMarkdown": "CV: 0.597, 0.600, 0.608, 0.613\nLB: 0.575, 0.586, 0.589, 0.585\n\nModel: small nn on top of frozen molformer\nVal Split: positive ratio 1% (random, len 1e5)",
      "votes": null
    },
    {
      "id": "2788485",
      "postDate": "05/02/2024 08:51:47",
      "content": "<p>your cv looks low.<br>\ni am not sure if it is the effect of different split or frozen (i think frozen is more likely)</p>\n<p>the good news is that CV and LB gap is very small, which means that you actually score very well in nonshare b block (you can confirm by local experiment or separate submission)</p>\n<p>and i think if you starts unfreeze, you will start to have different results.<br>\nyou can try ensemble of freeze and frozen models, or using freeze and frozen models for share and unshare blocks</p>",
      "rawMarkdown": "your cv looks low.\ni am not sure if it is the effect of different split or frozen (i think frozen is more likely)\n\nthe good news is that CV and LB gap is very small, which means that you actually score very well in nonshare b block (you can confirm by local experiment or separate submission)\n\nand i think if you starts unfreeze, you will start to have different results.\nyou can try ensemble of freeze and frozen models, or using freeze and frozen models for share and unshare blocks",
      "votes": null
    },
    {
      "id": "2788579",
      "postDate": "05/02/2024 09:44:57",
      "content": "<p>i suddenly have an idea.<br>\nhow about LORA like model?<br>\nuse some layers to modify molformer intermediate or/and final layer output?</p>",
      "rawMarkdown": "i suddenly have an idea.\nhow about LORA like model?\nuse some layers to modify molformer intermediate or/and final layer output?",
      "votes": null
    },
    {
      "id": "2788682",
      "postDate": "05/02/2024 10:54:42",
      "content": "<p>I think there are still some lower hanging fruits that can be picked. These models were trained (with 4090*1) in around 5-7 hours with a subset of (neg) train, only iterating over the train set about twice.</p>\n<p>Before I set the fixed positive ratio val, a different random split/sample had cv around 0.62, so yeah this split is lower.</p>\n<p>I don't think I'll unfreeze molformer just yet as that adds a massive amount of compute cost and will destroy my batch size.</p>",
      "rawMarkdown": "I think there are still some lower hanging fruits that can be picked. These models were trained (with 4090*1) in around 5-7 hours with a subset of (neg) train, only iterating over the train set about twice.\n\nBefore I set the fixed positive ratio val, a different random split/sample had cv around 0.62, so yeah this split is lower.\n\nI don't think I'll unfreeze molformer just yet as that adds a massive amount of compute cost and will destroy my batch size.",
      "votes": null
    },
    {
      "id": "2788934",
      "postDate": "05/02/2024 13:37:36",
      "content": "<p>Hmmm will have to look into loras, definitely seems like a possible path forward.</p>",
      "rawMarkdown": "Hmmm will have to look into loras, definitely seems like a possible path forward.",
      "votes": null
    },
    {
      "id": "2790293",
      "postDate": "05/03/2024 05:40:06",
      "content": "<p>i load my data,  and sample 1024 * 1024 molecules, with X% positive samples, and save it to a parquet file. I do this 100 times. For training i load one file at a time to save ram. You could also have a separate file for just postives , and multiple files for negative if you want to change +ve/-ve split on the fly.</p>",
      "rawMarkdown": "i load my data,  and sample 1024 * 1024 molecules, with X% positive samples, and save it to a parquet file. I do this 100 times. For training i load one file at a time to save ram. You could also have a separate file for just postives , and multiple files for negative if you want to change +ve/-ve split on the fly.",
      "votes": null
    },
    {
      "id": "2790563",
      "postDate": "05/03/2024 08:31:35",
      "content": "<p>molformer works slightly better for me.<br>\nWith Molformer i get .69 CV, .583 lb.<br>\nWith chemberta i get .65 CV, .56 lb. </p>\n<p>i tried using Selformer, ChemGPT (selfies models) but got much worse scores.</p>",
      "rawMarkdown": "molformer works slightly better for me.\nWith Molformer i get .69 CV, .583 lb.\nWith chemberta i get .65 CV, .56 lb. \n\ni tried using Selformer, ChemGPT (selfies models) but got much worse scores.",
      "votes": null
    },
    {
      "id": "2790810",
      "postDate": "05/03/2024 10:51:02",
      "content": "<p>Using a variation on <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> notebook, I get the following CV-LB scores.</p>\n<p>0.72 - 0.61<br>\n0.71 - 0.6<br>\n0.65 - 0.57</p>\n<p>I've only tried a few simple things here, but I am really surprised how well a 1D-CNN handles the in-distribution samples in such a simple yet elegant way. I really learned something from this notebook.</p>",
      "rawMarkdown": "Using a variation on @ahmedelfazouan notebook, I get the following CV-LB scores.\n\n0.72 - 0.61\n0.71 - 0.6\n0.65 - 0.57\n\nI've only tried a few simple things here, but I am really surprised how well a 1D-CNN handles the in-distribution samples in such a simple yet elegant way. I really learned something from this notebook.",
      "votes": null
    },
    {
      "id": "2790830",
      "postDate": "05/03/2024 10:58:04",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a>  !</p>",
      "rawMarkdown": "Nice work @chemdatafarmer  !",
      "votes": null
    },
    {
      "id": "2790833",
      "postDate": "05/03/2024 10:58:56",
      "content": "<p>It's very much thanks to learning from you, so thank you!</p>",
      "rawMarkdown": "It's very much thanks to learning from you, so thank you!",
      "votes": null
    },
    {
      "id": "2791018",
      "postDate": "05/03/2024 12:44:31",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fabe5677ac9f32a81dfa7bdc35868e0db%2FSelection_061.png?generation=1714740269070323&amp;alt=media\"></p>\n<p>some observation:</p>\n<ol>\n<li>i train on my local machines, batch size = 5k to 10k, single gpu in pytorch.<br>\nsurprising, adding batch norm don't work as well. (meaning some bias in data, or 10k is not representative\" enough)</li>\n</ol>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fabe5677ac9f32a81dfa7bdc35868e0db%2FSelection_061.png?generation=1714740269070323&alt=media)\n\n\nsome observation:\n1. i train on my local machines, batch size = 5k to 10k, single gpu in pytorch.\nsurprising, adding batch norm don't work as well. (meaning some bias in data, or 10k is not representative\" enough)",
      "votes": null
    },
    {
      "id": "2791747",
      "postDate": "05/03/2024 20:27:41",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> my best model uses batch norm, though I haven't systematically studied the impact of taking it out </p>",
      "rawMarkdown": "hengck23 my best model uses batch norm, though I haven't systematically studied the impact of taking it out",
      "votes": null
    },
    {
      "id": "2796173",
      "postDate": "05/06/2024 06:23:41",
      "content": "<p>\"best model uses batch norm, \"</p>\n<p>i think the reason for my poorer performance could be this<br>\n<a href=\"https://medium.com/deeplearningmadeeasy/ghost-batchnorm-explained-e0fa9d651e03\" target=\"_blank\">https://medium.com/deeplearningmadeeasy/ghost-batchnorm-explained-e0fa9d651e03</a></p>\n<p>\" It seems neural networks tends to do worse for unseen data when being trained on large batch sizes.\"</p>\n<p>if you are using 8xTPU, you don't see this problem because data is split into 8 chunks and bn is applied independently to each chunk (i.e natural ghost batchnorm)</p>\n<p>it is a problem for me because i am using single gpu.</p>\n<p>Now i change bn to ghost bn and it seems to work better</p>\n<hr>\n<p>for simple 3 layer cnn1d, it doesn't matter if you use batch norm or not. But if you use deep layer, or shalow/deep layer for rnn, trasnformer, then it is important for normalisation because of the use in exp family function in activation (e.g. tahn,sigmoid in rnn or exp in softmax in attnetion)</p>",
      "rawMarkdown": "\"best model uses batch norm, \"\n\ni think the reason for my poorer performance could be this\nhttps://medium.com/deeplearningmadeeasy/ghost-batchnorm-explained-e0fa9d651e03\n\n\" It seems neural networks tends to do worse for unseen data when being trained on large batch sizes.\"\n\nif you are using 8xTPU, you don't see this problem because data is split into 8 chunks and bn is applied independently to each chunk (i.e natural ghost batchnorm)\n\nit is a problem for me because i am using single gpu.\n\nNow i change bn to ghost bn and it seems to work better\n\n---\n\nfor simple 3 layer cnn1d, it doesn't matter if you use batch norm or not. But if you use deep layer, or shalow/deep layer for rnn, trasnformer, then it is important for normalisation because of the use in exp family function in activation (e.g. tahn,sigmoid in rnn or exp in softmax in attnetion)",
      "votes": null
    },
    {
      "id": "2796700",
      "postDate": "05/06/2024 11:31:23",
      "content": "<p>Ah, yes, I didn't appreciate the differences between GPU and TPU here. Indeed, my batch norms were added when I tried to make the network deeper.</p>",
      "rawMarkdown": "Ah, yes, I didn't appreciate the differences between GPU and TPU here. Indeed, my batch norms were added when I tried to make the network deeper.",
      "votes": null
    },
    {
      "id": "2799055",
      "postDate": "05/07/2024 15:17:44",
      "content": "<table>\n<thead>\n<tr>\n<th>CV-random</th>\n<th>CV-noshare</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.664</td>\n<td>0.022</td>\n</tr>\n</tbody>\n</table>\n<p>I still have not submitted it because I am currently toying around (or rather, fighting) with RDKit for molecule representation within a 3d space, because I think that it would be quite informative to have</p>",
      "rawMarkdown": "| CV-random | CV-noshare |\n| --- | --- |\n| 0.664 | 0.022 |\n\n\nI still have not submitted it because I am currently toying around (or rather, fighting) with RDKit for molecule representation within a 3d space, because I think that it would be quite informative to have",
      "votes": null
    },
    {
      "id": "2804645",
      "postDate": "05/10/2024 06:21:09",
      "content": "<p>set all sharing blocks to zero at submission.csv.<br>\nI get nonshare sore = lb0.055.<br>\n(actually 0.055 include score of both share(truth distrubution) and nonshare binds(prediction ap)).</p>\n<p>if anyone has nonshare score, please put it here.</p>",
      "rawMarkdown": "set all sharing blocks to zero at submission.csv.\nI get nonshare sore = lb0.055.\n(actually 0.055 include score of both share(truth distrubution) and nonshare binds(prediction ap)).\n\n\nif anyone has nonshare score, please put it here.",
      "votes": null
    },
    {
      "id": "2805029",
      "postDate": "05/10/2024 10:20:33",
      "content": "<p><br>\nThe new result is quite puzzling, not sure what happened here.</p>\n<table>\n<thead>\n<tr>\n<th>LB</th>\n<th>LB (mask noshare)</th>\n<th>LB (mask share)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.629</td>\n<td>0.596</td>\n<td>0.053</td>\n</tr>\n<tr>\n<td>0.663</td>\n<td>0.596</td>\n<td>0.058</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "~~My best submission is currently overwritten by a small sample experiment, so I checked one of my past submission.~~\nThe new result is quite puzzling, not sure what happened here.\n| LB | LB (mask noshare)  | LB (mask share)|\n| --- | --- | ---|\n|  0.629| 0.596 | 0.053 |\n|  0.663 | 0.596 | 0.058 |",
      "votes": null
    },
    {
      "id": "2806318",
      "postDate": "05/11/2024 03:23:07",
      "content": "<p><a href=\"https://www.kaggle.com/w5833946\" target=\"_blank\">@w5833946</a> </p>\n<p>thanks. your results are probably correct.<br>\nif we have perfect nonshare prediction, we can get nonshare lb=0.2+</p>\n<p>check the following<br>\n<a href=\"https://www.kaggle.com/code/hengck23/how-to-probe\" target=\"_blank\">https://www.kaggle.com/code/hengck23/how-to-probe</a><br>\n<a href=\"https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set/notebook\" target=\"_blank\">https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set/notebook</a></p>",
      "rawMarkdown": "w5833946 \n\nthanks. your results are probably correct.\nif we have perfect nonshare prediction, we can get nonshare lb=0.2+\n\ncheck the following\nhttps://www.kaggle.com/code/hengck23/how-to-probe\nhttps://www.kaggle.com/code/junkoda/unknowns-in-public-test-set/notebook",
      "votes": null
    },
    {
      "id": "2806321",
      "postDate": "05/11/2024 03:25:42",
      "content": "<p>just to double check, i attached \"submit_sharing.zip\" to mask share and nonshare in submit. you may check if it is correct.</p>",
      "rawMarkdown": "just to double check, i attached \"submit_sharing.zip\" to mask share and nonshare in submit. you may check if it is correct.",
      "votes": null
    },
    {
      "id": "2806324",
      "postDate": "05/11/2024 03:28:37",
      "content": "<p>Thank you, I will check it this night.</p>",
      "rawMarkdown": "Thank you, I will check it this night.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2785587,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "04/30/2024 22:56:03",
      "content": "<p>My CV LB split is very split dependent. The more samples I add to my experiments the lower my CV (using scaffold splitting) but higher LB. I'd put some examples here, but I am still experimenting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2785685,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/01/2024 01:45:03",
          "content": "<p>the best split is probably random, see also paper.<br>\ni suggest 2 separate splits and experiments</p>\n<p>my strategy:<br>\n[1] cv score versus num of share blocks</p>\n<ul>\n<li>train samples = share blocks</li>\n<li>valid samples with 0,1,2,3 share blocks<br>\nmake a plot of cv score versus num of share blocks<br>\nmake tsnet map or knn results for the blocks<br>\ntry some external dataset<br>\nmy conclusion here so far is that for non share blocks, results are really bad (need to solve later)</li>\n</ul>\n<p>[2] cv score for share blocks</p>\n<ul>\n<li>best results is probably using random split and all samples</li>\n</ul>\n<p>paper on cv split, datasize, etc:<br>\nUncovering Neural Scaling Laws in Molecular Representation Learning<br>\n<a href=\"https://arxiv.org/abs/2309.15123\" target=\"_blank\">https://arxiv.org/abs/2309.15123</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2786617,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "05/01/2024 11:27:41",
              "content": "<p>I agree there will likely need to be ways to handle compounds close to the train set and those that aren't. I have 64gb of RAM at home and it seems that even with np.packbits I can't train models with half of the train rows. So, now I'm focused on finding a subsampling strategy. I'm learning a lot about how to manage large datasets in this competition.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2786912,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "05/01/2024 13:47:12",
                  "content": "<p>actually i am think that for those kagglers that shown good performance, espeically those with good domain knowledge, google should get them better compute power, e.g. discount price at google cloud for better gpu, etc</p>\n<p>or kaggler ourselves can setup some fund.<br>\nif the solution is good, then we can all learn something valuable, rather than trial and error like a headless fly</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2787630,
                      "author_name": "chemdatafarmer",
                      "author_url": "",
                      "post_date": "05/01/2024 21:22:27",
                      "content": "<p>I've considered renting more compute power, but I do think exploring a subsampling strategy is a good learning experience for me. </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2790293,
                          "author_name": "amishm",
                          "author_url": "",
                          "post_date": "05/03/2024 05:40:06",
                          "content": "<p>i load my data,  and sample 1024 * 1024 molecules, with X% positive samples, and save it to a parquet file. I do this 100 times. For training i load one file at a time to save ram. You could also have a separate file for just postives , and multiple files for negative if you want to change +ve/-ve split on the fly.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2785690,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/01/2024 01:50:00",
      "content": "<p>\"I'm using my CV split, here are most recent results:\"</p>\n<p>both molformer and chemberta should be about to give you lb 0.575. (CV 0.670), i am using random splits<br>\nI think this is the limit of SMILES string transformer?<br>\n(i think some kaggler can get up to lb0.605(CV 0.700))?</p>\n<p>molformer = 12 layers<br>\nchemberta = 3 layers</p>\n<p>different pretrain chemberta 5,10,77M MLM,MTR has slight different results</p>\n<p>Note there a 2 different molformer in papers (one is based on SMILES string , another is a GNN)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2790563,
          "author_name": "amishm",
          "author_url": "",
          "post_date": "05/03/2024 08:31:35",
          "content": "<p>molformer works slightly better for me.<br>\nWith Molformer i get .69 CV, .583 lb.<br>\nWith chemberta i get .65 CV, .56 lb. </p>\n<p>i tried using Selformer, ChemGPT (selfies models) but got much worse scores.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2787584,
      "author_name": "ahmedelfazouan",
      "author_url": "",
      "post_date": "05/01/2024 20:36:32",
      "content": "<p>CV 0.69, LB 0.585<br>\nmodel : 1dcnn trained on all data<br>\nI made this model public : <a href=\"https://www.kaggle.com/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">https://www.kaggle.com/ahmedelfazouan/belka-1dcnn-starter-with-all-data</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2787594,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "05/01/2024 20:46:46",
          "content": "<p>This is pretty cool <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> ! Looking forward to reading through it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2790810,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "05/03/2024 10:51:02",
              "content": "<p>Using a variation on <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> notebook, I get the following CV-LB scores.</p>\n<p>0.72 - 0.61<br>\n0.71 - 0.6<br>\n0.65 - 0.57</p>\n<p>I've only tried a few simple things here, but I am really surprised how well a 1D-CNN handles the in-distribution samples in such a simple yet elegant way. I really learned something from this notebook.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2790830,
                  "author_name": "ahmedelfazouan",
                  "author_url": "",
                  "post_date": "05/03/2024 10:58:04",
                  "content": "<p>Nice work <a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a>  !</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2790833,
                      "author_name": "chemdatafarmer",
                      "author_url": "",
                      "post_date": "05/03/2024 10:58:56",
                      "content": "<p>It's very much thanks to learning from you, so thank you!</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2791018,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "05/03/2024 12:44:31",
                          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fabe5677ac9f32a81dfa7bdc35868e0db%2FSelection_061.png?generation=1714740269070323&amp;alt=media\"></p>\n<p>some observation:</p>\n<ol>\n<li>i train on my local machines, batch size = 5k to 10k, single gpu in pytorch.<br>\nsurprising, adding batch norm don't work as well. (meaning some bias in data, or 10k is not representative\" enough)</li>\n</ol>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2791747,
                              "author_name": "chemdatafarmer",
                              "author_url": "",
                              "post_date": "05/03/2024 20:27:41",
                              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> my best model uses batch norm, though I haven't systematically studied the impact of taking it out </p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2796173,
                                  "author_name": "hengck23",
                                  "author_url": "",
                                  "post_date": "05/06/2024 06:23:41",
                                  "content": "<p>\"best model uses batch norm, \"</p>\n<p>i think the reason for my poorer performance could be this<br>\n<a href=\"https://medium.com/deeplearningmadeeasy/ghost-batchnorm-explained-e0fa9d651e03\" target=\"_blank\">https://medium.com/deeplearningmadeeasy/ghost-batchnorm-explained-e0fa9d651e03</a></p>\n<p>\" It seems neural networks tends to do worse for unseen data when being trained on large batch sizes.\"</p>\n<p>if you are using 8xTPU, you don't see this problem because data is split into 8 chunks and bn is applied independently to each chunk (i.e natural ghost batchnorm)</p>\n<p>it is a problem for me because i am using single gpu.</p>\n<p>Now i change bn to ghost bn and it seems to work better</p>\n<hr>\n<p>for simple 3 layer cnn1d, it doesn't matter if you use batch norm or not. But if you use deep layer, or shalow/deep layer for rnn, trasnformer, then it is important for normalisation because of the use in exp family function in activation (e.g. tahn,sigmoid in rnn or exp in softmax in attnetion)</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2796700,
                                      "author_name": "chemdatafarmer",
                                      "author_url": "",
                                      "post_date": "05/06/2024 11:31:23",
                                      "content": "<p>Ah, yes, I didn't appreciate the differences between GPU and TPU here. Indeed, my batch norms were added when I tried to make the network deeper.</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2788308,
      "author_name": "sroger",
      "author_url": "",
      "post_date": "05/02/2024 07:17:13",
      "content": "<p>CV: 0.597, 0.600, 0.608, 0.613<br>\nLB: 0.575, 0.586, 0.589, 0.585</p>\n<p>Model: small nn on top of frozen molformer<br>\nVal Split: positive ratio 1% (random, len 1e5)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2788485,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/02/2024 08:51:47",
          "content": "<p>your cv looks low.<br>\ni am not sure if it is the effect of different split or frozen (i think frozen is more likely)</p>\n<p>the good news is that CV and LB gap is very small, which means that you actually score very well in nonshare b block (you can confirm by local experiment or separate submission)</p>\n<p>and i think if you starts unfreeze, you will start to have different results.<br>\nyou can try ensemble of freeze and frozen models, or using freeze and frozen models for share and unshare blocks</p>",
          "votes": null,
          "replies": [
            {
              "id": 2788579,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "05/02/2024 09:44:57",
              "content": "<p>i suddenly have an idea.<br>\nhow about LORA like model?<br>\nuse some layers to modify molformer intermediate or/and final layer output?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2788934,
                  "author_name": "sroger",
                  "author_url": "",
                  "post_date": "05/02/2024 13:37:36",
                  "content": "<p>Hmmm will have to look into loras, definitely seems like a possible path forward.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2788682,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "05/02/2024 10:54:42",
              "content": "<p>I think there are still some lower hanging fruits that can be picked. These models were trained (with 4090*1) in around 5-7 hours with a subset of (neg) train, only iterating over the train set about twice.</p>\n<p>Before I set the fixed positive ratio val, a different random split/sample had cv around 0.62, so yeah this split is lower.</p>\n<p>I don't think I'll unfreeze molformer just yet as that adds a massive amount of compute cost and will destroy my batch size.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2799055,
      "author_name": "giorgiomicaletto",
      "author_url": "",
      "post_date": "05/07/2024 15:17:44",
      "content": "<table>\n<thead>\n<tr>\n<th>CV-random</th>\n<th>CV-noshare</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.664</td>\n<td>0.022</td>\n</tr>\n</tbody>\n</table>\n<p>I still have not submitted it because I am currently toying around (or rather, fighting) with RDKit for molecule representation within a 3d space, because I think that it would be quite informative to have</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2804645,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/10/2024 06:21:09",
      "content": "<p>set all sharing blocks to zero at submission.csv.<br>\nI get nonshare sore = lb0.055.<br>\n(actually 0.055 include score of both share(truth distrubution) and nonshare binds(prediction ap)).</p>\n<p>if anyone has nonshare score, please put it here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2805029,
          "author_name": "w5833946",
          "author_url": "",
          "post_date": "05/10/2024 10:20:33",
          "content": "<p><br>\nThe new result is quite puzzling, not sure what happened here.</p>\n<table>\n<thead>\n<tr>\n<th>LB</th>\n<th>LB (mask noshare)</th>\n<th>LB (mask share)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.629</td>\n<td>0.596</td>\n<td>0.053</td>\n</tr>\n<tr>\n<td>0.663</td>\n<td>0.596</td>\n<td>0.058</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": [
            {
              "id": 2806318,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "05/11/2024 03:23:07",
              "content": "<p><a href=\"https://www.kaggle.com/w5833946\" target=\"_blank\">@w5833946</a> </p>\n<p>thanks. your results are probably correct.<br>\nif we have perfect nonshare prediction, we can get nonshare lb=0.2+</p>\n<p>check the following<br>\n<a href=\"https://www.kaggle.com/code/hengck23/how-to-probe\" target=\"_blank\">https://www.kaggle.com/code/hengck23/how-to-probe</a><br>\n<a href=\"https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set/notebook\" target=\"_blank\">https://www.kaggle.com/code/junkoda/unknowns-in-public-test-set/notebook</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2806321,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "05/11/2024 03:25:42",
                  "content": "<p>just to double check, i attached \"submit_sharing.zip\" to mask share and nonshare in submit. you may check if it is correct.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2806324,
                      "author_name": "w5833946",
                      "author_url": "",
                      "post_date": "05/11/2024 03:28:37",
                      "content": "<p>Thank you, I will check it this night.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2784340": "Please share your CV/LB results here. \n\nI'm using my [CV split](https://www.kaggle.com/datasets/thedrcat/belka-cv-split), here are most recent results:\n\n| Model | CV-random | CV-noshare | LB |\n| --- | --- | --- | --- |\n| Molformer | 0.64 | 0.008 | 0.564 |\n| Molformer | 0.62 | 0.017 | 0.554 |\n\n\nFrom other discussions and my initial results it's likely that public LB seems correlated with random split, but private results will most likely depend on how well you can model the no-share / test distribution. \n\nBelow more details by protein target (if you're not using [W&B](https://wandb.ai) for tracking, you're missing out):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1328185%2Fa57a0fd2ce86c30dd3b0345958eb2c4f%2FScreenshot%202024-04-30%20at%2009.46.49.png?generation=1714463680621641&alt=media)",
    "2785587": "My CV LB split is very split dependent. The more samples I add to my experiments the lower my CV (using scaffold splitting) but higher LB. I'd put some examples here, but I am still experimenting.",
    "2785685": "the best split is probably random, see also paper.\ni suggest 2 separate splits and experiments\n\nmy strategy:\n[1] cv score versus num of share blocks\n- train samples = share blocks\n- valid samples with 0,1,2,3 share blocks\nmake a plot of cv score versus num of share blocks\nmake tsnet map or knn results for the blocks\ntry some external dataset\nmy conclusion here so far is that for non share blocks, results are really bad (need to solve later)\n\n[2] cv score for share blocks\n- best results is probably using random split and all samples\n\n\n\npaper on cv split, datasize, etc:\nUncovering Neural Scaling Laws in Molecular Representation Learning\nhttps://arxiv.org/abs/2309.15123",
    "2785690": "\"I'm using my CV split, here are most recent results:\"\n\nboth molformer and chemberta should be about to give you lb 0.575. (CV 0.670), i am using random splits\nI think this is the limit of SMILES string transformer?\n(i think some kaggler can get up to lb0.605(CV 0.700))?\n\nmolformer = 12 layers\nchemberta = 3 layers\n\ndifferent pretrain chemberta 5,10,77M MLM,MTR has slight different results\n\nNote there a 2 different molformer in papers (one is based on SMILES string , another is a GNN)",
    "2786617": "I agree there will likely need to be ways to handle compounds close to the train set and those that aren't. I have 64gb of RAM at home and it seems that even with np.packbits I can't train models with half of the train rows. So, now I'm focused on finding a subsampling strategy. I'm learning a lot about how to manage large datasets in this competition.",
    "2786912": "actually i am think that for those kagglers that shown good performance, espeically those with good domain knowledge, google should get them better compute power, e.g. discount price at google cloud for better gpu, etc\n\nor kaggler ourselves can setup some fund.\nif the solution is good, then we can all learn something valuable, rather than trial and error like a headless fly",
    "2787584": "CV 0.69, LB 0.585\nmodel : 1dcnn trained on all data\nI made this model public : https://www.kaggle.com/ahmedelfazouan/belka-1dcnn-starter-with-all-data",
    "2787594": "This is pretty cool @ahmedelfazouan ! Looking forward to reading through it.",
    "2787630": "I've considered renting more compute power, but I do think exploring a subsampling strategy is a good learning experience for me.",
    "2788308": "CV: 0.597, 0.600, 0.608, 0.613\nLB: 0.575, 0.586, 0.589, 0.585\n\nModel: small nn on top of frozen molformer\nVal Split: positive ratio 1% (random, len 1e5)",
    "2788485": "your cv looks low.\ni am not sure if it is the effect of different split or frozen (i think frozen is more likely)\n\nthe good news is that CV and LB gap is very small, which means that you actually score very well in nonshare b block (you can confirm by local experiment or separate submission)\n\nand i think if you starts unfreeze, you will start to have different results.\nyou can try ensemble of freeze and frozen models, or using freeze and frozen models for share and unshare blocks",
    "2788579": "i suddenly have an idea.\nhow about LORA like model?\nuse some layers to modify molformer intermediate or/and final layer output?",
    "2788682": "I think there are still some lower hanging fruits that can be picked. These models were trained (with 4090*1) in around 5-7 hours with a subset of (neg) train, only iterating over the train set about twice.\n\nBefore I set the fixed positive ratio val, a different random split/sample had cv around 0.62, so yeah this split is lower.\n\nI don't think I'll unfreeze molformer just yet as that adds a massive amount of compute cost and will destroy my batch size.",
    "2788934": "Hmmm will have to look into loras, definitely seems like a possible path forward.",
    "2790293": "i load my data,  and sample 1024 * 1024 molecules, with X% positive samples, and save it to a parquet file. I do this 100 times. For training i load one file at a time to save ram. You could also have a separate file for just postives , and multiple files for negative if you want to change +ve/-ve split on the fly.",
    "2790563": "molformer works slightly better for me.\nWith Molformer i get .69 CV, .583 lb.\nWith chemberta i get .65 CV, .56 lb. \n\ni tried using Selformer, ChemGPT (selfies models) but got much worse scores.",
    "2790810": "Using a variation on @ahmedelfazouan notebook, I get the following CV-LB scores.\n\n0.72 - 0.61\n0.71 - 0.6\n0.65 - 0.57\n\nI've only tried a few simple things here, but I am really surprised how well a 1D-CNN handles the in-distribution samples in such a simple yet elegant way. I really learned something from this notebook.",
    "2790830": "Nice work @chemdatafarmer  !",
    "2790833": "It's very much thanks to learning from you, so thank you!",
    "2791018": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fabe5677ac9f32a81dfa7bdc35868e0db%2FSelection_061.png?generation=1714740269070323&alt=media)\n\n\nsome observation:\n1. i train on my local machines, batch size = 5k to 10k, single gpu in pytorch.\nsurprising, adding batch norm don't work as well. (meaning some bias in data, or 10k is not representative\" enough)",
    "2791747": "hengck23 my best model uses batch norm, though I haven't systematically studied the impact of taking it out",
    "2796173": "\"best model uses batch norm, \"\n\ni think the reason for my poorer performance could be this\nhttps://medium.com/deeplearningmadeeasy/ghost-batchnorm-explained-e0fa9d651e03\n\n\" It seems neural networks tends to do worse for unseen data when being trained on large batch sizes.\"\n\nif you are using 8xTPU, you don't see this problem because data is split into 8 chunks and bn is applied independently to each chunk (i.e natural ghost batchnorm)\n\nit is a problem for me because i am using single gpu.\n\nNow i change bn to ghost bn and it seems to work better\n\n---\n\nfor simple 3 layer cnn1d, it doesn't matter if you use batch norm or not. But if you use deep layer, or shalow/deep layer for rnn, trasnformer, then it is important for normalisation because of the use in exp family function in activation (e.g. tahn,sigmoid in rnn or exp in softmax in attnetion)",
    "2796700": "Ah, yes, I didn't appreciate the differences between GPU and TPU here. Indeed, my batch norms were added when I tried to make the network deeper.",
    "2799055": "| CV-random | CV-noshare |\n| --- | --- |\n| 0.664 | 0.022 |\n\n\nI still have not submitted it because I am currently toying around (or rather, fighting) with RDKit for molecule representation within a 3d space, because I think that it would be quite informative to have",
    "2804645": "set all sharing blocks to zero at submission.csv.\nI get nonshare sore = lb0.055.\n(actually 0.055 include score of both share(truth distrubution) and nonshare binds(prediction ap)).\n\n\nif anyone has nonshare score, please put it here.",
    "2805029": "~~My best submission is currently overwritten by a small sample experiment, so I checked one of my past submission.~~\nThe new result is quite puzzling, not sure what happened here.\n| LB | LB (mask noshare)  | LB (mask share)|\n| --- | --- | ---|\n|  0.629| 0.596 | 0.053 |\n|  0.663 | 0.596 | 0.058 |",
    "2806318": "w5833946 \n\nthanks. your results are probably correct.\nif we have perfect nonshare prediction, we can get nonshare lb=0.2+\n\ncheck the following\nhttps://www.kaggle.com/code/hengck23/how-to-probe\nhttps://www.kaggle.com/code/junkoda/unknowns-in-public-test-set/notebook",
    "2806321": "just to double check, i attached \"submit_sharing.zip\" to mask share and nonshare in submit. you may check if it is correct.",
    "2806324": "Thank you, I will check it this night."
  },
  "source": "meta"
}