{
  "id": 519135,
  "title": "988th Place solution (4th on Public) - GNN and domain adaption. 800Kb model.",
  "url": "/competitions/leash-BELKA/discussion/519135",
  "author_name": "yamu_duck",
  "post_date": "2024-07-09T20:37:48.568000",
  "votes": 29,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I was holding 1st place for a week and dropped to 4th in the end and then to 988th in private LB, so this is a story full of sadness.</p>\n<p>**Baseline **<br>\nAs everyone else for the first month I played with ECFP and other fingerprints, XGboost et al. Then I concluded that it will not generalize well. After some poking around I found a solution that quickly took took me up - GNNs.</p>\n<p>The trick was to generate a training set with rich atom / bond properties annotation - charges, molecular weights, bond types, aromaticity, etc - basically everything you can squeeze from RDKit without heavy computation. Pack all this into pre-processed and pre-splitted CV folds and store on SSD. Took like 12 hours per split to generate and 1TB of storage, but made training GNNs quite fast on a single GPU that I used and model size was like 800kb.</p>\n<p><strong>Shared BB</strong></p>\n<p>Best result on shared BB was 0.344 for a single GAT based GNN (3 class classification). It was training on a CV split and then finetuned on full set. I just used one split, seems pretraining on others was giving worse results. It was a \"lucky\" fold. Seems not so lucky after all.</p>\n<p>** non-shared BB*</p>\n<p>For non-shared BB I used a domain adaptation loss function from <a href=\"https://arxiv.org/pdf/2306.04979\" target=\"_blank\">https://arxiv.org/pdf/2306.04979</a> paper. I was doing classification loss on full set and domain adaption (basically soft labels generated based on embedding cosines) and this gave me additional public LB boost from 0.498 -&gt; 0.517</p>\n<p>I was quite sure that this domain adaptation protocol will save my *** from shake up. </p>",
  "messages": [
    {
      "id": 2914294,
      "postDate": "2024-07-09T20:37:48.570Z",
      "content": "<p>I was holding 1st place for a week and dropped to 4th in the end and then to 988th in private LB, so this is a story full of sadness.</p>\n<p>**Baseline **<br>\nAs everyone else for the first month I played with ECFP and other fingerprints, XGboost et al. Then I concluded that it will not generalize well. After some poking around I found a solution that quickly took took me up - GNNs.</p>\n<p>The trick was to generate a training set with rich atom / bond properties annotation - charges, molecular weights, bond types, aromaticity, etc - basically everything you can squeeze from RDKit without heavy computation. Pack all this into pre-processed and pre-splitted CV folds and store on SSD. Took like 12 hours per split to generate and 1TB of storage, but made training GNNs quite fast on a single GPU that I used and model size was like 800kb.</p>\n<p><strong>Shared BB</strong></p>\n<p>Best result on shared BB was 0.344 for a single GAT based GNN (3 class classification). It was training on a CV split and then finetuned on full set. I just used one split, seems pretraining on others was giving worse results. It was a \"lucky\" fold. Seems not so lucky after all.</p>\n<p>** non-shared BB*</p>\n<p>For non-shared BB I used a domain adaptation loss function from <a href=\"https://arxiv.org/pdf/2306.04979\" target=\"_blank\">https://arxiv.org/pdf/2306.04979</a> paper. I was doing classification loss on full set and domain adaption (basically soft labels generated based on embedding cosines) and this gave me additional public LB boost from 0.498 -&gt; 0.517</p>\n<p>I was quite sure that this domain adaptation protocol will save my *** from shake up. </p>",
      "rawMarkdown": "I was holding 1st place for a week and dropped to 4th in the end and then to 988th in private LB, so this is a story full of sadness.\n\n**Baseline **\nAs everyone else for the first month I played with ECFP and other fingerprints, XGboost et al. Then I concluded that it will not generalize well. After some poking around I found a solution that quickly took took me up - GNNs.\n\nThe trick was to generate a training set with rich atom / bond properties annotation - charges, molecular weights, bond types, aromaticity, etc - basically everything you can squeeze from RDKit without heavy computation. Pack all this into pre-processed and pre-splitted CV folds and store on SSD. Took like 12 hours per split to generate and 1TB of storage, but made training GNNs quite fast on a single GPU that I used and model size was like 800kb.\n\n**Shared BB**\n\nBest result on shared BB was 0.344 for a single GAT based GNN (3 class classification). It was training on a CV split and then finetuned on full set. I just used one split, seems pretraining on others was giving worse results. It was a \"lucky\" fold. Seems not so lucky after all.\n\n** non-shared BB*\n\nFor non-shared BB I used a domain adaptation loss function from https://arxiv.org/pdf/2306.04979 paper. I was doing classification loss on full set and domain adaption (basically soft labels generated based on embedding cosines) and this gave me additional public LB boost from 0.498 -> 0.517\n\nI was quite sure that this domain adaptation protocol will save my *** from shake up. ",
      "votes": 29
    },
    {
      "id": 2946074,
      "postDate": "2024-08-04T00:04:07.883Z",
      "content": "<p>This is very interesting, you may check this <a href=\"https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group\" target=\"_blank\">notebook</a> or <a href=\"https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group\" target=\"_blank\">this one</a> to see which group in the submission you get the highest score. <br>\nWill you share your code?</p>",
      "rawMarkdown": "This is very interesting, you may check this [notebook](https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group) or [this one](https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group) to see which group in the submission you get the highest score. \nWill you share your code?"
    },
    {
      "id": 2914729,
      "postDate": "2024-07-10T06:14:22.410Z",
      "content": "<p>Thanks for your sharing! Did you try use the atom properties only? I have a GNN model using the atom properties only, for the bond I just use the adjacency matrix. This model have only 0.398 in public LB. I'm curious about if it is cause by the loss of bond features. It seems that there are many GNNs with the bond features achieve a very good score on public LB. <br>\n Another thing is that can I ask what kind of strategies do you use to accerate the data IO? I tried to cache the preprocessed features using numpy.memmap or lmdb, but the speed is still very slow and even slower than instant calculation in dataloader of pytorch. For training one epoch (whole dataset) of GNN in A100, it takes 6-8 hours. </p>",
      "rawMarkdown": "Thanks for your sharing! Did you try use the atom properties only? I have a GNN model using the atom properties only, for the bond I just use the adjacency matrix. This model have only 0.398 in public LB. I'm curious about if it is cause by the loss of bond features. It seems that there are many GNNs with the bond features achieve a very good score on public LB. \n Another thing is that can I ask what kind of strategies do you use to accerate the data IO? I tried to cache the preprocessed features using numpy.memmap or lmdb, but the speed is still very slow and even slower than instant calculation in dataloader of pytorch. For training one epoch (whole dataset) of GNN in A100, it takes 6-8 hours. ",
      "replies": [
        {
          "id": 2914860,
          "postDate": "2024-07-10T08:01:19.453Z",
          "content": "<p>I tried this model.  It's very slow to train.  I am not even sure this is real GNN.  There is no real graph in the model.  I think \"adjacency matrix\" might get updated after a few layers.  Zeros may not be kept as zeros.  That would break the graph.  But I did n't look at the edge adjacent matrices of the intermediate layers to know if this is indeed the case.  </p>",
          "rawMarkdown": "I tried this model.  It's very slow to train.  I am not even sure this is real GNN.  There is no real graph in the model.  I think \"adjacency matrix\" might get updated after a few layers.  Zeros may not be kept as zeros.  That would break the graph.  But I did n't look at the edge adjacent matrices of the intermediate layers to know if this is indeed the case.  ",
          "replies": [
            {
              "id": 2915129,
              "postDate": "2024-07-10T10:58:38.943Z",
              "content": "<p>My GNN is normal GCN, which use the adjacent matrices directly and the adjacent matrices are always adjacent matrices. In my experiments, the Transformer/GNN-based model can have a 100% usage of GPU, which means they are almost the fastest (6-8 hours/one epoch). For RNN/CNN/MLP-like model, the usage is less than 30%, so if the IO can be speed, then these models will be faster. I remember there is a SMILES-based CNN model released in Discussion before, which need only hundreds of seconds (~5min) for one epoch. It use some techniques like TPU-based data cluster storage. I spend half of the time to optimize the speed, like using numpy.memmap/lmdb to cache fingerprint/atom/bond features, but all didn't work well. </p>",
              "rawMarkdown": "My GNN is normal GCN, which use the adjacent matrices directly and the adjacent matrices are always adjacent matrices. In my experiments, the Transformer/GNN-based model can have a 100% usage of GPU, which means they are almost the fastest (6-8 hours/one epoch). For RNN/CNN/MLP-like model, the usage is less than 30%, so if the IO can be speed, then these models will be faster. I remember there is a SMILES-based CNN model released in Discussion before, which need only hundreds of seconds (~5min) for one epoch. It use some techniques like TPU-based data cluster storage. I spend half of the time to optimize the speed, like using numpy.memmap/lmdb to cache fingerprint/atom/bond features, but all didn't work well. ",
              "votes": 1
            },
            {
              "id": 2916059,
              "postDate": "2024-07-10T18:44:18.563Z",
              "content": "<p>I guess I tried a different model.  dhttps://docs.dgl.ai/generated/dgl.nn.pytorch.gt.EGTLayer.html#dgl.nn.pytorch.gt.EGTLayer</p>",
              "rawMarkdown": "I guess I tried a different model.  dhttps://docs.dgl.ai/generated/dgl.nn.pytorch.gt.EGTLayer.html#dgl.nn.pytorch.gt.EGTLayer"
            },
            {
              "id": 2917960,
              "postDate": "2024-07-11T23:20:57.413Z",
              "content": "<p><a href=\"https://github.com/crazyleg/belka-kaggle/blob/master/systematic/mega_ml/model.py\" target=\"_blank\">https://github.com/crazyleg/belka-kaggle/blob/master/systematic/mega_ml/model.py</a> -&lt; GATBasedMolecularGraphNeuralNetwork194 is my main model</p>\n<p><a href=\"https://github.com/crazyleg/belka-kaggle/blob/master/systematic/split_generation/generate_datasets.py\" target=\"_blank\">https://github.com/crazyleg/belka-kaggle/blob/master/systematic/split_generation/generate_datasets.py</a> -&lt; that's the dataset generation and atom/edge annotation</p>",
              "rawMarkdown": "https://github.com/crazyleg/belka-kaggle/blob/master/systematic/mega_ml/model.py -< GATBasedMolecularGraphNeuralNetwork194 is my main model\n\nhttps://github.com/crazyleg/belka-kaggle/blob/master/systematic/split_generation/generate_datasets.py -< that's the dataset generation and atom/edge annotation\n",
              "votes": 1
            },
            {
              "id": 2917962,
              "postDate": "2024-07-11T23:22:04.470Z",
              "content": "<p>it is super fast to train if you do data proprocessing right, take 1.4h per epoch on 2080 Ti. </p>",
              "rawMarkdown": "it is super fast to train if you do data proprocessing right, take 1.4h per epoch on 2080 Ti. ",
              "votes": 1
            },
            {
              "id": 2918249,
              "postDate": "2024-07-12T05:38:14.100Z",
              "content": "<p>Do you store the characteristics of each molecule in a single file or all in one file? Use something like pickle.dump(…)?</p>",
              "rawMarkdown": "Do you store the characteristics of each molecule in a single file or all in one file? Use something like pickle.dump(...)?"
            },
            {
              "id": 2918537,
              "postDate": "2024-07-12T10:14:51.033Z",
              "content": "<p>I stored a full pre-processed batch of 1024 molecules and labels in a torch_geometric.Batch and used torch.save to save/load</p>",
              "rawMarkdown": "I stored a full pre-processed batch of 1024 molecules and labels in a torch_geometric.Batch and used torch.save to save/load",
              "votes": 1
            },
            {
              "id": 2921151,
              "postDate": "2024-07-14T05:59:40.043Z",
              "content": "<p>OK. I think that is the case. Dumping each batch of preprocessed data into a separate file is the most efficient way for IO. I was considering using lmdb/memmap to dump all preprocessed data into one file because I was afraid that dumping each batch separately would make shuffle insufficient and affect model performance. But lmdb/memmap is still very slow. Thank you for your tricks of IO. </p>",
              "rawMarkdown": "OK. I think that is the case. Dumping each batch of preprocessed data into a separate file is the most efficient way for IO. I was considering using lmdb/memmap to dump all preprocessed data into one file because I was afraid that dumping each batch separately would make shuffle insufficient and affect model performance. But lmdb/memmap is still very slow. Thank you for your tricks of IO. "
            }
          ]
        }
      ]
    },
    {
      "id": 2914304,
      "postDate": "2024-07-09T20:56:21.747Z",
      "content": "<p>Interesting story.  How did you get public LB 0.498 score?  </p>",
      "rawMarkdown": "Interesting story.  How did you get public LB 0.498 score?  ",
      "replies": [
        {
          "id": 2914307,
          "postDate": "2024-07-09T21:01:03.947Z",
          "content": "<p>0.498 was a single GNN with no domain adaptation. 0.498 -&gt; 0.517 was when I added a Cross domain loss function from the paper mentioned. 0.344 mention is the same GNN result on a shared bb (non-shared was masked)</p>",
          "rawMarkdown": "0.498 was a single GNN with no domain adaptation. 0.498 -> 0.517 was when I added a Cross domain loss function from the paper mentioned. 0.344 mention is the same GNN result on a shared bb (non-shared was masked)",
          "votes": 2,
          "replies": [
            {
              "id": 2914437,
              "postDate": "2024-07-09T23:10:47.850Z",
              "content": "<p>That's pretty impressive to get 0.498.  Is this from one fold or the whole training data or averaging of a few folds?  </p>",
              "rawMarkdown": "That's pretty impressive to get 0.498.  Is this from one fold or the whole training data or averaging of a few folds?  "
            },
            {
              "id": 2917963,
              "postDate": "2024-07-11T23:22:21.760Z",
              "content": "<p>pretrain on split, finetune on full</p>",
              "rawMarkdown": "pretrain on split, finetune on full"
            }
          ]
        }
      ]
    },
    {
      "id": 2914524,
      "postDate": "2024-07-10T02:20:03.587Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2946074,
      "author_name": "Swikwislkdjc",
      "author_url": "",
      "post_date": "2024-08-04T00:04:07.883000",
      "content": "<p>This is very interesting, you may check this <a href=\"https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group\" target=\"_blank\">notebook</a> or <a href=\"https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group\" target=\"_blank\">this one</a> to see which group in the submission you get the highest score. <br>\nWill you share your code?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2914729,
      "author_name": "Yifan Wu",
      "author_url": "",
      "post_date": "2024-07-10T06:14:22.410000",
      "content": "<p>Thanks for your sharing! Did you try use the atom properties only? I have a GNN model using the atom properties only, for the bond I just use the adjacency matrix. This model have only 0.398 in public LB. I'm curious about if it is cause by the loss of bond features. It seems that there are many GNNs with the bond features achieve a very good score on public LB. <br>\n Another thing is that can I ask what kind of strategies do you use to accerate the data IO? I tried to cache the preprocessed features using numpy.memmap or lmdb, but the speed is still very slow and even slower than instant calculation in dataloader of pytorch. For training one epoch (whole dataset) of GNN in A100, it takes 6-8 hours. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2914860,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "2024-07-10T08:01:19.453000",
          "content": "<p>I tried this model.  It's very slow to train.  I am not even sure this is real GNN.  There is no real graph in the model.  I think \"adjacency matrix\" might get updated after a few layers.  Zeros may not be kept as zeros.  That would break the graph.  But I did n't look at the edge adjacent matrices of the intermediate layers to know if this is indeed the case.  </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2915129,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-10T10:58:38.943000",
              "content": "<p>My GNN is normal GCN, which use the adjacent matrices directly and the adjacent matrices are always adjacent matrices. In my experiments, the Transformer/GNN-based model can have a 100% usage of GPU, which means they are almost the fastest (6-8 hours/one epoch). For RNN/CNN/MLP-like model, the usage is less than 30%, so if the IO can be speed, then these models will be faster. I remember there is a SMILES-based CNN model released in Discussion before, which need only hundreds of seconds (~5min) for one epoch. It use some techniques like TPU-based data cluster storage. I spend half of the time to optimize the speed, like using numpy.memmap/lmdb to cache fingerprint/atom/bond features, but all didn't work well. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2916059,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2024-07-10T18:44:18.563000",
              "content": "<p>I guess I tried a different model.  dhttps://docs.dgl.ai/generated/dgl.nn.pytorch.gt.EGTLayer.html#dgl.nn.pytorch.gt.EGTLayer</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2917960,
              "author_name": "yamu_duck",
              "author_url": "",
              "post_date": "2024-07-11T23:20:57.413000",
              "content": "<p><a href=\"https://github.com/crazyleg/belka-kaggle/blob/master/systematic/mega_ml/model.py\" target=\"_blank\">https://github.com/crazyleg/belka-kaggle/blob/master/systematic/mega_ml/model.py</a> -&lt; GATBasedMolecularGraphNeuralNetwork194 is my main model</p>\n<p><a href=\"https://github.com/crazyleg/belka-kaggle/blob/master/systematic/split_generation/generate_datasets.py\" target=\"_blank\">https://github.com/crazyleg/belka-kaggle/blob/master/systematic/split_generation/generate_datasets.py</a> -&lt; that's the dataset generation and atom/edge annotation</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2917962,
              "author_name": "yamu_duck",
              "author_url": "",
              "post_date": "2024-07-11T23:22:04.470000",
              "content": "<p>it is super fast to train if you do data proprocessing right, take 1.4h per epoch on 2080 Ti. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2918249,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-12T05:38:14.100000",
              "content": "<p>Do you store the characteristics of each molecule in a single file or all in one file? Use something like pickle.dump(…)?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2918537,
              "author_name": "yamu_duck",
              "author_url": "",
              "post_date": "2024-07-12T10:14:51.033000",
              "content": "<p>I stored a full pre-processed batch of 1024 molecules and labels in a torch_geometric.Batch and used torch.save to save/load</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2921151,
              "author_name": "Yifan Wu",
              "author_url": "",
              "post_date": "2024-07-14T05:59:40.043000",
              "content": "<p>OK. I think that is the case. Dumping each batch of preprocessed data into a separate file is the most efficient way for IO. I was considering using lmdb/memmap to dump all preprocessed data into one file because I was afraid that dumping each batch separately would make shuffle insufficient and affect model performance. But lmdb/memmap is still very slow. Thank you for your tricks of IO. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2914304,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2024-07-09T20:56:21.747000",
      "content": "<p>Interesting story.  How did you get public LB 0.498 score?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2914307,
          "author_name": "yamu_duck",
          "author_url": "",
          "post_date": "2024-07-09T21:01:03.947000",
          "content": "<p>0.498 was a single GNN with no domain adaptation. 0.498 -&gt; 0.517 was when I added a Cross domain loss function from the paper mentioned. 0.344 mention is the same GNN result on a shared bb (non-shared was masked)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2914437,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2024-07-09T23:10:47.850000",
              "content": "<p>That's pretty impressive to get 0.498.  Is this from one fold or the whole training data or averaging of a few folds?  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2917963,
              "author_name": "yamu_duck",
              "author_url": "",
              "post_date": "2024-07-11T23:22:21.760000",
              "content": "<p>pretrain on split, finetune on full</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2914524,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-10T02:20:03.587000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2914294": "I was holding 1st place for a week and dropped to 4th in the end and then to 988th in private LB, so this is a story full of sadness.\n\n**Baseline **\nAs everyone else for the first month I played with ECFP and other fingerprints, XGboost et al. Then I concluded that it will not generalize well. After some poking around I found a solution that quickly took took me up - GNNs.\n\nThe trick was to generate a training set with rich atom / bond properties annotation - charges, molecular weights, bond types, aromaticity, etc - basically everything you can squeeze from RDKit without heavy computation. Pack all this into pre-processed and pre-splitted CV folds and store on SSD. Took like 12 hours per split to generate and 1TB of storage, but made training GNNs quite fast on a single GPU that I used and model size was like 800kb.\n\n**Shared BB**\n\nBest result on shared BB was 0.344 for a single GAT based GNN (3 class classification). It was training on a CV split and then finetuned on full set. I just used one split, seems pretraining on others was giving worse results. It was a \"lucky\" fold. Seems not so lucky after all.\n\n** non-shared BB*\n\nFor non-shared BB I used a domain adaptation loss function from https://arxiv.org/pdf/2306.04979 paper. I was doing classification loss on full set and domain adaption (basically soft labels generated based on embedding cosines) and this gave me additional public LB boost from 0.498 -> 0.517\n\nI was quite sure that this domain adaptation protocol will save my *** from shake up. ",
    "2946074": "This is very interesting, you may check this [notebook](https://www.kaggle.com/code/kirkdco/precision-by-protein-and-group) or [this one](https://www.kaggle.com/code/lililycai/ap-score-by-protein-and-group) to see which group in the submission you get the highest score. \nWill you share your code?",
    "2914729": "Thanks for your sharing! Did you try use the atom properties only? I have a GNN model using the atom properties only, for the bond I just use the adjacency matrix. This model have only 0.398 in public LB. I'm curious about if it is cause by the loss of bond features. It seems that there are many GNNs with the bond features achieve a very good score on public LB. \n Another thing is that can I ask what kind of strategies do you use to accerate the data IO? I tried to cache the preprocessed features using numpy.memmap or lmdb, but the speed is still very slow and even slower than instant calculation in dataloader of pytorch. For training one epoch (whole dataset) of GNN in A100, it takes 6-8 hours. ",
    "2914304": "Interesting story.  How did you get public LB 0.498 score?  ",
    "2914524": ""
  }
}