{
  "id": 500591,
  "title": "My GNN scores 0.5 CV / 0.42 LB - can I do better? ",
  "url": "/competitions/leash-BELKA/discussion/500591",
  "author_name": "",
  "post_date": "2024-05-06T08:53:09.297688800Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I'm learning about GNNs and first larger scale training resulted in 0.5 CV / 0.42 LB. It's much easier to get a better score with other methods and I wonder if I should keep spending time on GNNs. Are you able to get a better score with this method? I might be making rookie mistakes so my score is probably far from the upper bound. </p>",
  "messages": [
    {
      "id": "2796486",
      "postDate": "05/06/2024 08:53:09",
      "content": "<p>I'm learning about GNNs and first larger scale training resulted in 0.5 CV / 0.42 LB. It's much easier to get a better score with other methods and I wonder if I should keep spending time on GNNs. Are you able to get a better score with this method? I might be making rookie mistakes so my score is probably far from the upper bound. </p>",
      "rawMarkdown": "I'm learning about GNNs and first larger scale training resulted in 0.5 CV / 0.42 LB. It's much easier to get a better score with other methods and I wonder if I should keep spending time on GNNs. Are you able to get a better score with this method? I might be making rookie mistakes so my score is probably far from the upper bound.",
      "votes": null
    },
    {
      "id": "2796511",
      "postDate": "05/06/2024 09:21:28",
      "content": "<p>Your results are somehow consistent with several studies showing GNNs having inferior performance relative to traditional descriptor-based or string-based methods. Personally, I haven't tested GNNs though. </p>\n<p><a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00479-8\" target=\"_blank\">https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00479-8</a></p>\n<p><a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/60c74f590f50db94793973b5\" target=\"_blank\">https://chemrxiv.org/engage/chemrxiv/article-details/60c74f590f50db94793973b5</a></p>",
      "rawMarkdown": "Your results are somehow consistent with several studies showing GNNs having inferior performance relative to traditional descriptor-based or string-based methods. Personally, I haven't tested GNNs though. \n\nhttps://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00479-8\n\nhttps://chemrxiv.org/engage/chemrxiv/article-details/60c74f590f50db94793973b5",
      "votes": null
    },
    {
      "id": "2796628",
      "postDate": "05/06/2024 10:41:10",
      "content": "<p>my suggestion, subsample train data to some small set of bb1 block.<br>\nthen you can make fast comparison of several methods.</p>\n<p>i think at the end it would be ensemble of several methods:<br>\nfast methods trained on all data,  more accurate (slower training) trained on more difficult samples.</p>\n<hr>\n<p>you should start off with the fastest methods first because these are benchmark methods that left you know the dataset and task more. they also let you debug more complex models like gnn, transformer later.</p>\n<hr>\n<p>in the end, methods that uses 3d information should be the best (but most difficult to implement and train) i think</p>",
      "rawMarkdown": "my suggestion, subsample train data to some small set of bb1 block.\nthen you can make fast comparison of several methods.\n\ni think at the end it would be ensemble of several methods:\nfast methods trained on all data,  more accurate (slower training) trained on more difficult samples.\n\n---\n\nyou should start off with the fastest methods first because these are benchmark methods that left you know the dataset and task more. they also let you debug more complex models like gnn, transformer later.\n\n---\n\nin the end, methods that uses 3d information should be the best (but most difficult to implement and train) i think",
      "votes": null
    },
    {
      "id": "2797574",
      "postDate": "05/06/2024 19:27:41",
      "content": "<p>Yes you can. My GCN is my model with the least gap between CV and LB : CV=0.55x, LB= 0.54x. I only used 3M smiles.</p>",
      "rawMarkdown": "Yes you can. My GCN is my model with the least gap between CV and LB : CV=0.55x, LB= 0.54x. I only used 3M smiles.",
      "votes": null
    },
    {
      "id": "2799047",
      "postDate": "05/07/2024 15:14:34",
      "content": "<p>Methods that uses 3d information should be \"the best\"  but just for half of the data in the puzzle. <br>\nBut for the other half (test data without triazine cores)……</p>\n<p><em>One of the goals of this competition is to explore and compare many different ways of representing molecules. Small molecules have been represented with SMILES, graphs, 3D structures, and more, including more esoteric methods such as spherical convolutional neural nets.<strong>We encourage competitors to explore not only different methods of making predictions but also to try different ways of representing the molecules.</strong></em></p>",
      "rawMarkdown": "Methods that uses 3d information should be \"the best\"  but just for half of the data in the puzzle. \nBut for the other half (test data without triazine cores)……\n\n*One of the goals of this competition is to explore and compare many different ways of representing molecules. Small molecules have been represented with SMILES, graphs, 3D structures, and more, including more esoteric methods such as spherical convolutional neural nets.**We encourage competitors to explore not only different methods of making predictions but also to try different ways of representing the molecules.***",
      "votes": null
    },
    {
      "id": "2799109",
      "postDate": "05/07/2024 15:30:10",
      "content": "<p>I believe GCN is slow to train?</p>",
      "rawMarkdown": "I believe GCN is slow to train?",
      "votes": null
    },
    {
      "id": "2799685",
      "postDate": "05/07/2024 22:15:34",
      "content": "<p>yes but</p>\n<ol>\n<li>we do not know if you need simple CGN (eg only 2 to 3 layers) or complicated ones?</li>\n<li>we may not need to apply GCN on all samples (active learning) if it is very strong</li>\n<li>we can always do knowledge distillation (teacher and student learning). note that pytorch 2.0 has acceleration for GCN</li>\n</ol>\n<hr>\n<p>There is a trick. you can implement GCN as transformer. just set the src mask to enable only interaction for k-hop neighbours of graph.</p>",
      "rawMarkdown": "yes but\n1. we do not know if you need simple CGN (eg only 2 to 3 layers) or complicated ones?\n2. we may not need to apply GCN on all samples (active learning) if it is very strong\n3. we can always do knowledge distillation (teacher and student learning). note that pytorch 2.0 has acceleration for GCN\n\n---\nThere is a trick. you can implement GCN as transformer. just set the src mask to enable only interaction for k-hop neighbours of graph.",
      "votes": null
    },
    {
      "id": "2800617",
      "postDate": "05/08/2024 09:17:31",
      "content": "<p>That's true, there are many similar models.</p>",
      "rawMarkdown": "That's true, there are many similar models.",
      "votes": null
    },
    {
      "id": "2802554",
      "postDate": "05/09/2024 05:14:03",
      "content": "<p>Lately, came across <a href=\"https://github.com/masashitsubaki/molecularGNN_smiles\" target=\"_blank\">this</a> repo, it says: <br>\n<strong>Important: this repository will not be further developed and maintained because we have shown and believe that graph neural networks or graph convolutional networks are incorrect and useless for modeling molecules (see our paper in NeurIPS 2020).</strong><br>\nA <a href=\"https://proceedings.neurips.cc/paper/2020/hash/1534b76d325a8f591b52d302e7181331-Abstract.html\" target=\"_blank\">paper</a> published in NeurIPS 2020. This might have your answers. </p>\n<p>From the abstract: <br>\n<strong><em>GCNs involve unnecessary nonlinearity and deep architecture. We also verify that molecular GCNs are based on a poor basis function set compared with the standard one used in theoretical calculations or quantum chemical simulations.</em></strong></p>",
      "rawMarkdown": "Lately, came across [this](https://github.com/masashitsubaki/molecularGNN_smiles) repo, it says: \n**Important: this repository will not be further developed and maintained because we have shown and believe that graph neural networks or graph convolutional networks are incorrect and useless for modeling molecules (see our paper in NeurIPS 2020).**\nA [paper](https://proceedings.neurips.cc/paper/2020/hash/1534b76d325a8f591b52d302e7181331-Abstract.html) published in NeurIPS 2020. This might have your answers. \n\nFrom the abstract: \n***GCNs involve unnecessary nonlinearity and deep architecture. We also verify that molecular GCNs are based on a poor basis function set compared with the standard one used in theoretical calculations or quantum chemical simulations.***",
      "votes": null
    },
    {
      "id": "2805916",
      "postDate": "05/10/2024 19:47:32",
      "content": "<p>I haven't been following GNN's for a long time. GNN's were all the rage a few years ago. Does this apply only to the molecular field or do you think GNN's are facing this problem in other fields?</p>",
      "rawMarkdown": "I haven't been following GNN's for a long time. GNN's were all the rage a few years ago. Does this apply only to the molecular field or do you think GNN's are facing this problem in other fields?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2796511,
      "author_name": "ruelcedeno",
      "author_url": "",
      "post_date": "05/06/2024 09:21:28",
      "content": "<p>Your results are somehow consistent with several studies showing GNNs having inferior performance relative to traditional descriptor-based or string-based methods. Personally, I haven't tested GNNs though. </p>\n<p><a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00479-8\" target=\"_blank\">https://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00479-8</a></p>\n<p><a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/60c74f590f50db94793973b5\" target=\"_blank\">https://chemrxiv.org/engage/chemrxiv/article-details/60c74f590f50db94793973b5</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2796628,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/06/2024 10:41:10",
      "content": "<p>my suggestion, subsample train data to some small set of bb1 block.<br>\nthen you can make fast comparison of several methods.</p>\n<p>i think at the end it would be ensemble of several methods:<br>\nfast methods trained on all data,  more accurate (slower training) trained on more difficult samples.</p>\n<hr>\n<p>you should start off with the fastest methods first because these are benchmark methods that left you know the dataset and task more. they also let you debug more complex models like gnn, transformer later.</p>\n<hr>\n<p>in the end, methods that uses 3d information should be the best (but most difficult to implement and train) i think</p>",
      "votes": null,
      "replies": [
        {
          "id": 2799047,
          "author_name": "ricopue",
          "author_url": "",
          "post_date": "05/07/2024 15:14:34",
          "content": "<p>Methods that uses 3d information should be \"the best\"  but just for half of the data in the puzzle. <br>\nBut for the other half (test data without triazine cores)……</p>\n<p><em>One of the goals of this competition is to explore and compare many different ways of representing molecules. Small molecules have been represented with SMILES, graphs, 3D structures, and more, including more esoteric methods such as spherical convolutional neural nets.<strong>We encourage competitors to explore not only different methods of making predictions but also to try different ways of representing the molecules.</strong></em></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2797574,
      "author_name": "ulrich07",
      "author_url": "",
      "post_date": "05/06/2024 19:27:41",
      "content": "<p>Yes you can. My GCN is my model with the least gap between CV and LB : CV=0.55x, LB= 0.54x. I only used 3M smiles.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2799109,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "05/07/2024 15:30:10",
      "content": "<p>I believe GCN is slow to train?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2799685,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/07/2024 22:15:34",
          "content": "<p>yes but</p>\n<ol>\n<li>we do not know if you need simple CGN (eg only 2 to 3 layers) or complicated ones?</li>\n<li>we may not need to apply GCN on all samples (active learning) if it is very strong</li>\n<li>we can always do knowledge distillation (teacher and student learning). note that pytorch 2.0 has acceleration for GCN</li>\n</ol>\n<hr>\n<p>There is a trick. you can implement GCN as transformer. just set the src mask to enable only interaction for k-hop neighbours of graph.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2800617,
              "author_name": "yuanzhezhou",
              "author_url": "",
              "post_date": "05/08/2024 09:17:31",
              "content": "<p>That's true, there are many similar models.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2802554,
      "author_name": "ahsuna123",
      "author_url": "",
      "post_date": "05/09/2024 05:14:03",
      "content": "<p>Lately, came across <a href=\"https://github.com/masashitsubaki/molecularGNN_smiles\" target=\"_blank\">this</a> repo, it says: <br>\n<strong>Important: this repository will not be further developed and maintained because we have shown and believe that graph neural networks or graph convolutional networks are incorrect and useless for modeling molecules (see our paper in NeurIPS 2020).</strong><br>\nA <a href=\"https://proceedings.neurips.cc/paper/2020/hash/1534b76d325a8f591b52d302e7181331-Abstract.html\" target=\"_blank\">paper</a> published in NeurIPS 2020. This might have your answers. </p>\n<p>From the abstract: <br>\n<strong><em>GCNs involve unnecessary nonlinearity and deep architecture. We also verify that molecular GCNs are based on a poor basis function set compared with the standard one used in theoretical calculations or quantum chemical simulations.</em></strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 2805916,
          "author_name": "verracodeguacas",
          "author_url": "",
          "post_date": "05/10/2024 19:47:32",
          "content": "<p>I haven't been following GNN's for a long time. GNN's were all the rage a few years ago. Does this apply only to the molecular field or do you think GNN's are facing this problem in other fields?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2796486": "I'm learning about GNNs and first larger scale training resulted in 0.5 CV / 0.42 LB. It's much easier to get a better score with other methods and I wonder if I should keep spending time on GNNs. Are you able to get a better score with this method? I might be making rookie mistakes so my score is probably far from the upper bound.",
    "2796511": "Your results are somehow consistent with several studies showing GNNs having inferior performance relative to traditional descriptor-based or string-based methods. Personally, I haven't tested GNNs though. \n\nhttps://jcheminf.biomedcentral.com/articles/10.1186/s13321-020-00479-8\n\nhttps://chemrxiv.org/engage/chemrxiv/article-details/60c74f590f50db94793973b5",
    "2796628": "my suggestion, subsample train data to some small set of bb1 block.\nthen you can make fast comparison of several methods.\n\ni think at the end it would be ensemble of several methods:\nfast methods trained on all data,  more accurate (slower training) trained on more difficult samples.\n\n---\n\nyou should start off with the fastest methods first because these are benchmark methods that left you know the dataset and task more. they also let you debug more complex models like gnn, transformer later.\n\n---\n\nin the end, methods that uses 3d information should be the best (but most difficult to implement and train) i think",
    "2797574": "Yes you can. My GCN is my model with the least gap between CV and LB : CV=0.55x, LB= 0.54x. I only used 3M smiles.",
    "2799047": "Methods that uses 3d information should be \"the best\"  but just for half of the data in the puzzle. \nBut for the other half (test data without triazine cores)……\n\n*One of the goals of this competition is to explore and compare many different ways of representing molecules. Small molecules have been represented with SMILES, graphs, 3D structures, and more, including more esoteric methods such as spherical convolutional neural nets.**We encourage competitors to explore not only different methods of making predictions but also to try different ways of representing the molecules.***",
    "2799109": "I believe GCN is slow to train?",
    "2799685": "yes but\n1. we do not know if you need simple CGN (eg only 2 to 3 layers) or complicated ones?\n2. we may not need to apply GCN on all samples (active learning) if it is very strong\n3. we can always do knowledge distillation (teacher and student learning). note that pytorch 2.0 has acceleration for GCN\n\n---\nThere is a trick. you can implement GCN as transformer. just set the src mask to enable only interaction for k-hop neighbours of graph.",
    "2800617": "That's true, there are many similar models.",
    "2802554": "Lately, came across [this](https://github.com/masashitsubaki/molecularGNN_smiles) repo, it says: \n**Important: this repository will not be further developed and maintained because we have shown and believe that graph neural networks or graph convolutional networks are incorrect and useless for modeling molecules (see our paper in NeurIPS 2020).**\nA [paper](https://proceedings.neurips.cc/paper/2020/hash/1534b76d325a8f591b52d302e7181331-Abstract.html) published in NeurIPS 2020. This might have your answers. \n\nFrom the abstract: \n***GCNs involve unnecessary nonlinearity and deep architecture. We also verify that molecular GCNs are based on a poor basis function set compared with the standard one used in theoretical calculations or quantum chemical simulations.***",
    "2805916": "I haven't been following GNN's for a long time. GNN's were all the rage a few years ago. Does this apply only to the molecular field or do you think GNN's are facing this problem in other fields?"
  },
  "source": "meta"
}