{
  "id": 461015,
  "title": "9th solution write-up: Pure NN model",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/northern-star-9th-solution-write-up-pure-nn-model",
  "author_name": "",
  "post_date": "2023-12-12T08:44:11.117Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>9th solution: pure NN model</h1>\n<p>Hi everyone! I am Dave, this is my first time completing a kaggle competition, and I feel honored to win a gold. I want to use a pure NN model to solve this problem. I would be happy if you find this solution interesting and helpful.</p>\n<h2>Problem definition</h2>\n<p>I have seen many good solutions using each row of the data as one sample, however, I view this problem in a different way. I extracted (cell,sm,gene,value) pairs from the dataset and for each time, the model will predict the object value for a given cell type, sm type, and gene type.</p>\n<h2>The architecture of Model</h2>\n<p>The overview of the model is shown as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F244e8db36e50bc5eeb7ae04d1b4657b1%2F2023-12-12%204.07.55.png?generation=1702368531932768&amp;alt=media\" alt=\"\"><br>\nYou can see that I used three kinds of features: sm features, gene features and cell features. Let’s first check how to get the sm features:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F0c1cf4496cea2ad183a03910b46793ce%2F2023-12-12%202.30.23.png?generation=1702365244366440&amp;alt=media\" alt=\"\"><br>\nHere I involved many different kinds of features:</p>\n<ul>\n<li>MACC: Molecular ACCess System keys, are one of the most commonly used structural keys, check details <a href=\"https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors\" target=\"_blank\">here</a>.</li>\n<li>ECFP: extended-connectivity fingerprints, are generated using a variant of the Morgan algorithm, check details <a href=\"https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors\" target=\"_blank\">here</a>.</li>\n<li>WHIM: Weighted Holistic Invariant Molecular descriptors, are geometrical descriptors based on statistical indices calculated on the projections of the atoms along principal axes, check details <a href=\"https://chemgps.bmc.uu.se/help/dragonx/WHIMdecriptors1.html\" target=\"_blank\">here</a>.</li>\n<li>sm type: the types of small molecules.</li>\n<li>sm hba: the number of H-bond acceptors for a molecule.</li>\n<li>sm hbd: the number of H-bond donors for a molecule.</li>\n<li>sm rotb: the number of rotatable bonds for a molecule.</li>\n<li>sm mw: the molecular weight for a molecule.</li>\n<li>sm psa: the Polar surface area for a molecule.</li>\n<li>sm logp: the log of the partition coefficient of a solute between octanol and water.</li>\n</ul>\n<p>Now let's check how to get the gene features:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Ffcadf4701630dbb7774bc4d4616ba1c7%2F2023-12-12%202.46.18.png?generation=1702365793278185&amp;alt=media\" alt=\"\"><br>\nI have involved the PCA of the genes and the additional features. To calculate the PCA feature, I set the n_components as 10 so that I can get a 10-dimensional vector to represent each gene.</p>\n<p>Then let's check how to get the cell features:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Fd0c22eba36b90129dd4ff4053bd52e88%2F2023-12-12%202.50.06.png?generation=1702365903775022&amp;alt=media\" alt=\"\"><br>\nSomething new here is that I used a GCN layer to help extract the feature contained in the relationship between different cells. The graph used here is very simple: cells are denoted as nodes, so there are only 6 nodes in this graph. Each pair of two nodes has an un-directed edge. I found it helpful in LB.<br>\nWhat's more, given this problem is unbalanced, we need to improve the generalization on different cell types, so you may notice that I have involved random noise in the cell features.</p>\n<p>Another important part is the attention layers, the structure is shown as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F7ab9b5fff9e6588b8d3b3d17020a87e0%2F2023-12-12%203.43.25.png?generation=1702367027176571&amp;alt=media\" alt=\"\"></p>\n<p>Now we know the structure of the model, but before directly training them, to improve the performance, I built a larger model based on 3 models with the same structure mentioned earlier:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F9269c6e48312dcde4897c6a95fe4c783%2F2023-12-12%204.17.34.png?generation=1702369066405973&amp;alt=media\" alt=\"\"></p>\n<h2>Others</h2>\n<p>I used MSE as the loss function, AdamW as the optimizer, the learning rate is 3e-4, weight decay is 1e-3, and the batch size is 128.</p>\n<ul>\n<li>loss function: MSE</li>\n<li>optimizer: AdamW</li>\n<li>learning rate: 3e-4</li>\n<li>weight decay: 1e-3</li>\n<li>batch size: 128</li>\n<li>CV: 5-fold cross-validation</li>\n</ul>",
  "messages": [
    {
      "id": "2558578",
      "postDate": "12/12/2023 08:31:10",
      "content": "<h1>9th solution: pure NN model</h1>\n<p>Hi everyone! I am Dave, this is my first time completing a kaggle competition, and I feel honored to win a gold. I want to use a pure NN model to solve this problem. I would be happy if you find this solution interesting and helpful.</p>\n<h2>Problem definition</h2>\n<p>I have seen many good solutions using each row of the data as one sample, however, I view this problem in a different way. I extracted (cell,sm,gene,value) pairs from the dataset and for each time, the model will predict the object value for a given cell type, sm type, and gene type.</p>\n<h2>The architecture of Model</h2>\n<p>The overview of the model is shown as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F244e8db36e50bc5eeb7ae04d1b4657b1%2F2023-12-12%204.07.55.png?generation=1702368531932768&amp;alt=media\" alt=\"\"><br>\nYou can see that I used three kinds of features: sm features, gene features and cell features. Let’s first check how to get the sm features:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F0c1cf4496cea2ad183a03910b46793ce%2F2023-12-12%202.30.23.png?generation=1702365244366440&amp;alt=media\" alt=\"\"><br>\nHere I involved many different kinds of features:</p>\n<ul>\n<li>MACC: Molecular ACCess System keys, are one of the most commonly used structural keys, check details <a href=\"https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors\" target=\"_blank\">here</a>.</li>\n<li>ECFP: extended-connectivity fingerprints, are generated using a variant of the Morgan algorithm, check details <a href=\"https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors\" target=\"_blank\">here</a>.</li>\n<li>WHIM: Weighted Holistic Invariant Molecular descriptors, are geometrical descriptors based on statistical indices calculated on the projections of the atoms along principal axes, check details <a href=\"https://chemgps.bmc.uu.se/help/dragonx/WHIMdecriptors1.html\" target=\"_blank\">here</a>.</li>\n<li>sm type: the types of small molecules.</li>\n<li>sm hba: the number of H-bond acceptors for a molecule.</li>\n<li>sm hbd: the number of H-bond donors for a molecule.</li>\n<li>sm rotb: the number of rotatable bonds for a molecule.</li>\n<li>sm mw: the molecular weight for a molecule.</li>\n<li>sm psa: the Polar surface area for a molecule.</li>\n<li>sm logp: the log of the partition coefficient of a solute between octanol and water.</li>\n</ul>\n<p>Now let's check how to get the gene features:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Ffcadf4701630dbb7774bc4d4616ba1c7%2F2023-12-12%202.46.18.png?generation=1702365793278185&amp;alt=media\" alt=\"\"><br>\nI have involved the PCA of the genes and the additional features. To calculate the PCA feature, I set the n_components as 10 so that I can get a 10-dimensional vector to represent each gene.</p>\n<p>Then let's check how to get the cell features:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Fd0c22eba36b90129dd4ff4053bd52e88%2F2023-12-12%202.50.06.png?generation=1702365903775022&amp;alt=media\" alt=\"\"><br>\nSomething new here is that I used a GCN layer to help extract the feature contained in the relationship between different cells. The graph used here is very simple: cells are denoted as nodes, so there are only 6 nodes in this graph. Each pair of two nodes has an un-directed edge. I found it helpful in LB.<br>\nWhat's more, given this problem is unbalanced, we need to improve the generalization on different cell types, so you may notice that I have involved random noise in the cell features.</p>\n<p>Another important part is the attention layers, the structure is shown as follows:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F7ab9b5fff9e6588b8d3b3d17020a87e0%2F2023-12-12%203.43.25.png?generation=1702367027176571&amp;alt=media\" alt=\"\"></p>\n<p>Now we know the structure of the model, but before directly training them, to improve the performance, I built a larger model based on 3 models with the same structure mentioned earlier:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F9269c6e48312dcde4897c6a95fe4c783%2F2023-12-12%204.17.34.png?generation=1702369066405973&amp;alt=media\" alt=\"\"></p>\n<h2>Others</h2>\n<p>I used MSE as the loss function, AdamW as the optimizer, the learning rate is 3e-4, weight decay is 1e-3, and the batch size is 128.</p>\n<ul>\n<li>loss function: MSE</li>\n<li>optimizer: AdamW</li>\n<li>learning rate: 3e-4</li>\n<li>weight decay: 1e-3</li>\n<li>batch size: 128</li>\n<li>CV: 5-fold cross-validation</li>\n</ul>",
      "rawMarkdown": "# 9th solution: pure NN model\n\nHi everyone! I am Dave, this is my first time completing a kaggle competition, and I feel honored to win a gold. I want to use a pure NN model to solve this problem. I would be happy if you find this solution interesting and helpful.\n\n## Problem definition\nI have seen many good solutions using each row of the data as one sample, however, I view this problem in a different way. I extracted (cell,sm,gene,value) pairs from the dataset and for each time, the model will predict the object value for a given cell type, sm type, and gene type.\n\n## The architecture of Model\nThe overview of the model is shown as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F244e8db36e50bc5eeb7ae04d1b4657b1%2F2023-12-12%204.07.55.png?generation=1702368531932768&alt=media)\nYou can see that I used three kinds of features: sm features, gene features and cell features. Let’s first check how to get the sm features:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F0c1cf4496cea2ad183a03910b46793ce%2F2023-12-12%202.30.23.png?generation=1702365244366440&alt=media)\nHere I involved many different kinds of features:\n- MACC: Molecular ACCess System keys, are one of the most commonly used structural keys, check details [here](https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors).\n- ECFP: extended-connectivity fingerprints, are generated using a variant of the Morgan algorithm, check details [here](https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors).\n- WHIM: Weighted Holistic Invariant Molecular descriptors, are geometrical descriptors based on statistical indices calculated on the projections of the atoms along principal axes, check details [here](https://chemgps.bmc.uu.se/help/dragonx/WHIMdecriptors1.html).\n- sm type: the types of small molecules.\n- sm hba: the number of H-bond acceptors for a molecule.\n- sm hbd: the number of H-bond donors for a molecule.\n- sm rotb: the number of rotatable bonds for a molecule.\n- sm mw: the molecular weight for a molecule.\n- sm psa: the Polar surface area for a molecule.\n- sm logp: the log of the partition coefficient of a solute between octanol and water.\n\nNow let's check how to get the gene features:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Ffcadf4701630dbb7774bc4d4616ba1c7%2F2023-12-12%202.46.18.png?generation=1702365793278185&alt=media)\nI have involved the PCA of the genes and the additional features. To calculate the PCA feature, I set the n_components as 10 so that I can get a 10-dimensional vector to represent each gene.\n\nThen let's check how to get the cell features:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Fd0c22eba36b90129dd4ff4053bd52e88%2F2023-12-12%202.50.06.png?generation=1702365903775022&alt=media)\nSomething new here is that I used a GCN layer to help extract the feature contained in the relationship between different cells. The graph used here is very simple: cells are denoted as nodes, so there are only 6 nodes in this graph. Each pair of two nodes has an un-directed edge. I found it helpful in LB.\nWhat's more, given this problem is unbalanced, we need to improve the generalization on different cell types, so you may notice that I have involved random noise in the cell features.\n\nAnother important part is the attention layers, the structure is shown as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F7ab9b5fff9e6588b8d3b3d17020a87e0%2F2023-12-12%203.43.25.png?generation=1702367027176571&alt=media)\n\nNow we know the structure of the model, but before directly training them, to improve the performance, I built a larger model based on 3 models with the same structure mentioned earlier:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F9269c6e48312dcde4897c6a95fe4c783%2F2023-12-12%204.17.34.png?generation=1702369066405973&alt=media)\n\n## Others\nI used MSE as the loss function, AdamW as the optimizer, the learning rate is 3e-4, weight decay is 1e-3, and the batch size is 128.\n- loss function: MSE\n- optimizer: AdamW\n- learning rate: 3e-4\n- weight decay: 1e-3\n- batch size: 128\n- CV: 5-fold cross-validation",
      "votes": null
    },
    {
      "id": "2562197",
      "postDate": "12/15/2023 07:25:03",
      "content": "<p>Very interesting, for linear mixture layer of the 3 models did you constrain the weights to sum to 1 or were they unconstrained?</p>",
      "rawMarkdown": "Very interesting, for linear mixture layer of the 3 models did you constrain the weights to sum to 1 or were they unconstrained?",
      "votes": null
    },
    {
      "id": "2562367",
      "postDate": "12/15/2023 09:36:29",
      "content": "<p>Thank you very much! In this competition, I found that unconstrainted weight and without bias for linear mixture layer can help me to improve the score, and the score will drop if I use softmax to constraint the weights, that's interesting😆</p>",
      "rawMarkdown": "Thank you very much! In this competition, I found that unconstrainted weight and without bias for linear mixture layer can help me to improve the score, and the score will drop if I use softmax to constraint the weights, that's interesting😆",
      "votes": null
    },
    {
      "id": "2563045",
      "postDate": "12/16/2023 01:06:49",
      "content": "<p>Thank you for sharing this detailed description of your model and congratulations on making it into the top ten! </p>\n<p>I also converted the data into long format with <code>cell_type</code>, <code>sm_name</code>, and <code>gene</code> as features and a single target <code>value</code> per row. Yours is the first solution I read that did that, too. </p>",
      "rawMarkdown": "Thank you for sharing this detailed description of your model and congratulations on making it into the top ten! \n\nI also converted the data into long format with ```cell_type```, ```sm_name```, and ```gene``` as features and a single target ```value``` per row. Yours is the first solution I read that did that, too.",
      "votes": null
    },
    {
      "id": "2563129",
      "postDate": "12/16/2023 03:42:04",
      "content": "<p>Thank you! I'm glad that we share the same view on this problem, I think this way has more freedom, do you think so? Actually, by checking the private score of the history submissions, I found I could get the prize with a score of 0.732, but I didn't choose that submission… </p>",
      "rawMarkdown": "Thank you! I'm glad that we share the same view on this problem, I think this way has more freedom, do you think so? Actually, by checking the private score of the history submissions, I found I could get the prize with a score of 0.732, but I didn't choose that submission...",
      "votes": null
    },
    {
      "id": "2566184",
      "postDate": "12/18/2023 14:40:15",
      "content": "<p>I'm not sure about the superiority of our approach. Probably, an ensemble using both approaches would be best. If you're interested in my solution, <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461649\" target=\"_blank\">here</a> is the write-up.</p>",
      "rawMarkdown": "I'm not sure about the superiority of our approach. Probably, an ensemble using both approaches would be best. If you're interested in my solution, [here](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461649) is the write-up.",
      "votes": null
    },
    {
      "id": "2696470",
      "postDate": "03/14/2024 10:34:03",
      "content": "<p>Thanks for sharing and congratulations with gold medal ! <br>\nWould it be possible to share code ? Kaggle notebooks (prefarbaly) or github ? </p>",
      "rawMarkdown": "Thanks for sharing and congratulations with gold medal ! \nWould it be possible to share code ? Kaggle notebooks (prefarbaly) or github ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2562197,
      "author_name": "maxleverage",
      "author_url": "",
      "post_date": "12/15/2023 07:25:03",
      "content": "<p>Very interesting, for linear mixture layer of the 3 models did you constrain the weights to sum to 1 or were they unconstrained?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2562367,
          "author_name": "davezzq",
          "author_url": "",
          "post_date": "12/15/2023 09:36:29",
          "content": "<p>Thank you very much! In this competition, I found that unconstrainted weight and without bias for linear mixture layer can help me to improve the score, and the score will drop if I use softmax to constraint the weights, that's interesting😆</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2563045,
      "author_name": "frenio",
      "author_url": "",
      "post_date": "12/16/2023 01:06:49",
      "content": "<p>Thank you for sharing this detailed description of your model and congratulations on making it into the top ten! </p>\n<p>I also converted the data into long format with <code>cell_type</code>, <code>sm_name</code>, and <code>gene</code> as features and a single target <code>value</code> per row. Yours is the first solution I read that did that, too. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2563129,
          "author_name": "davezzq",
          "author_url": "",
          "post_date": "12/16/2023 03:42:04",
          "content": "<p>Thank you! I'm glad that we share the same view on this problem, I think this way has more freedom, do you think so? Actually, by checking the private score of the history submissions, I found I could get the prize with a score of 0.732, but I didn't choose that submission… </p>",
          "votes": null,
          "replies": [
            {
              "id": 2566184,
              "author_name": "frenio",
              "author_url": "",
              "post_date": "12/18/2023 14:40:15",
              "content": "<p>I'm not sure about the superiority of our approach. Probably, an ensemble using both approaches would be best. If you're interested in my solution, <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461649\" target=\"_blank\">here</a> is the write-up.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2696470,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "03/14/2024 10:34:03",
      "content": "<p>Thanks for sharing and congratulations with gold medal ! <br>\nWould it be possible to share code ? Kaggle notebooks (prefarbaly) or github ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2558578": "# 9th solution: pure NN model\n\nHi everyone! I am Dave, this is my first time completing a kaggle competition, and I feel honored to win a gold. I want to use a pure NN model to solve this problem. I would be happy if you find this solution interesting and helpful.\n\n## Problem definition\nI have seen many good solutions using each row of the data as one sample, however, I view this problem in a different way. I extracted (cell,sm,gene,value) pairs from the dataset and for each time, the model will predict the object value for a given cell type, sm type, and gene type.\n\n## The architecture of Model\nThe overview of the model is shown as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F244e8db36e50bc5eeb7ae04d1b4657b1%2F2023-12-12%204.07.55.png?generation=1702368531932768&alt=media)\nYou can see that I used three kinds of features: sm features, gene features and cell features. Let’s first check how to get the sm features:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F0c1cf4496cea2ad183a03910b46793ce%2F2023-12-12%202.30.23.png?generation=1702365244366440&alt=media)\nHere I involved many different kinds of features:\n- MACC: Molecular ACCess System keys, are one of the most commonly used structural keys, check details [here](https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors).\n- ECFP: extended-connectivity fingerprints, are generated using a variant of the Morgan algorithm, check details [here](https://chem.libretexts.org/Courses/Intercollegiate_Courses/Cheminformatics/06%3A_Molecular_Similarity/6.01%3A_Molecular_Descriptors).\n- WHIM: Weighted Holistic Invariant Molecular descriptors, are geometrical descriptors based on statistical indices calculated on the projections of the atoms along principal axes, check details [here](https://chemgps.bmc.uu.se/help/dragonx/WHIMdecriptors1.html).\n- sm type: the types of small molecules.\n- sm hba: the number of H-bond acceptors for a molecule.\n- sm hbd: the number of H-bond donors for a molecule.\n- sm rotb: the number of rotatable bonds for a molecule.\n- sm mw: the molecular weight for a molecule.\n- sm psa: the Polar surface area for a molecule.\n- sm logp: the log of the partition coefficient of a solute between octanol and water.\n\nNow let's check how to get the gene features:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Ffcadf4701630dbb7774bc4d4616ba1c7%2F2023-12-12%202.46.18.png?generation=1702365793278185&alt=media)\nI have involved the PCA of the genes and the additional features. To calculate the PCA feature, I set the n_components as 10 so that I can get a 10-dimensional vector to represent each gene.\n\nThen let's check how to get the cell features:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2Fd0c22eba36b90129dd4ff4053bd52e88%2F2023-12-12%202.50.06.png?generation=1702365903775022&alt=media)\nSomething new here is that I used a GCN layer to help extract the feature contained in the relationship between different cells. The graph used here is very simple: cells are denoted as nodes, so there are only 6 nodes in this graph. Each pair of two nodes has an un-directed edge. I found it helpful in LB.\nWhat's more, given this problem is unbalanced, we need to improve the generalization on different cell types, so you may notice that I have involved random noise in the cell features.\n\nAnother important part is the attention layers, the structure is shown as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F7ab9b5fff9e6588b8d3b3d17020a87e0%2F2023-12-12%203.43.25.png?generation=1702367027176571&alt=media)\n\nNow we know the structure of the model, but before directly training them, to improve the performance, I built a larger model based on 3 models with the same structure mentioned earlier:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4090067%2F9269c6e48312dcde4897c6a95fe4c783%2F2023-12-12%204.17.34.png?generation=1702369066405973&alt=media)\n\n## Others\nI used MSE as the loss function, AdamW as the optimizer, the learning rate is 3e-4, weight decay is 1e-3, and the batch size is 128.\n- loss function: MSE\n- optimizer: AdamW\n- learning rate: 3e-4\n- weight decay: 1e-3\n- batch size: 128\n- CV: 5-fold cross-validation",
    "2562197": "Very interesting, for linear mixture layer of the 3 models did you constrain the weights to sum to 1 or were they unconstrained?",
    "2562367": "Thank you very much! In this competition, I found that unconstrainted weight and without bias for linear mixture layer can help me to improve the score, and the score will drop if I use softmax to constraint the weights, that's interesting😆",
    "2563045": "Thank you for sharing this detailed description of your model and congratulations on making it into the top ten! \n\nI also converted the data into long format with ```cell_type```, ```sm_name```, and ```gene``` as features and a single target ```value``` per row. Yours is the first solution I read that did that, too.",
    "2563129": "Thank you! I'm glad that we share the same view on this problem, I think this way has more freedom, do you think so? Actually, by checking the private score of the history submissions, I found I could get the prize with a score of 0.732, but I didn't choose that submission...",
    "2566184": "I'm not sure about the superiority of our approach. Probably, an ensemble using both approaches would be best. If you're interested in my solution, [here](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/461649) is the write-up.",
    "2696470": "Thanks for sharing and congratulations with gold medal ! \nWould it be possible to share code ? Kaggle notebooks (prefarbaly) or github ?"
  },
  "source": "meta"
}