{
  "id": 459961,
  "title": "10th Place Solution for the Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/nm-r-nm-10th-place-solution-for-the-open-problems-",
  "author_name": "",
  "post_date": "2023-12-14T13:53:06.450Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to Kaggle for organizing such an interesting competition.<br>\nThanks to the teammates who fought side by side. And other Kagglers who share various ideas.</p>\n<h1>Context</h1>\n<p>•    <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Competition Overview</a></p>\n<p>•    <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Competition Data</a></p>\n<h1>Overview of the approach</h1>\n<p>Overall, our final submitted result is an ensemble of two parts. </p>\n<p><strong>Final submission = Part A×0.7 + Part B×0.3</strong></p>\n<p>We will explain the composition of Part A and Part B respectively.</p>\n<h1>Part A</h1>\n<p>It is an ensemble composed of neural networks with different structures.</p>\n<h3>Feature engineering</h3>\n<p>After many attempts, we finally adopted the following two features from the public notebook <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-feature-engineering\" target=\"_blank\">\"OP2: feature engineering\"</a> as our training and testing feature.</p>\n<p><em>(1). PCA followed by target encoding (cell type + drug) without noise (pca_target_encoded_features)</em>     <br>\n<em>(2). PCA followed by target encoding (cell type + drug) with noise (pca_target_encoded_features_s0.1)</em></p>\n<p>(1) is subjected to PCA on 18,211 target variables and produced features representing cell type means for drugs and cell type means for compounds. And by using features (2) that add random noise, we believe this will make the model more generalizable.</p>\n<h3>Models</h3>\n<ul>\n<li>NN with Fully Connected Layers, as well as BatchNormalization, Dropout, ReLU, and Linear Activation Functions.</li>\n<li>The initial seed is set to 42 and is fixed. The loss function is mae, the optimizer is Adam.</li>\n<li>The structure of NN is shown below.</li>\n</ul>\n<pre><code>tf.random.set_seed()\n\nmodel = Sequential([ \n    Dense(),\n    BatchNormalization(),\n    Activation(),\n    Dropout(),\n    Dense(),\n    BatchNormalization(),\n    Activation(),\n    Dropout(),\n    Dense(, activation=),\n    Dropout(),\n    Dense(, activation=),\n    BatchNormalization(),\n    Dropout(),\n    Dense(, activation=),\n    Dropout(),\n    Dense(,activation= )\n])\n\nmodel.(loss=, \n                optimizer=tf.keras.optimizers.Adam(),\n                metrics=[custom_mean_rowwise_rmse])\n\nhistory = model.fit(full_features, labels, epochs=, verbose=)\n</code></pre>\n<p>A simple Part A model training process is shown in this notebook <a href=\"https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training\" target=\"_blank\">\"Single-Cell Perturbations(Part A-Model Training)\"</a> .</p>\n<h3>Model ensemble</h3>\n<p>We use feature (1) and feature (2) to train the network respectively. By changing the feature, the number of network layers, the number of nodes, and the number of training epochs, we successfully obtained a set of individual models scoring 0.567-0.582 on LB.  <br>\nSubsequently, we ensembled 7 models and use it as <strong>Part A</strong> (LB: 0.556/PB 0.741).</p>\n<p>In fact, if the weight of each model is determined based on CV during ensemble, we find that the above results can be further optimized to LB0.556/PB0.74.</p>\n<p>Moreover, the highest score we obtained was (LB: 0.557/PB 0.737) after combining 8 models through this method.</p>\n<h1>Part B</h1>\n<p>This part is also composed of NN and an ensemble of different models.</p>\n<h3>Feature engineering</h3>\n<p>The following two approaches are used to perform feature engineering.</p>\n<p>・One-hot encoding on cell_type and sm_name  <br>\n・SMILES(ChemBERTa-77M-MLM)</p>\n<h3>Models</h3>\n<pre><code> (nn.Module):\n     ():\n        (DnnV5, self).__init__() \n        self.conv1d1 = nn.Conv1d(\n            in_channels=,\n            out_channels=,\n            kernel_size=,\n            stride=,\n            padding=,\n            bias=)\n        self.batch_norm1 = nn.BatchNorm1d()\n        self.dense1 = nn.utils.weight_norm(nn.Linear(, ))\n\n        self.batch_norm2 = nn.BatchNorm1d()\n        self.dropout2 = nn.Dropout()\n        self.dense2 = nn.utils.weight_norm(nn.Linear(, ))\n\n\n        self.batch_norm4 = nn.BatchNorm1d()\n        self.dropout4 = nn.Dropout()\n\n        self.dense4 = nn.utils.weight_norm(nn.Linear(, num_targets))\n\n     ():\n        b,w = x.shape\n        x = x.reshape(b,w,)\n        x = self.conv1d1(x)\n        x = x.reshape(b,)\n        x = self.batch_norm1(x)\n        x = F.leaky_relu(self.dense1(x))\n\n        x = self.batch_norm2(x)\n        x = self.dropout2(x)\n        x = F.leaky_relu(self.dense2(x))\n\n\n        x = self.batch_norm4(x)\n        x = self.dropout4(x)\n        y = self.dense4(x)\n         y\n</code></pre>\n<p>The highest score on LB when using model = DnnV5() for each fold is 0.579.  <br>\nFor an ensemble of 5 fold + 3 seed, the best LB 0.568 for the above model structure can be obtained.</p>\n<p>Additionally, <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">\"SCP Quickstart\"</a>  was referred to for fold creation.</p>\n<h3>Model ensemble</h3>\n<p>Multiple models were created with different smiles and model structures, and the final ensemble was created with the following ratio.</p>\n<pre><code>ensemble_submission = * (sub0568* +sub0570*+ sub0571*+ sub0571_2*+ sub0573*+sub0576_1* + sub0576_2*)+ lbsub567*\n</code></pre>\n<p>Among them, lbsub567 referred to the public notebook <a href=\"https://www.kaggle.com/code/misakimatsutomo/blend-for-single-cell-perturbations\" target=\"_blank\">\"Blend for Single-Cell Perturbations\"</a>.   <br>\nWith the above model ensemble, we get a score of 0.560 on LB. We use this ensemble as <strong>part B</strong>.</p>\n<p>Based on the above process, our final submission is <strong>Part A×0.7 + Part B×0.3</strong>,     and <strong>(LB: 0.554/PB 0.741)</strong> is obtained.</p>\n<h1>Things that didn't work</h1>\n<p>•    Pseudo labels</p>\n<p>•    Feature normalization</p>",
  "messages": [
    {
      "id": "2552135",
      "postDate": "12/07/2023 08:21:42",
      "content": "<p>Thanks to Kaggle for organizing such an interesting competition.<br>\nThanks to the teammates who fought side by side. And other Kagglers who share various ideas.</p>\n<h1>Context</h1>\n<p>•    <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Competition Overview</a></p>\n<p>•    <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Competition Data</a></p>\n<h1>Overview of the approach</h1>\n<p>Overall, our final submitted result is an ensemble of two parts. </p>\n<p><strong>Final submission = Part A×0.7 + Part B×0.3</strong></p>\n<p>We will explain the composition of Part A and Part B respectively.</p>\n<h1>Part A</h1>\n<p>It is an ensemble composed of neural networks with different structures.</p>\n<h3>Feature engineering</h3>\n<p>After many attempts, we finally adopted the following two features from the public notebook <a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-feature-engineering\" target=\"_blank\">\"OP2: feature engineering\"</a> as our training and testing feature.</p>\n<p><em>(1). PCA followed by target encoding (cell type + drug) without noise (pca_target_encoded_features)</em>     <br>\n<em>(2). PCA followed by target encoding (cell type + drug) with noise (pca_target_encoded_features_s0.1)</em></p>\n<p>(1) is subjected to PCA on 18,211 target variables and produced features representing cell type means for drugs and cell type means for compounds. And by using features (2) that add random noise, we believe this will make the model more generalizable.</p>\n<h3>Models</h3>\n<ul>\n<li>NN with Fully Connected Layers, as well as BatchNormalization, Dropout, ReLU, and Linear Activation Functions.</li>\n<li>The initial seed is set to 42 and is fixed. The loss function is mae, the optimizer is Adam.</li>\n<li>The structure of NN is shown below.</li>\n</ul>\n<pre><code>tf.random.set_seed()\n\nmodel = Sequential([ \n    Dense(),\n    BatchNormalization(),\n    Activation(),\n    Dropout(),\n    Dense(),\n    BatchNormalization(),\n    Activation(),\n    Dropout(),\n    Dense(, activation=),\n    Dropout(),\n    Dense(, activation=),\n    BatchNormalization(),\n    Dropout(),\n    Dense(, activation=),\n    Dropout(),\n    Dense(,activation= )\n])\n\nmodel.(loss=, \n                optimizer=tf.keras.optimizers.Adam(),\n                metrics=[custom_mean_rowwise_rmse])\n\nhistory = model.fit(full_features, labels, epochs=, verbose=)\n</code></pre>\n<p>A simple Part A model training process is shown in this notebook <a href=\"https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training\" target=\"_blank\">\"Single-Cell Perturbations(Part A-Model Training)\"</a> .</p>\n<h3>Model ensemble</h3>\n<p>We use feature (1) and feature (2) to train the network respectively. By changing the feature, the number of network layers, the number of nodes, and the number of training epochs, we successfully obtained a set of individual models scoring 0.567-0.582 on LB.  <br>\nSubsequently, we ensembled 7 models and use it as <strong>Part A</strong> (LB: 0.556/PB 0.741).</p>\n<p>In fact, if the weight of each model is determined based on CV during ensemble, we find that the above results can be further optimized to LB0.556/PB0.74.</p>\n<p>Moreover, the highest score we obtained was (LB: 0.557/PB 0.737) after combining 8 models through this method.</p>\n<h1>Part B</h1>\n<p>This part is also composed of NN and an ensemble of different models.</p>\n<h3>Feature engineering</h3>\n<p>The following two approaches are used to perform feature engineering.</p>\n<p>・One-hot encoding on cell_type and sm_name  <br>\n・SMILES(ChemBERTa-77M-MLM)</p>\n<h3>Models</h3>\n<pre><code> (nn.Module):\n     ():\n        (DnnV5, self).__init__() \n        self.conv1d1 = nn.Conv1d(\n            in_channels=,\n            out_channels=,\n            kernel_size=,\n            stride=,\n            padding=,\n            bias=)\n        self.batch_norm1 = nn.BatchNorm1d()\n        self.dense1 = nn.utils.weight_norm(nn.Linear(, ))\n\n        self.batch_norm2 = nn.BatchNorm1d()\n        self.dropout2 = nn.Dropout()\n        self.dense2 = nn.utils.weight_norm(nn.Linear(, ))\n\n\n        self.batch_norm4 = nn.BatchNorm1d()\n        self.dropout4 = nn.Dropout()\n\n        self.dense4 = nn.utils.weight_norm(nn.Linear(, num_targets))\n\n     ():\n        b,w = x.shape\n        x = x.reshape(b,w,)\n        x = self.conv1d1(x)\n        x = x.reshape(b,)\n        x = self.batch_norm1(x)\n        x = F.leaky_relu(self.dense1(x))\n\n        x = self.batch_norm2(x)\n        x = self.dropout2(x)\n        x = F.leaky_relu(self.dense2(x))\n\n\n        x = self.batch_norm4(x)\n        x = self.dropout4(x)\n        y = self.dense4(x)\n         y\n</code></pre>\n<p>The highest score on LB when using model = DnnV5() for each fold is 0.579.  <br>\nFor an ensemble of 5 fold + 3 seed, the best LB 0.568 for the above model structure can be obtained.</p>\n<p>Additionally, <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart\" target=\"_blank\">\"SCP Quickstart\"</a>  was referred to for fold creation.</p>\n<h3>Model ensemble</h3>\n<p>Multiple models were created with different smiles and model structures, and the final ensemble was created with the following ratio.</p>\n<pre><code>ensemble_submission = * (sub0568* +sub0570*+ sub0571*+ sub0571_2*+ sub0573*+sub0576_1* + sub0576_2*)+ lbsub567*\n</code></pre>\n<p>Among them, lbsub567 referred to the public notebook <a href=\"https://www.kaggle.com/code/misakimatsutomo/blend-for-single-cell-perturbations\" target=\"_blank\">\"Blend for Single-Cell Perturbations\"</a>.   <br>\nWith the above model ensemble, we get a score of 0.560 on LB. We use this ensemble as <strong>part B</strong>.</p>\n<p>Based on the above process, our final submission is <strong>Part A×0.7 + Part B×0.3</strong>,     and <strong>(LB: 0.554/PB 0.741)</strong> is obtained.</p>\n<h1>Things that didn't work</h1>\n<p>•    Pseudo labels</p>\n<p>•    Feature normalization</p>",
      "rawMarkdown": "Thanks to Kaggle for organizing such an interesting competition.\nThanks to the teammates who fought side by side. And other Kagglers who share various ideas.\n\n# Context\n•\t[Competition Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n\n•\t[Competition Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n# Overview of the approach\nOverall, our final submitted result is an ensemble of two parts. \n\n**Final submission = Part A×0.7 + Part B×0.3**\n\nWe will explain the composition of Part A and Part B respectively.\n\n# Part A\nIt is an ensemble composed of neural networks with different structures.\n\n### Feature engineering\nAfter many attempts, we finally adopted the following two features from the public notebook [\"OP2: feature engineering\"](https://www.kaggle.com/code/antoninadolgorukova/op2-feature-engineering) as our training and testing feature.\n\n*(1). PCA followed by target encoding (cell type + drug) without noise (pca_target_encoded_features)*     \n*(2). PCA followed by target encoding (cell type + drug) with noise (pca_target_encoded_features_s0.1)*\n\n(1) is subjected to PCA on 18,211 target variables and produced features representing cell type means for drugs and cell type means for compounds. And by using features (2) that add random noise, we believe this will make the model more generalizable.\n\n### Models\n- NN with Fully Connected Layers, as well as BatchNormalization, Dropout, ReLU, and Linear Activation Functions.\n- The initial seed is set to 42 and is fixed. The loss function is mae, the optimizer is Adam.\n- The structure of NN is shown below.\n\n```python\ntf.random.set_seed(42)\n\nmodel = Sequential([ \n    Dense(1228),\n    BatchNormalization(),\n    Activation(\"relu\"),\n    Dropout(0.2),\n    Dense(614),\n    BatchNormalization(),\n    Activation(\"relu\"),\n    Dropout(0.2),\n    Dense(512, activation=\"relu\"),\n    Dropout(0.2),\n    Dense(256, activation=\"relu\"),\n    BatchNormalization(),\n    Dropout(0.2),\n    Dense(128, activation=\"relu\"),\n    Dropout(0.1),\n    Dense(18211,activation= \"linear\")\n])\n\nmodel.compile(loss=\"mae\", \n                optimizer=tf.keras.optimizers.Adam(),\n                metrics=[custom_mean_rowwise_rmse])\n\nhistory = model.fit(full_features, labels, epochs=450, verbose=1)\n\n```\nA simple Part A model training process is shown in this notebook [\"Single-Cell Perturbations(Part A-Model Training)\"](https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training) .\n\n### Model ensemble\nWe use feature (1) and feature (2) to train the network respectively. By changing the feature, the number of network layers, the number of nodes, and the number of training epochs, we successfully obtained a set of individual models scoring 0.567-0.582 on LB.  \nSubsequently, we ensembled 7 models and use it as **Part A** (LB: 0.556/PB 0.741).\n\nIn fact, if the weight of each model is determined based on CV during ensemble, we find that the above results can be further optimized to LB0.556/PB0.74.\n\nMoreover, the highest score we obtained was (LB: 0.557/PB 0.737) after combining 8 models through this method.\n\n# Part B\nThis part is also composed of NN and an ensemble of different models.\n\n### Feature engineering\nThe following two approaches are used to perform feature engineering.\n\n・One-hot encoding on cell_type and sm_name  \n・SMILES(ChemBERTa-77M-MLM)\n\n### Models\n```python\nclass DnnV5(nn.Module):\n    def __init__(self, num_features, num_targets, hidden_size):\n        super(DnnV5, self).__init__() \n        self.conv1d1 = nn.Conv1d(\n            in_channels=752,\n            out_channels=256,\n            kernel_size=1,\n            stride=1,\n            padding=0,\n            bias=True)\n        self.batch_norm1 = nn.BatchNorm1d(256)\n        self.dense1 = nn.utils.weight_norm(nn.Linear(256, 256))\n\n        self.batch_norm2 = nn.BatchNorm1d(256)\n        self.dropout2 = nn.Dropout(0.3)\n        self.dense2 = nn.utils.weight_norm(nn.Linear(256, 512))\n\n\n        self.batch_norm4 = nn.BatchNorm1d(512)\n        self.dropout4 = nn.Dropout(0.1)\n\n        self.dense4 = nn.utils.weight_norm(nn.Linear(512, num_targets))\n\n    def forward(self, x):\n        b,w = x.shape\n        x = x.reshape(b,w,1)\n        x = self.conv1d1(x)\n        x = x.reshape(b,256)\n        x = self.batch_norm1(x)\n        x = F.leaky_relu(self.dense1(x))\n\n        x = self.batch_norm2(x)\n        x = self.dropout2(x)\n        x = F.leaky_relu(self.dense2(x))\n\n\n        x = self.batch_norm4(x)\n        x = self.dropout4(x)\n        y = self.dense4(x)\n        return y\n\n```\n\nThe highest score on LB when using model = DnnV5() for each fold is 0.579.  \nFor an ensemble of 5 fold + 3 seed, the best LB 0.568 for the above model structure can be obtained.\n\nAdditionally, [\"SCP Quickstart\"](https://www.kaggle.com/code/ambrosm/scp-quickstart)  was referred to for fold creation.\n\n### Model ensemble\nMultiple models were created with different smiles and model structures, and the final ensemble was created with the following ratio.\n\n```python\nensemble_submission = 0.5* (sub0568*0.3 +sub0570*0.2+ sub0571*0.1+ sub0571_2*0.2+ sub0573*0.1+sub0576_1*0.05 + sub0576_2*0.05)+ lbsub567*0.5\n```\nAmong them, lbsub567 referred to the public notebook [\"Blend for Single-Cell Perturbations\"](https://www.kaggle.com/code/misakimatsutomo/blend-for-single-cell-perturbations).   \nWith the above model ensemble, we get a score of 0.560 on LB. We use this ensemble as **part B**.\n\nBased on the above process, our final submission is **Part A×0.7 + Part B×0.3**,     and **(LB: 0.554/PB 0.741)** is obtained.\n\n# Things that didn't work\n•\tPseudo labels\n\n•\tFeature normalization",
      "votes": null
    },
    {
      "id": "2552723",
      "postDate": "12/07/2023 17:15:25",
      "content": "<p>Congratulation! It's quite amazing since I did use the same features and a very similar model as your part A (<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution\" target=\"_blank\">this notebook</a> ). But only with train augmentation (simple multiplication of rows by 50 with Gaussian noise), did I achieve min 0.569 score, and with ensembling 8 models - 0.566. Would be great if you publish a notebook with your part A solution, I have so much to learn! </p>",
      "rawMarkdown": "Congratulation! It's quite amazing since I did use the same features and a very similar model as your part A ([this notebook](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution) ). But only with train augmentation (simple multiplication of rows by 50 with Gaussian noise), did I achieve min 0.569 score, and with ensembling 8 models - 0.566. Would be great if you publish a notebook with your part A solution, I have so much to learn!",
      "votes": null
    },
    {
      "id": "2553360",
      "postDate": "12/08/2023 07:34:49",
      "content": "<p>Thank you! Your feature engineering ideas are great! And I really appreciate you sharing your ideas and implementation process with us. I also have a lot to learn. I just updated the solution page and added a simple instruction notebook about training Part A: <a href=\"https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training\" target=\"_blank\">\"Single-Cell Perturbations(Part A-Model Training)\"</a>.</p>",
      "rawMarkdown": "Thank you! Your feature engineering ideas are great! And I really appreciate you sharing your ideas and implementation process with us. I also have a lot to learn. I just updated the solution page and added a simple instruction notebook about training Part A: [\"Single-Cell Perturbations(Part A-Model Training)\"](https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training).",
      "votes": null
    },
    {
      "id": "2696431",
      "postDate": "03/14/2024 10:16:47",
      "content": "<p>Congratulations with gold medal ! <br>\nWould it be possible to share the notebook(s) for Part \"B\" of your solution  ? </p>",
      "rawMarkdown": "Congratulations with gold medal ! \nWould it be possible to share the notebook(s) for Part \"B\" of your solution  ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2552723,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "12/07/2023 17:15:25",
      "content": "<p>Congratulation! It's quite amazing since I did use the same features and a very similar model as your part A (<a href=\"https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution\" target=\"_blank\">this notebook</a> ). But only with train augmentation (simple multiplication of rows by 50 with Gaussian noise), did I achieve min 0.569 score, and with ensembling 8 models - 0.566. Would be great if you publish a notebook with your part A solution, I have so much to learn! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2553360,
          "author_name": "mori123",
          "author_url": "",
          "post_date": "12/08/2023 07:34:49",
          "content": "<p>Thank you! Your feature engineering ideas are great! And I really appreciate you sharing your ideas and implementation process with us. I also have a lot to learn. I just updated the solution page and added a simple instruction notebook about training Part A: <a href=\"https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training\" target=\"_blank\">\"Single-Cell Perturbations(Part A-Model Training)\"</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2696431,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "03/14/2024 10:16:47",
      "content": "<p>Congratulations with gold medal ! <br>\nWould it be possible to share the notebook(s) for Part \"B\" of your solution  ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2552135": "Thanks to Kaggle for organizing such an interesting competition.\nThanks to the teammates who fought side by side. And other Kagglers who share various ideas.\n\n# Context\n•\t[Competition Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n\n•\t[Competition Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n# Overview of the approach\nOverall, our final submitted result is an ensemble of two parts. \n\n**Final submission = Part A×0.7 + Part B×0.3**\n\nWe will explain the composition of Part A and Part B respectively.\n\n# Part A\nIt is an ensemble composed of neural networks with different structures.\n\n### Feature engineering\nAfter many attempts, we finally adopted the following two features from the public notebook [\"OP2: feature engineering\"](https://www.kaggle.com/code/antoninadolgorukova/op2-feature-engineering) as our training and testing feature.\n\n*(1). PCA followed by target encoding (cell type + drug) without noise (pca_target_encoded_features)*     \n*(2). PCA followed by target encoding (cell type + drug) with noise (pca_target_encoded_features_s0.1)*\n\n(1) is subjected to PCA on 18,211 target variables and produced features representing cell type means for drugs and cell type means for compounds. And by using features (2) that add random noise, we believe this will make the model more generalizable.\n\n### Models\n- NN with Fully Connected Layers, as well as BatchNormalization, Dropout, ReLU, and Linear Activation Functions.\n- The initial seed is set to 42 and is fixed. The loss function is mae, the optimizer is Adam.\n- The structure of NN is shown below.\n\n```python\ntf.random.set_seed(42)\n\nmodel = Sequential([ \n    Dense(1228),\n    BatchNormalization(),\n    Activation(\"relu\"),\n    Dropout(0.2),\n    Dense(614),\n    BatchNormalization(),\n    Activation(\"relu\"),\n    Dropout(0.2),\n    Dense(512, activation=\"relu\"),\n    Dropout(0.2),\n    Dense(256, activation=\"relu\"),\n    BatchNormalization(),\n    Dropout(0.2),\n    Dense(128, activation=\"relu\"),\n    Dropout(0.1),\n    Dense(18211,activation= \"linear\")\n])\n\nmodel.compile(loss=\"mae\", \n                optimizer=tf.keras.optimizers.Adam(),\n                metrics=[custom_mean_rowwise_rmse])\n\nhistory = model.fit(full_features, labels, epochs=450, verbose=1)\n\n```\nA simple Part A model training process is shown in this notebook [\"Single-Cell Perturbations(Part A-Model Training)\"](https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training) .\n\n### Model ensemble\nWe use feature (1) and feature (2) to train the network respectively. By changing the feature, the number of network layers, the number of nodes, and the number of training epochs, we successfully obtained a set of individual models scoring 0.567-0.582 on LB.  \nSubsequently, we ensembled 7 models and use it as **Part A** (LB: 0.556/PB 0.741).\n\nIn fact, if the weight of each model is determined based on CV during ensemble, we find that the above results can be further optimized to LB0.556/PB0.74.\n\nMoreover, the highest score we obtained was (LB: 0.557/PB 0.737) after combining 8 models through this method.\n\n# Part B\nThis part is also composed of NN and an ensemble of different models.\n\n### Feature engineering\nThe following two approaches are used to perform feature engineering.\n\n・One-hot encoding on cell_type and sm_name  \n・SMILES(ChemBERTa-77M-MLM)\n\n### Models\n```python\nclass DnnV5(nn.Module):\n    def __init__(self, num_features, num_targets, hidden_size):\n        super(DnnV5, self).__init__() \n        self.conv1d1 = nn.Conv1d(\n            in_channels=752,\n            out_channels=256,\n            kernel_size=1,\n            stride=1,\n            padding=0,\n            bias=True)\n        self.batch_norm1 = nn.BatchNorm1d(256)\n        self.dense1 = nn.utils.weight_norm(nn.Linear(256, 256))\n\n        self.batch_norm2 = nn.BatchNorm1d(256)\n        self.dropout2 = nn.Dropout(0.3)\n        self.dense2 = nn.utils.weight_norm(nn.Linear(256, 512))\n\n\n        self.batch_norm4 = nn.BatchNorm1d(512)\n        self.dropout4 = nn.Dropout(0.1)\n\n        self.dense4 = nn.utils.weight_norm(nn.Linear(512, num_targets))\n\n    def forward(self, x):\n        b,w = x.shape\n        x = x.reshape(b,w,1)\n        x = self.conv1d1(x)\n        x = x.reshape(b,256)\n        x = self.batch_norm1(x)\n        x = F.leaky_relu(self.dense1(x))\n\n        x = self.batch_norm2(x)\n        x = self.dropout2(x)\n        x = F.leaky_relu(self.dense2(x))\n\n\n        x = self.batch_norm4(x)\n        x = self.dropout4(x)\n        y = self.dense4(x)\n        return y\n\n```\n\nThe highest score on LB when using model = DnnV5() for each fold is 0.579.  \nFor an ensemble of 5 fold + 3 seed, the best LB 0.568 for the above model structure can be obtained.\n\nAdditionally, [\"SCP Quickstart\"](https://www.kaggle.com/code/ambrosm/scp-quickstart)  was referred to for fold creation.\n\n### Model ensemble\nMultiple models were created with different smiles and model structures, and the final ensemble was created with the following ratio.\n\n```python\nensemble_submission = 0.5* (sub0568*0.3 +sub0570*0.2+ sub0571*0.1+ sub0571_2*0.2+ sub0573*0.1+sub0576_1*0.05 + sub0576_2*0.05)+ lbsub567*0.5\n```\nAmong them, lbsub567 referred to the public notebook [\"Blend for Single-Cell Perturbations\"](https://www.kaggle.com/code/misakimatsutomo/blend-for-single-cell-perturbations).   \nWith the above model ensemble, we get a score of 0.560 on LB. We use this ensemble as **part B**.\n\nBased on the above process, our final submission is **Part A×0.7 + Part B×0.3**,     and **(LB: 0.554/PB 0.741)** is obtained.\n\n# Things that didn't work\n•\tPseudo labels\n\n•\tFeature normalization",
    "2552723": "Congratulation! It's quite amazing since I did use the same features and a very similar model as your part A ([this notebook](https://www.kaggle.com/code/antoninadolgorukova/op2-simple-mlp-part-of-13th-place-solution) ). But only with train augmentation (simple multiplication of rows by 50 with Gaussian noise), did I achieve min 0.569 score, and with ensembling 8 models - 0.566. Would be great if you publish a notebook with your part A solution, I have so much to learn!",
    "2553360": "Thank you! Your feature engineering ideas are great! And I really appreciate you sharing your ideas and implementation process with us. I also have a lot to learn. I just updated the solution page and added a simple instruction notebook about training Part A: [\"Single-Cell Perturbations(Part A-Model Training)\"](https://www.kaggle.com/code/mori123/single-cell-perturbations-part-a-model-training).",
    "2696431": "Congratulations with gold medal ! \nWould it be possible to share the notebook(s) for Part \"B\" of your solution  ?"
  },
  "source": "meta"
}