{
  "id": 459618,
  "title": "8th Place Solution for the Open Problems – Single-Cell Perturbations Competition",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/459618",
  "author_name": "aper",
  "post_date": "2023-12-06T02:24:03.801000",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Congrats to all the winners, and thank you to Kaggle for organizing such an interesting competition. And also thanks to the other kagglers who shared their ideas and notebooks.</p>\n<h1>Context</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Competition Overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Competition Data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<h3>Model</h3>\n<p>For model development, I designed and fine-tuned a simple neural network through a series of experiments, aiming to reduce both CV and LB scores. I used augment features from the <a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles?scriptVersionId=150423767&amp;cellId=24\" target=\"_blank\">notebook - [3] OP2 - Feature Augment &amp; Fragments of SMILES</a> as model inputs.</p>\n<p>Here is the model architecture:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.ce_layer = nn.Linear(labels, dim_size)\n        self.sm_layer = nn.Linear(labels, dim_size)\n\n        hidden_size = dim_size * \n        self.fc1 = nn.Linear(hidden_size, hidden_size*mul_ratio)\n        self.fc2 = nn.Linear(hidden_size*mul_ratio, hidden_size)\n\n        self.act = nn.GELU()\n        self.out = nn.Linear(hidden_size, labels)\n\n     ():\n        x1 = self.act(self.ce_layer(cell_type))\n        x2 = self.act(self.sm_layer(sm_name))\n\n        x = torch.concat([x1, x2], dim=-)\n\n        x = self.act(self.fc1(x))\n        x = self.act(self.fc2(x))\n\n        x = self.out(x)\n         x\n</code></pre>\n<h3>Data Augmentation</h3>\n<p>The training process invloved a strategic approach to data augmentation. Initially, I employed only mean values for the cell_type and sm_name, respectively. Subsequently, I explored various statistical values such as median, min, max and quantiles. And I found out that median values significantly improves the LB score.</p>\n<p>Moreover, I experimented with combinations of these features. I implemented 50% random selection between mean and median, and 25% random selection among mean, median, Q1 and Q2 for both cell_type and sm_name.  </p>\n<h3>Validation Strategy</h3>\n<p>I used K-Fold cross-validation stratify on cell_type, trained 5, 10, 15, 20 splits. I reviewed that LB score increases with 10 and 15 splits.</p>\n<h1>Details of the submission</h1>\n<p>The results of augmented models are summarized below, along with their respective LB scores.</p>\n<ol>\n<li>median / 0.549</li>\n<li>mean and median / 0.549</li>\n<li>mean, median, Q1 and Q2 / 0.551</li>\n</ol>\n<p>The final submission was a weighted average of these models by 0.35/0.35/0.3, which boosted up the LB score to 0.547. Since they had different prediction distributions with similiar LB scores, I believed that the ensemble would generalize well in the private.  </p>\n<h3>Things that didn't work</h3>\n<ul>\n<li>pseudo labels</li>\n<li>dropout</li>\n<li>normalization</li>\n<li>data selection (control, etc.)</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.nature.com/articles/s41592-023-01969-x\" target=\"_blank\">Learning single-cell perturbation responses using neural optimal transport</a></li>\n<li><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a></li>\n</ul>",
  "messages": [
    {
      "id": 2550385,
      "postDate": "2023-12-06T02:24:03.800Z",
      "content": "<p>Congrats to all the winners, and thank you to Kaggle for organizing such an interesting competition. And also thanks to the other kagglers who shared their ideas and notebooks.</p>\n<h1>Context</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Competition Overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Competition Data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<h3>Model</h3>\n<p>For model development, I designed and fine-tuned a simple neural network through a series of experiments, aiming to reduce both CV and LB scores. I used augment features from the <a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles?scriptVersionId=150423767&amp;cellId=24\" target=\"_blank\">notebook - [3] OP2 - Feature Augment &amp; Fragments of SMILES</a> as model inputs.</p>\n<p>Here is the model architecture:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.ce_layer = nn.Linear(labels, dim_size)\n        self.sm_layer = nn.Linear(labels, dim_size)\n\n        hidden_size = dim_size * \n        self.fc1 = nn.Linear(hidden_size, hidden_size*mul_ratio)\n        self.fc2 = nn.Linear(hidden_size*mul_ratio, hidden_size)\n\n        self.act = nn.GELU()\n        self.out = nn.Linear(hidden_size, labels)\n\n     ():\n        x1 = self.act(self.ce_layer(cell_type))\n        x2 = self.act(self.sm_layer(sm_name))\n\n        x = torch.concat([x1, x2], dim=-)\n\n        x = self.act(self.fc1(x))\n        x = self.act(self.fc2(x))\n\n        x = self.out(x)\n         x\n</code></pre>\n<h3>Data Augmentation</h3>\n<p>The training process invloved a strategic approach to data augmentation. Initially, I employed only mean values for the cell_type and sm_name, respectively. Subsequently, I explored various statistical values such as median, min, max and quantiles. And I found out that median values significantly improves the LB score.</p>\n<p>Moreover, I experimented with combinations of these features. I implemented 50% random selection between mean and median, and 25% random selection among mean, median, Q1 and Q2 for both cell_type and sm_name.  </p>\n<h3>Validation Strategy</h3>\n<p>I used K-Fold cross-validation stratify on cell_type, trained 5, 10, 15, 20 splits. I reviewed that LB score increases with 10 and 15 splits.</p>\n<h1>Details of the submission</h1>\n<p>The results of augmented models are summarized below, along with their respective LB scores.</p>\n<ol>\n<li>median / 0.549</li>\n<li>mean and median / 0.549</li>\n<li>mean, median, Q1 and Q2 / 0.551</li>\n</ol>\n<p>The final submission was a weighted average of these models by 0.35/0.35/0.3, which boosted up the LB score to 0.547. Since they had different prediction distributions with similiar LB scores, I believed that the ensemble would generalize well in the private.  </p>\n<h3>Things that didn't work</h3>\n<ul>\n<li>pseudo labels</li>\n<li>dropout</li>\n<li>normalization</li>\n<li>data selection (control, etc.)</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.nature.com/articles/s41592-023-01969-x\" target=\"_blank\">Learning single-cell perturbation responses using neural optimal transport</a></li>\n<li><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a></li>\n</ul>",
      "rawMarkdown": "Congrats to all the winners, and thank you to Kaggle for organizing such an interesting competition. And also thanks to the other kagglers who shared their ideas and notebooks.\n\n# Context\n- [Competition Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n- [Competition Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n# Overview of the Approach\n\n### Model\nFor model development, I designed and fine-tuned a simple neural network through a series of experiments, aiming to reduce both CV and LB scores. I used augment features from the [notebook - [3] OP2 - Feature Augment & Fragments of SMILES](https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles?scriptVersionId=150423767&cellId=24) as model inputs.\n\nHere is the model architecture:\n```python\nclass SingleCellModel(nn.Module):\n    def __init__(self, dim_size, mul_ratio, labels=18211):\n        super().__init__()\n        self.ce_layer = nn.Linear(labels, dim_size)\n        self.sm_layer = nn.Linear(labels, dim_size)\n\n        hidden_size = dim_size * 2\n        self.fc1 = nn.Linear(hidden_size, hidden_size*mul_ratio)\n        self.fc2 = nn.Linear(hidden_size*mul_ratio, hidden_size)\n        \n        self.act = nn.GELU()\n        self.out = nn.Linear(hidden_size, labels)\n \n    def forward(self, cell_type, sm_name):\n        x1 = self.act(self.ce_layer(cell_type))\n        x2 = self.act(self.sm_layer(sm_name))\n\n        x = torch.concat([x1, x2], dim=-1)\n\n        x = self.act(self.fc1(x))\n        x = self.act(self.fc2(x))\n\n        x = self.out(x)\n        return x\n```\n\n### Data Augmentation\nThe training process invloved a strategic approach to data augmentation. Initially, I employed only mean values for the cell_type and sm_name, respectively. Subsequently, I explored various statistical values such as median, min, max and quantiles. And I found out that median values significantly improves the LB score.\n\nMoreover, I experimented with combinations of these features. I implemented 50% random selection between mean and median, and 25% random selection among mean, median, Q1 and Q2 for both cell_type and sm_name.  \n\n### Validation Strategy\nI used K-Fold cross-validation stratify on cell_type, trained 5, 10, 15, 20 splits. I reviewed that LB score increases with 10 and 15 splits.\n\n# Details of the submission\n\nThe results of augmented models are summarized below, along with their respective LB scores.\n1. median / 0.549\n2. mean and median / 0.549\n3. mean, median, Q1 and Q2 / 0.551\n\nThe final submission was a weighted average of these models by 0.35/0.35/0.3, which boosted up the LB score to 0.547. Since they had different prediction distributions with similiar LB scores, I believed that the ensemble would generalize well in the private.  \n\n### Things that didn't work\n- pseudo labels\n- dropout\n- normalization\n- data selection (control, etc.)\n\n# Sources\n\n- [Learning single-cell perturbation responses using neural optimal transport](https://www.nature.com/articles/s41592-023-01969-x)\n- https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles",
      "votes": 7
    },
    {
      "id": 2691898,
      "postDate": "2024-03-11T14:14:28.080Z",
      "content": "<p>Congratulations with great results ! <br>\nWould it be possible to share more detailed code - something like Kaggle notebook (preferably) or github ? </p>",
      "rawMarkdown": "Congratulations with great results ! \nWould it be possible to share more detailed code - something like Kaggle notebook (preferably) or github ? ",
      "votes": 3,
      "replies": [
        {
          "id": 2703346,
          "postDate": "2024-03-18T05:45:24.443Z",
          "content": "<p>No problem. Here's the <a href=\"https://www.kaggle.com/code/todaya/op2-notebook\" target=\"_blank\">link</a> to the notebook.</p>",
          "rawMarkdown": "No problem. Here's the [link](https://www.kaggle.com/code/todaya/op2-notebook) to the notebook.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2559140,
      "postDate": "2023-12-12T16:13:41.903Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2559538,
          "postDate": "2023-12-12T22:53:44.823Z",
          "content": "<p>Thanks, glad you find it interesting.</p>",
          "rawMarkdown": "Thanks, glad you find it interesting.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2553847,
      "postDate": "2023-12-08T14:50:14.403Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2691898,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2024-03-11T14:14:28.080000",
      "content": "<p>Congratulations with great results ! <br>\nWould it be possible to share more detailed code - something like Kaggle notebook (preferably) or github ? </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2703346,
          "author_name": "aper",
          "author_url": "",
          "post_date": "2024-03-18T05:45:24.443000",
          "content": "<p>No problem. Here's the <a href=\"https://www.kaggle.com/code/todaya/op2-notebook\" target=\"_blank\">link</a> to the notebook.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2559140,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-12T16:13:41.903000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2559538,
          "author_name": "aper",
          "author_url": "",
          "post_date": "2023-12-12T22:53:44.823000",
          "content": "<p>Thanks, glad you find it interesting.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2553847,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T14:50:14.403000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2550385": "Congrats to all the winners, and thank you to Kaggle for organizing such an interesting competition. And also thanks to the other kagglers who shared their ideas and notebooks.\n\n# Context\n- [Competition Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n- [Competition Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n# Overview of the Approach\n\n### Model\nFor model development, I designed and fine-tuned a simple neural network through a series of experiments, aiming to reduce both CV and LB scores. I used augment features from the [notebook - [3] OP2 - Feature Augment & Fragments of SMILES](https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles?scriptVersionId=150423767&cellId=24) as model inputs.\n\nHere is the model architecture:\n```python\nclass SingleCellModel(nn.Module):\n    def __init__(self, dim_size, mul_ratio, labels=18211):\n        super().__init__()\n        self.ce_layer = nn.Linear(labels, dim_size)\n        self.sm_layer = nn.Linear(labels, dim_size)\n\n        hidden_size = dim_size * 2\n        self.fc1 = nn.Linear(hidden_size, hidden_size*mul_ratio)\n        self.fc2 = nn.Linear(hidden_size*mul_ratio, hidden_size)\n        \n        self.act = nn.GELU()\n        self.out = nn.Linear(hidden_size, labels)\n \n    def forward(self, cell_type, sm_name):\n        x1 = self.act(self.ce_layer(cell_type))\n        x2 = self.act(self.sm_layer(sm_name))\n\n        x = torch.concat([x1, x2], dim=-1)\n\n        x = self.act(self.fc1(x))\n        x = self.act(self.fc2(x))\n\n        x = self.out(x)\n        return x\n```\n\n### Data Augmentation\nThe training process invloved a strategic approach to data augmentation. Initially, I employed only mean values for the cell_type and sm_name, respectively. Subsequently, I explored various statistical values such as median, min, max and quantiles. And I found out that median values significantly improves the LB score.\n\nMoreover, I experimented with combinations of these features. I implemented 50% random selection between mean and median, and 25% random selection among mean, median, Q1 and Q2 for both cell_type and sm_name.  \n\n### Validation Strategy\nI used K-Fold cross-validation stratify on cell_type, trained 5, 10, 15, 20 splits. I reviewed that LB score increases with 10 and 15 splits.\n\n# Details of the submission\n\nThe results of augmented models are summarized below, along with their respective LB scores.\n1. median / 0.549\n2. mean and median / 0.549\n3. mean, median, Q1 and Q2 / 0.551\n\nThe final submission was a weighted average of these models by 0.35/0.35/0.3, which boosted up the LB score to 0.547. Since they had different prediction distributions with similiar LB scores, I believed that the ensemble would generalize well in the private.  \n\n### Things that didn't work\n- pseudo labels\n- dropout\n- normalization\n- data selection (control, etc.)\n\n# Sources\n\n- [Learning single-cell perturbation responses using neural optimal transport](https://www.nature.com/articles/s41592-023-01969-x)\n- https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles",
    "2691898": "Congratulations with great results ! \nWould it be possible to share more detailed code - something like Kaggle notebook (preferably) or github ? ",
    "2559140": "",
    "2553847": ""
  }
}