{
  "id": 460911,
  "title": "34th Place Solution Writeup for Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/460911",
  "author_name": "Octopus210",
  "post_date": "2023-12-11T17:58:53.632000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to express my gratitude to the organizers and kagglers of this competition. I learned a lot from the open notebooks and discussions.<br>\nEspecially, I would like to express my big thanks to the following people.<br>\n<a href=\"https://www.kaggle.com/mehrankazeminia\" target=\"_blank\">@mehrankazeminia</a><br>\n<a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a><br>\n<a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a><br>\n<a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a><br>\n<a href=\"https://www.kaggle.com/kishanvavdara\" target=\"_blank\">@kishanvavdara</a><br>\n<a href=\"https://www.kaggle.com/pablormier\" target=\"_blank\">@pablormier</a></p>\n<p><strong>34th Solution</strong></p>\n<p>We blended the results of three tasks to find one result. Being different approaches, blending these models was very effective.</p>\n<p>The three approaches<br>\nTask1 : test set as categorical variable<br>\nTask2 : test set as continuous variable ( genes as sampled)<br>\nTask3 : test set as continuous variable ( compounds as sample)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2646279%2Fb2df8d249afa3fa8af194ba5d14f3438%2FOC2_34th.png?generation=1702317451136015&amp;alt=media\" alt=\"34th_solution\"></p>\n<p>Basically, it is a simple solution that blends public notebook's score and sklearn's model. If you combine it with a better performing model (like PY-BOOST), the performance will be even better.</p>\n<ol>\n<li>integration of biological knowledge<br>\nIn Task 3, we incorporated the decoupler. A small improvement in public scores was obtained. </li>\n</ol>\n<p>2.Exploration of the problem<br>\nIn Task 2 and Task 3, CD8 data were excluded. <a href=\"https://www.kaggle.com/code/yoshifumimiya/op2-about-positive-control\" target=\"_blank\">op2-about-positive-control</a></p>\n<p>3.Model design<br>\nI used sklearn's MLP, lightgbm, and Ridge. </p>\n<p>4.Robustness<br>\nRobustness is considered to be high, because I sought a single result from three different perspectives. </p>\n<p>5.Documentation &amp; code style<br>\nIn preparation</p>\n<p>6.Reproducibility<br>\nIn preparation</p>",
  "messages": [
    {
      "id": 2557896,
      "postDate": "2023-12-11T17:58:53.633Z",
      "content": "<p>First of all, I would like to express my gratitude to the organizers and kagglers of this competition. I learned a lot from the open notebooks and discussions.<br>\nEspecially, I would like to express my big thanks to the following people.<br>\n<a href=\"https://www.kaggle.com/mehrankazeminia\" target=\"_blank\">@mehrankazeminia</a><br>\n<a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a><br>\n<a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a><br>\n<a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a><br>\n<a href=\"https://www.kaggle.com/kishanvavdara\" target=\"_blank\">@kishanvavdara</a><br>\n<a href=\"https://www.kaggle.com/pablormier\" target=\"_blank\">@pablormier</a></p>\n<p><strong>34th Solution</strong></p>\n<p>We blended the results of three tasks to find one result. Being different approaches, blending these models was very effective.</p>\n<p>The three approaches<br>\nTask1 : test set as categorical variable<br>\nTask2 : test set as continuous variable ( genes as sampled)<br>\nTask3 : test set as continuous variable ( compounds as sample)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2646279%2Fb2df8d249afa3fa8af194ba5d14f3438%2FOC2_34th.png?generation=1702317451136015&amp;alt=media\" alt=\"34th_solution\"></p>\n<p>Basically, it is a simple solution that blends public notebook's score and sklearn's model. If you combine it with a better performing model (like PY-BOOST), the performance will be even better.</p>\n<ol>\n<li>integration of biological knowledge<br>\nIn Task 3, we incorporated the decoupler. A small improvement in public scores was obtained. </li>\n</ol>\n<p>2.Exploration of the problem<br>\nIn Task 2 and Task 3, CD8 data were excluded. <a href=\"https://www.kaggle.com/code/yoshifumimiya/op2-about-positive-control\" target=\"_blank\">op2-about-positive-control</a></p>\n<p>3.Model design<br>\nI used sklearn's MLP, lightgbm, and Ridge. </p>\n<p>4.Robustness<br>\nRobustness is considered to be high, because I sought a single result from three different perspectives. </p>\n<p>5.Documentation &amp; code style<br>\nIn preparation</p>\n<p>6.Reproducibility<br>\nIn preparation</p>",
      "rawMarkdown": "First of all, I would like to express my gratitude to the organizers and kagglers of this competition. I learned a lot from the open notebooks and discussions.\nEspecially, I would like to express my big thanks to the following people.\n@mehrankazeminia\n@somayyehgholami\n@alexandervc\n@antoninadolgorukova\n@kishanvavdara\n@pablormier\n\n\n\n**34th Solution**\n\nWe blended the results of three tasks to find one result. Being different approaches, blending these models was very effective.\n\nThe three approaches\nTask1 : test set as categorical variable\nTask2 : test set as continuous variable ( genes as sampled)\nTask3 : test set as continuous variable ( compounds as sample)\n\n![34th_solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2646279%2Fb2df8d249afa3fa8af194ba5d14f3438%2FOC2_34th.png?generation=1702317451136015&alt=media)\n\nBasically, it is a simple solution that blends public notebook's score and sklearn's model. If you combine it with a better performing model (like PY-BOOST), the performance will be even better.\n\n1. integration of biological knowledge\nIn Task 3, we incorporated the decoupler. A small improvement in public scores was obtained. \n\n2.Exploration of the problem\nIn Task 2 and Task 3, CD8 data were excluded. [op2-about-positive-control](https://www.kaggle.com/code/yoshifumimiya/op2-about-positive-control)\n\n3.Model design\nI used sklearn's MLP, lightgbm, and Ridge. \n\n4.Robustness\nRobustness is considered to be high, because I sought a single result from three different perspectives. \n\n5.Documentation & code style\nIn preparation\n\n6.Reproducibility\nIn preparation",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2557896": "First of all, I would like to express my gratitude to the organizers and kagglers of this competition. I learned a lot from the open notebooks and discussions.\nEspecially, I would like to express my big thanks to the following people.\n@mehrankazeminia\n@somayyehgholami\n@alexandervc\n@antoninadolgorukova\n@kishanvavdara\n@pablormier\n\n\n\n**34th Solution**\n\nWe blended the results of three tasks to find one result. Being different approaches, blending these models was very effective.\n\nThe three approaches\nTask1 : test set as categorical variable\nTask2 : test set as continuous variable ( genes as sampled)\nTask3 : test set as continuous variable ( compounds as sample)\n\n![34th_solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2646279%2Fb2df8d249afa3fa8af194ba5d14f3438%2FOC2_34th.png?generation=1702317451136015&alt=media)\n\nBasically, it is a simple solution that blends public notebook's score and sklearn's model. If you combine it with a better performing model (like PY-BOOST), the performance will be even better.\n\n1. integration of biological knowledge\nIn Task 3, we incorporated the decoupler. A small improvement in public scores was obtained. \n\n2.Exploration of the problem\nIn Task 2 and Task 3, CD8 data were excluded. [op2-about-positive-control](https://www.kaggle.com/code/yoshifumimiya/op2-about-positive-control)\n\n3.Model design\nI used sklearn's MLP, lightgbm, and Ridge. \n\n4.Robustness\nRobustness is considered to be high, because I sought a single result from three different perspectives. \n\n5.Documentation & code style\nIn preparation\n\n6.Reproducibility\nIn preparation"
  }
}