{
  "id": 451245,
  "title": "Feature Augmentation",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/451245",
  "author_name": "Mehran Kazeminia",
  "post_date": "2023-10-27T18:05:37.283000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<ul>\n<li><p>Two available features; are \"cell_type\" and \"sm_name\" and we added two new columns (two new features) to them in <a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augmentation\" target=\"_blank\">our new notebook</a>.</p></li>\n<li><p>If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. </p></li>\n<li><p>Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. </p></li>\n<li><p>Obviously, to add these two features, TrainData and TestData must be customized for each y column, and this may seem a bit complicated. For this reason, we first performed all the calculations only on column zero and then continued the main calculations in a loop with \"range(y.shape[1])\".</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffbb8a9d4412a1f052e0cdeb70aace49f%2Fop104.png?generation=1698429062467385&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F263b16596fdff184807767dc9cdc482d%2Fop103.png?generation=1698428872023115&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>By adding these two new features, the score of this notebook improved and probably the score of all notebooks that use the usual methods in machine learning (such as neural network, etc.) will be better.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ff2293928ebe057cb7def207718344499%2Fop105.png?generation=1698429275436133&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>Of course, since the beginning of this challenge, many public notebooks have used averaging methods, but in these notebooks, the obtained values are directly considered as the answer.</p></li>\n<li><p>It should be noted that in the notebooks mentioned above, guesses are made to find the effect of averages or their combination, and these guesses will probably cause instability in the model as well as the risk of overfitting.</p></li>\n<li><p>Ensembling Histograms for two different columns chosen at random:</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F93370e0b6008dc614d97ad06d773d455%2Fop101.png?generation=1698429541605565&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F0621480492a22b280e74ff6e56cf5d04%2Fop102.png?generation=1698429568102236&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2501868,
      "postDate": "2023-10-27T18:05:37.283Z",
      "content": "<ul>\n<li><p>Two available features; are \"cell_type\" and \"sm_name\" and we added two new columns (two new features) to them in <a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augmentation\" target=\"_blank\">our new notebook</a>.</p></li>\n<li><p>If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. </p></li>\n<li><p>Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. </p></li>\n<li><p>Obviously, to add these two features, TrainData and TestData must be customized for each y column, and this may seem a bit complicated. For this reason, we first performed all the calculations only on column zero and then continued the main calculations in a loop with \"range(y.shape[1])\".</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffbb8a9d4412a1f052e0cdeb70aace49f%2Fop104.png?generation=1698429062467385&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F263b16596fdff184807767dc9cdc482d%2Fop103.png?generation=1698428872023115&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>By adding these two new features, the score of this notebook improved and probably the score of all notebooks that use the usual methods in machine learning (such as neural network, etc.) will be better.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ff2293928ebe057cb7def207718344499%2Fop105.png?generation=1698429275436133&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>Of course, since the beginning of this challenge, many public notebooks have used averaging methods, but in these notebooks, the obtained values are directly considered as the answer.</p></li>\n<li><p>It should be noted that in the notebooks mentioned above, guesses are made to find the effect of averages or their combination, and these guesses will probably cause instability in the model as well as the risk of overfitting.</p></li>\n<li><p>Ensembling Histograms for two different columns chosen at random:</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F93370e0b6008dc614d97ad06d773d455%2Fop101.png?generation=1698429541605565&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F0621480492a22b280e74ff6e56cf5d04%2Fop102.png?generation=1698429568102236&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "- Two available features; are \"cell_type\" and \"sm_name\" and we added two new columns (two new features) to them in [our new notebook](https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augmentation).\n\n- If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. \n\n- Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. \n\n- Obviously, to add these two features, TrainData and TestData must be customized for each y column, and this may seem a bit complicated. For this reason, we first performed all the calculations only on column zero and then continued the main calculations in a loop with \"range(y.shape[1])\".\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffbb8a9d4412a1f052e0cdeb70aace49f%2Fop104.png?generation=1698429062467385&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F263b16596fdff184807767dc9cdc482d%2Fop103.png?generation=1698428872023115&alt=media)\n\n- By adding these two new features, the score of this notebook improved and probably the score of all notebooks that use the usual methods in machine learning (such as neural network, etc.) will be better.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ff2293928ebe057cb7def207718344499%2Fop105.png?generation=1698429275436133&alt=media)\n\n- Of course, since the beginning of this challenge, many public notebooks have used averaging methods, but in these notebooks, the obtained values are directly considered as the answer.\n\n- It should be noted that in the notebooks mentioned above, guesses are made to find the effect of averages or their combination, and these guesses will probably cause instability in the model as well as the risk of overfitting.\n\n- Ensembling Histograms for two different columns chosen at random:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F93370e0b6008dc614d97ad06d773d455%2Fop101.png?generation=1698429541605565&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F0621480492a22b280e74ff6e56cf5d04%2Fop102.png?generation=1698429568102236&alt=media)\n",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2501868": "- Two available features; are \"cell_type\" and \"sm_name\" and we added two new columns (two new features) to them in [our new notebook](https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augmentation).\n\n- If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. \n\n- Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. \n\n- Obviously, to add these two features, TrainData and TestData must be customized for each y column, and this may seem a bit complicated. For this reason, we first performed all the calculations only on column zero and then continued the main calculations in a loop with \"range(y.shape[1])\".\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffbb8a9d4412a1f052e0cdeb70aace49f%2Fop104.png?generation=1698429062467385&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F263b16596fdff184807767dc9cdc482d%2Fop103.png?generation=1698428872023115&alt=media)\n\n- By adding these two new features, the score of this notebook improved and probably the score of all notebooks that use the usual methods in machine learning (such as neural network, etc.) will be better.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ff2293928ebe057cb7def207718344499%2Fop105.png?generation=1698429275436133&alt=media)\n\n- Of course, since the beginning of this challenge, many public notebooks have used averaging methods, but in these notebooks, the obtained values are directly considered as the answer.\n\n- It should be noted that in the notebooks mentioned above, guesses are made to find the effect of averages or their combination, and these guesses will probably cause instability in the model as well as the risk of overfitting.\n\n- Ensembling Histograms for two different columns chosen at random:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F93370e0b6008dc614d97ad06d773d455%2Fop101.png?generation=1698429541605565&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F0621480492a22b280e74ff6e56cf5d04%2Fop102.png?generation=1698429568102236&alt=media)\n"
  }
}