{
  "id": 459588,
  "title": "Public 7th and Private 15th solution (Nothing but just multiplied a factor of 1.2)",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/459588",
  "author_name": "Jie Wu",
  "post_date": "2023-12-05T23:35:55.666000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First many thanks go to those who created 0.574 and 0.577 public notebook.</p>\n<p>While I started this competition 2 months ago, I found the public LB is very variant, it’s very different from CV. Ensemble some worse LB results like everyone (or most people) found that ensemble results with 0.702 LB will boost the result. So I thought this would be a shake competition. Hence, I didn’t spend too much time on how to build a more diversity or robust model (I didn’t think I could), but I focused on some LB probing or tricks. For instance, multiplying a factor to any of my results (the best one is 1.2 on public LB), I could boost my results. This made me more confident that there would be some big shake up, but I also believed some of the gold medal teams will be quite stable, they will stay there.</p>\n<p>I used public 0.574 and 0.577 results + some of my own models (Public LB 0.578).</p>\n<p>For my own models:<br>\n•    Conv1D NN<br>\n•    LSTM<br>\n•    MLP<br>\n•    LGBM</p>\n<p>Features:<br>\n•    Standard Scaler train label columns for NN models<br>\n•    One-hot encoded cell_type,sm_name <br>\n•    Split SMILES to character and use TFIDF to get embedding.</p>\n<p>Final results:<br>\nStep1<br>\n•    sub_pub[:128] = 0.55<em>Public 0.574[:128] + 0.45</em>public 0.577[:128]<br>\n•    sub_pub [128:] = 0.6<em>Public 0.574[128:] + 0.4</em>public 0.577[128:]<br>\n•    then postprocess it using what’s done in <a href=\"https://www.kaggle.com/code/jeffreylihkust/op2-eda-lb\" target=\"_blank\">https://www.kaggle.com/code/jeffreylihkust/op2-eda-lb</a> </p>\n<p>Step2<br>\n•    final_sub = 1.2<em>(0.95</em>sub_pub + 0.05*my_0578)</p>\n<p>About the factor:<br>\nI tried factors of 0.95, 1.05, 1.1, 1.15, 1.2, 1.25, 1.3, 1.4, 1.5, but 1.2 gives best public LB.</p>\n<p>A bit pity, I have a few results in the gold area, but their public LBs is 0.02 worse.</p>",
  "messages": [
    {
      "id": 2550239,
      "postDate": "2023-12-05T23:35:55.667Z",
      "content": "<p>First many thanks go to those who created 0.574 and 0.577 public notebook.</p>\n<p>While I started this competition 2 months ago, I found the public LB is very variant, it’s very different from CV. Ensemble some worse LB results like everyone (or most people) found that ensemble results with 0.702 LB will boost the result. So I thought this would be a shake competition. Hence, I didn’t spend too much time on how to build a more diversity or robust model (I didn’t think I could), but I focused on some LB probing or tricks. For instance, multiplying a factor to any of my results (the best one is 1.2 on public LB), I could boost my results. This made me more confident that there would be some big shake up, but I also believed some of the gold medal teams will be quite stable, they will stay there.</p>\n<p>I used public 0.574 and 0.577 results + some of my own models (Public LB 0.578).</p>\n<p>For my own models:<br>\n•    Conv1D NN<br>\n•    LSTM<br>\n•    MLP<br>\n•    LGBM</p>\n<p>Features:<br>\n•    Standard Scaler train label columns for NN models<br>\n•    One-hot encoded cell_type,sm_name <br>\n•    Split SMILES to character and use TFIDF to get embedding.</p>\n<p>Final results:<br>\nStep1<br>\n•    sub_pub[:128] = 0.55<em>Public 0.574[:128] + 0.45</em>public 0.577[:128]<br>\n•    sub_pub [128:] = 0.6<em>Public 0.574[128:] + 0.4</em>public 0.577[128:]<br>\n•    then postprocess it using what’s done in <a href=\"https://www.kaggle.com/code/jeffreylihkust/op2-eda-lb\" target=\"_blank\">https://www.kaggle.com/code/jeffreylihkust/op2-eda-lb</a> </p>\n<p>Step2<br>\n•    final_sub = 1.2<em>(0.95</em>sub_pub + 0.05*my_0578)</p>\n<p>About the factor:<br>\nI tried factors of 0.95, 1.05, 1.1, 1.15, 1.2, 1.25, 1.3, 1.4, 1.5, but 1.2 gives best public LB.</p>\n<p>A bit pity, I have a few results in the gold area, but their public LBs is 0.02 worse.</p>",
      "rawMarkdown": "First many thanks go to those who created 0.574 and 0.577 public notebook.\n\nWhile I started this competition 2 months ago, I found the public LB is very variant, it’s very different from CV. Ensemble some worse LB results like everyone (or most people) found that ensemble results with 0.702 LB will boost the result. So I thought this would be a shake competition. Hence, I didn’t spend too much time on how to build a more diversity or robust model (I didn’t think I could), but I focused on some LB probing or tricks. For instance, multiplying a factor to any of my results (the best one is 1.2 on public LB), I could boost my results. This made me more confident that there would be some big shake up, but I also believed some of the gold medal teams will be quite stable, they will stay there.\n\n\nI used public 0.574 and 0.577 results + some of my own models (Public LB 0.578).\n\nFor my own models:\n•\tConv1D NN\n•\tLSTM\n•\tMLP\n•\tLGBM\n\nFeatures:\n•\tStandard Scaler train label columns for NN models\n•\tOne-hot encoded cell_type,sm_name \n•\tSplit SMILES to character and use TFIDF to get embedding.\n\n\nFinal results:\nStep1\n•\tsub_pub[:128] = 0.55*Public 0.574[:128] + 0.45*public 0.577[:128]\n•\tsub_pub [128:] = 0.6*Public 0.574[128:] + 0.4*public 0.577[128:]\n•\tthen postprocess it using what’s done in https://www.kaggle.com/code/jeffreylihkust/op2-eda-lb \n\nStep2\n•\tfinal_sub = 1.2*(0.95*sub_pub + 0.05*my_0578)\n\nAbout the factor:\nI tried factors of 0.95, 1.05, 1.1, 1.15, 1.2, 1.25, 1.3, 1.4, 1.5, but 1.2 gives best public LB.\n\nA bit pity, I have a few results in the gold area, but their public LBs is 0.02 worse.\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2550239": "First many thanks go to those who created 0.574 and 0.577 public notebook.\n\nWhile I started this competition 2 months ago, I found the public LB is very variant, it’s very different from CV. Ensemble some worse LB results like everyone (or most people) found that ensemble results with 0.702 LB will boost the result. So I thought this would be a shake competition. Hence, I didn’t spend too much time on how to build a more diversity or robust model (I didn’t think I could), but I focused on some LB probing or tricks. For instance, multiplying a factor to any of my results (the best one is 1.2 on public LB), I could boost my results. This made me more confident that there would be some big shake up, but I also believed some of the gold medal teams will be quite stable, they will stay there.\n\n\nI used public 0.574 and 0.577 results + some of my own models (Public LB 0.578).\n\nFor my own models:\n•\tConv1D NN\n•\tLSTM\n•\tMLP\n•\tLGBM\n\nFeatures:\n•\tStandard Scaler train label columns for NN models\n•\tOne-hot encoded cell_type,sm_name \n•\tSplit SMILES to character and use TFIDF to get embedding.\n\n\nFinal results:\nStep1\n•\tsub_pub[:128] = 0.55*Public 0.574[:128] + 0.45*public 0.577[:128]\n•\tsub_pub [128:] = 0.6*Public 0.574[128:] + 0.4*public 0.577[128:]\n•\tthen postprocess it using what’s done in https://www.kaggle.com/code/jeffreylihkust/op2-eda-lb \n\nStep2\n•\tfinal_sub = 1.2*(0.95*sub_pub + 0.05*my_0578)\n\nAbout the factor:\nI tried factors of 0.95, 1.05, 1.1, 1.15, 1.2, 1.25, 1.3, 1.4, 1.5, but 1.2 gives best public LB.\n\nA bit pity, I have a few results in the gold area, but their public LBs is 0.02 worse.\n"
  }
}