{
  "id": 518941,
  "title": "Target Encoding",
  "url": "/competitions/leash-BELKA/discussion/518941",
  "author_name": "Robert Hatch",
  "post_date": "2024-07-09T00:35:13.191000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Early on, a week or two(?) into the competition, I got … 5th? on the leaderboard with just taking a public baseline model (KNN because I thought it might generalize better to non-shared BBs) and predicting ONLY shared BBs with my own model.</p>\n<p>My model was just a simple target encoding of the building blocks and building block pairs. 6 features, plus the protein. Importantly, for bb2 and bb3; and also for bb1+2 and bb1+3, I rearranged the target encoding results to be bb2/3_max and bb2/3min, and likewise bb1+2_max and bb1+3_min.</p>\n<p>Anyways, really the impact of that wasn't about clever feature engineering or not so clever. The issue was the model did too well! Target encoding by itself was so far ahead of ECFP, which spelled trouble for the competition. It was roughly on par with the 1dCNN that came later (and ensembled well with it on public LB). Because of that, from early on, I was very pessimistic about the chances of any real generalization on non-shared BBs. I'm happy a few people managed to create robust models that were at least partially effective, but it remains that the majority of the data in the train set is best at predicting itself.</p>",
  "messages": [
    {
      "id": 2912516,
      "postDate": "2024-07-09T00:35:13.190Z",
      "content": "<p>Early on, a week or two(?) into the competition, I got … 5th? on the leaderboard with just taking a public baseline model (KNN because I thought it might generalize better to non-shared BBs) and predicting ONLY shared BBs with my own model.</p>\n<p>My model was just a simple target encoding of the building blocks and building block pairs. 6 features, plus the protein. Importantly, for bb2 and bb3; and also for bb1+2 and bb1+3, I rearranged the target encoding results to be bb2/3_max and bb2/3min, and likewise bb1+2_max and bb1+3_min.</p>\n<p>Anyways, really the impact of that wasn't about clever feature engineering or not so clever. The issue was the model did too well! Target encoding by itself was so far ahead of ECFP, which spelled trouble for the competition. It was roughly on par with the 1dCNN that came later (and ensembled well with it on public LB). Because of that, from early on, I was very pessimistic about the chances of any real generalization on non-shared BBs. I'm happy a few people managed to create robust models that were at least partially effective, but it remains that the majority of the data in the train set is best at predicting itself.</p>",
      "rawMarkdown": "Early on, a week or two(?) into the competition, I got ... 5th? on the leaderboard with just taking a public baseline model (KNN because I thought it might generalize better to non-shared BBs) and predicting ONLY shared BBs with my own model.\n\nMy model was just a simple target encoding of the building blocks and building block pairs. 6 features, plus the protein. Importantly, for bb2 and bb3; and also for bb1+2 and bb1+3, I rearranged the target encoding results to be bb2/3_max and bb2/3min, and likewise bb1+2_max and bb1+3_min.\n\nAnyways, really the impact of that wasn't about clever feature engineering or not so clever. The issue was the model did too well! Target encoding by itself was so far ahead of ECFP, which spelled trouble for the competition. It was roughly on par with the 1dCNN that came later (and ensembled well with it on public LB). Because of that, from early on, I was very pessimistic about the chances of any real generalization on non-shared BBs. I'm happy a few people managed to create robust models that were at least partially effective, but it remains that the majority of the data in the train set is best at predicting itself.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2912516": "Early on, a week or two(?) into the competition, I got ... 5th? on the leaderboard with just taking a public baseline model (KNN because I thought it might generalize better to non-shared BBs) and predicting ONLY shared BBs with my own model.\n\nMy model was just a simple target encoding of the building blocks and building block pairs. 6 features, plus the protein. Importantly, for bb2 and bb3; and also for bb1+2 and bb1+3, I rearranged the target encoding results to be bb2/3_max and bb2/3min, and likewise bb1+2_max and bb1+3_min.\n\nAnyways, really the impact of that wasn't about clever feature engineering or not so clever. The issue was the model did too well! Target encoding by itself was so far ahead of ECFP, which spelled trouble for the competition. It was roughly on par with the 1dCNN that came later (and ensembled well with it on public LB). Because of that, from early on, I was very pessimistic about the chances of any real generalization on non-shared BBs. I'm happy a few people managed to create robust models that were at least partially effective, but it remains that the majority of the data in the train set is best at predicting itself."
  }
}