{
  "id": 457511,
  "title": "Should we use a blended data frame or any its combination with custom solution in final submission?",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/457511",
  "author_name": "",
  "post_date": "2023-11-25T07:38:15.397290500Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I wonder if the blended data frame can be useful in final submission. </p>",
  "messages": [
    {
      "id": "2537481",
      "postDate": "11/25/2023 07:38:15",
      "content": "<p>I wonder if the blended data frame can be useful in final submission. </p>",
      "rawMarkdown": "I wonder if the blended data frame can be useful in final submission.",
      "votes": null
    },
    {
      "id": "2538226",
      "postDate": "11/25/2023 22:40:03",
      "content": "<p>While I have enjoyed learning a few new things in this competition, IMO it's one of the poorer datasets I have seen in my 7 years on kaggle.  614 rows, with only 34 of them of the cell type that we need to predict, and than only predicting 255.  Some have made the argument that having 18K labels helps makes some statistical sense, but I am not buying that bridge.</p>\n<p>Reading the background literature on this measurement process there are huge batch effects, so we are modeling small data set with lots of known noise.  Not going to even think about the label - predict a probability value who's sign is based on comparison to a control.</p>\n<p>With that in mind I placing both of my bets on use of massive model with both blended data and lots of augmentation.</p>",
      "rawMarkdown": "While I have enjoyed learning a few new things in this competition, IMO it's one of the poorer datasets I have seen in my 7 years on kaggle.  614 rows, with only 34 of them of the cell type that we need to predict, and than only predicting 255.  Some have made the argument that having 18K labels helps makes some statistical sense, but I am not buying that bridge.\n\nReading the background literature on this measurement process there are huge batch effects, so we are modeling small data set with lots of known noise.  Not going to even think about the label - predict a probability value who's sign is based on comparison to a control.\n\nWith that in mind I placing both of my bets on use of massive model with both blended data and lots of augmentation.",
      "votes": null
    },
    {
      "id": "2539039",
      "postDate": "11/26/2023 16:27:10",
      "content": "<p>totally agree, one of the worst dataset ever</p>",
      "rawMarkdown": "totally agree, one of the worst dataset ever",
      "votes": null
    },
    {
      "id": "2544665",
      "postDate": "12/01/2023 02:23:29",
      "content": "<p>Congratulations on winning the 3rd prize! Looks like you've made the right decision in the end. I look forward to seeing what you've ended up submitting.</p>",
      "rawMarkdown": "Congratulations on winning the 3rd prize! Looks like you've made the right decision in the end. I look forward to seeing what you've ended up submitting.",
      "votes": null
    },
    {
      "id": "2545188",
      "postDate": "12/01/2023 10:55:18",
      "content": "<p>Thanks, my solution and notebook are available. </p>",
      "rawMarkdown": "Thanks, my solution and notebook are available.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2538226,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "11/25/2023 22:40:03",
      "content": "<p>While I have enjoyed learning a few new things in this competition, IMO it's one of the poorer datasets I have seen in my 7 years on kaggle.  614 rows, with only 34 of them of the cell type that we need to predict, and than only predicting 255.  Some have made the argument that having 18K labels helps makes some statistical sense, but I am not buying that bridge.</p>\n<p>Reading the background literature on this measurement process there are huge batch effects, so we are modeling small data set with lots of known noise.  Not going to even think about the label - predict a probability value who's sign is based on comparison to a control.</p>\n<p>With that in mind I placing both of my bets on use of massive model with both blended data and lots of augmentation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2539039,
          "author_name": "ermalossi",
          "author_url": "",
          "post_date": "11/26/2023 16:27:10",
          "content": "<p>totally agree, one of the worst dataset ever</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2544665,
      "author_name": "frenio",
      "author_url": "",
      "post_date": "12/01/2023 02:23:29",
      "content": "<p>Congratulations on winning the 3rd prize! Looks like you've made the right decision in the end. I look forward to seeing what you've ended up submitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2545188,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "12/01/2023 10:55:18",
          "content": "<p>Thanks, my solution and notebook are available. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2537481": "I wonder if the blended data frame can be useful in final submission.",
    "2538226": "While I have enjoyed learning a few new things in this competition, IMO it's one of the poorer datasets I have seen in my 7 years on kaggle.  614 rows, with only 34 of them of the cell type that we need to predict, and than only predicting 255.  Some have made the argument that having 18K labels helps makes some statistical sense, but I am not buying that bridge.\n\nReading the background literature on this measurement process there are huge batch effects, so we are modeling small data set with lots of known noise.  Not going to even think about the label - predict a probability value who's sign is based on comparison to a control.\n\nWith that in mind I placing both of my bets on use of massive model with both blended data and lots of augmentation.",
    "2539039": "totally agree, one of the worst dataset ever",
    "2544665": "Congratulations on winning the 3rd prize! Looks like you've made the right decision in the end. I look forward to seeing what you've ended up submitting.",
    "2545188": "Thanks, my solution and notebook are available."
  },
  "source": "meta"
}