{
  "id": 222300,
  "title": "Can external data improve the performance？",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/222300",
  "author_name": "seefun",
  "post_date": "2021-02-26T08:34:26.485000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We can use huge <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">public external data</a> (more than 2TB) in this competition. But, will the use of additional data lead to significant improvements？ Could you share it？</p>\n<p>And should we add  some <code>negative</code> class from external data into our training set?</p>",
  "messages": [
    {
      "id": 1218836,
      "postDate": "2021-02-26T08:34:26.487Z",
      "content": "<p>We can use huge <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">public external data</a> (more than 2TB) in this competition. But, will the use of additional data lead to significant improvements？ Could you share it？</p>\n<p>And should we add  some <code>negative</code> class from external data into our training set?</p>",
      "rawMarkdown": "We can use huge [public external data](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg) (more than 2TB) in this competition. But, will the use of additional data lead to significant improvements？ Could you share it？\n\nAnd should we add  some ```negative``` class from external data into our training set?",
      "votes": 5
    },
    {
      "id": 1219370,
      "postDate": "2021-02-26T18:45:51.860Z",
      "content": "<p>I have plans to add \"some\" of the additional data for the under-represented classes.  The host recently updated the list on the public site so we can avoid over-lap.</p>\n<p>Will it help ?  My current best model scores 0.310.   IMO - Added data is not likely to move this model into the gold range if I just grab a big pile, but should help a little if I am selective and address those classes that a confusion matrix tells me are troublesome.</p>\n<p>A huge download of 16 bit images also available for the current train set - will that help?</p>\n<p>The latest change to my model has been running for close to 30 hours of training - right now I don't feel I can afford the compute that comes with extra data.  I also don't feel like I have a good grasp on how to best handle weak training data - more of the same will not be helpful.</p>\n<p>When I segmented the current training data I got almost 2 million single cell images.  With my first pass of cleaning up that mess I am down to around 400K of single cell images - so at first glance I am thinking I have way too many \"negative\" class already.  </p>\n<p>It seems there are two types of negative classes - <br>\ntype 1)  The full external set has 35 classes - the hosts did some magic and gave us only 18 - with the other's that are MIA not part of our competition put into the negative class.  <br>\ntype 2) single cell images that have no stain (no Green).</p>\n<p>IMO most of our negative's will be type 2.  When my model score passes 0.50 than I will start to worry about the type 1.</p>",
      "rawMarkdown": "I have plans to add \"some\" of the additional data for the under-represented classes.  The host recently updated the list on the public site so we can avoid over-lap.\n\nWill it help ?  My current best model scores 0.310.   IMO - Added data is not likely to move this model into the gold range if I just grab a big pile, but should help a little if I am selective and address those classes that a confusion matrix tells me are troublesome.\n\nA huge download of 16 bit images also available for the current train set - will that help?\n\nThe latest change to my model has been running for close to 30 hours of training - right now I don't feel I can afford the compute that comes with extra data.  I also don't feel like I have a good grasp on how to best handle weak training data - more of the same will not be helpful.\n\nWhen I segmented the current training data I got almost 2 million single cell images.  With my first pass of cleaning up that mess I am down to around 400K of single cell images - so at first glance I am thinking I have way too many \"negative\" class already.  \n\nIt seems there are two types of negative classes - \ntype 1)  The full external set has 35 classes - the hosts did some magic and gave us only 18 - with the other's that are MIA not part of our competition put into the negative class.  \ntype 2) single cell images that have no stain (no Green).\n\nIMO most of our negative's will be type 2.  When my model score passes 0.50 than I will start to worry about the type 1.\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 1219730,
          "postDate": "2021-02-27T07:20:47.660Z",
          "content": "<p>Thank you very much for sharing！</p>",
          "rawMarkdown": "Thank you very much for sharing！"
        },
        {
          "id": 1232090,
          "postDate": "2021-03-09T13:34:33.977Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>,<br>\nI don't think classes that are not present in the 18 classes are Negative. The hosts mentioned that some classes present in the public dataset have been merged together. For example, in <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns#9.-Actin-filaments\" target=\"_blank\">this</a> notebook, it is mentioned that the Actin filaments class</p>\n<blockquote>\n  <p>includes Actin filaments and Focal adhesion sites.</p>\n</blockquote>\n<p>Focal adhesion sites are not a part of the 18 classes, but they instead merged into the Actin filament class.</p>",
          "rawMarkdown": "Hey @pcjimmmy,\nI don't think classes that are not present in the 18 classes are Negative. The hosts mentioned that some classes present in the public dataset have been merged together. For example, in [this](https://www.kaggle.com/lnhtrang/single-cell-patterns#9.-Actin-filaments) notebook, it is mentioned that the Actin filaments class\n\n> includes Actin filaments and Focal adhesion sites.\n \nFocal adhesion sites are not a part of the 18 classes, but they instead merged into the Actin filament class.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1221992,
      "postDate": "2021-03-01T13:09:57.797Z",
      "content": "<p>A note from the competition authors on one of the other threads where I asked about the public data download speeds and how cumbersome it was.</p>\n<hr>\n<blockquote>\n  <p><strong><em>\"Hi! Thanks for your question. Yes - this is still in progress. We'll be making a trimmed version of the public data available as a Kaggle dataset, which should be quicker to access via Kaggle notebooks (and also be a smaller download elsewhere).\"</em></strong></p>\n</blockquote>\n<hr>\n<p>I know this isn't an answer to your question, but it seemed relevant. Once they have made it a Kaggle dataset I will be adding it to my training regime to benchmark comparative performance.</p>",
      "rawMarkdown": "A note from the competition authors on one of the other threads where I asked about the public data download speeds and how cumbersome it was.\n\n---\n\n> ***\"Hi! Thanks for your question. Yes - this is still in progress. We'll be making a trimmed version of the public data available as a Kaggle dataset, which should be quicker to access via Kaggle notebooks (and also be a smaller download elsewhere).\"***\n\n---\n\nI know this isn't an answer to your question, but it seemed relevant. Once they have made it a Kaggle dataset I will be adding it to my training regime to benchmark comparative performance.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1219370,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-26T18:45:51.860000",
      "content": "<p>I have plans to add \"some\" of the additional data for the under-represented classes.  The host recently updated the list on the public site so we can avoid over-lap.</p>\n<p>Will it help ?  My current best model scores 0.310.   IMO - Added data is not likely to move this model into the gold range if I just grab a big pile, but should help a little if I am selective and address those classes that a confusion matrix tells me are troublesome.</p>\n<p>A huge download of 16 bit images also available for the current train set - will that help?</p>\n<p>The latest change to my model has been running for close to 30 hours of training - right now I don't feel I can afford the compute that comes with extra data.  I also don't feel like I have a good grasp on how to best handle weak training data - more of the same will not be helpful.</p>\n<p>When I segmented the current training data I got almost 2 million single cell images.  With my first pass of cleaning up that mess I am down to around 400K of single cell images - so at first glance I am thinking I have way too many \"negative\" class already.  </p>\n<p>It seems there are two types of negative classes - <br>\ntype 1)  The full external set has 35 classes - the hosts did some magic and gave us only 18 - with the other's that are MIA not part of our competition put into the negative class.  <br>\ntype 2) single cell images that have no stain (no Green).</p>\n<p>IMO most of our negative's will be type 2.  When my model score passes 0.50 than I will start to worry about the type 1.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1219730,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2021-02-27T07:20:47.660000",
          "content": "<p>Thank you very much for sharing！</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1232090,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "2021-03-09T13:34:33.977000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>,<br>\nI don't think classes that are not present in the 18 classes are Negative. The hosts mentioned that some classes present in the public dataset have been merged together. For example, in <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns#9.-Actin-filaments\" target=\"_blank\">this</a> notebook, it is mentioned that the Actin filaments class</p>\n<blockquote>\n  <p>includes Actin filaments and Focal adhesion sites.</p>\n</blockquote>\n<p>Focal adhesion sites are not a part of the 18 classes, but they instead merged into the Actin filament class.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1221992,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-03-01T13:09:57.797000",
      "content": "<p>A note from the competition authors on one of the other threads where I asked about the public data download speeds and how cumbersome it was.</p>\n<hr>\n<blockquote>\n  <p><strong><em>\"Hi! Thanks for your question. Yes - this is still in progress. We'll be making a trimmed version of the public data available as a Kaggle dataset, which should be quicker to access via Kaggle notebooks (and also be a smaller download elsewhere).\"</em></strong></p>\n</blockquote>\n<hr>\n<p>I know this isn't an answer to your question, but it seemed relevant. Once they have made it a Kaggle dataset I will be adding it to my training regime to benchmark comparative performance.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1218836": "We can use huge [public external data](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg) (more than 2TB) in this competition. But, will the use of additional data lead to significant improvements？ Could you share it？\n\nAnd should we add  some ```negative``` class from external data into our training set?",
    "1219370": "I have plans to add \"some\" of the additional data for the under-represented classes.  The host recently updated the list on the public site so we can avoid over-lap.\n\nWill it help ?  My current best model scores 0.310.   IMO - Added data is not likely to move this model into the gold range if I just grab a big pile, but should help a little if I am selective and address those classes that a confusion matrix tells me are troublesome.\n\nA huge download of 16 bit images also available for the current train set - will that help?\n\nThe latest change to my model has been running for close to 30 hours of training - right now I don't feel I can afford the compute that comes with extra data.  I also don't feel like I have a good grasp on how to best handle weak training data - more of the same will not be helpful.\n\nWhen I segmented the current training data I got almost 2 million single cell images.  With my first pass of cleaning up that mess I am down to around 400K of single cell images - so at first glance I am thinking I have way too many \"negative\" class already.  \n\nIt seems there are two types of negative classes - \ntype 1)  The full external set has 35 classes - the hosts did some magic and gave us only 18 - with the other's that are MIA not part of our competition put into the negative class.  \ntype 2) single cell images that have no stain (no Green).\n\nIMO most of our negative's will be type 2.  When my model score passes 0.50 than I will start to worry about the type 1.\n\n",
    "1221992": "A note from the competition authors on one of the other threads where I asked about the public data download speeds and how cumbersome it was.\n\n---\n\n> ***\"Hi! Thanks for your question. Yes - this is still in progress. We'll be making a trimmed version of the public data available as a Kaggle dataset, which should be quicker to access via Kaggle notebooks (and also be a smaller download elsewhere).\"***\n\n---\n\nI know this isn't an answer to your question, but it seemed relevant. Once they have made it a Kaggle dataset I will be adding it to my training regime to benchmark comparative performance."
  }
}