{
  "id": 233475,
  "title": "[external dataset] 16bit 1024x1024 HPA exteranl training data with cell instance mask",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/233475",
  "author_name": "seefun",
  "post_date": "2021-04-19T13:49:43.291000",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<h2>public dataset:</h2>\n<p><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839\" target=\"_blank\">Here</a> is public training data with 1024x1024 size ,16bit depth, with cell level mask.</p>\n<h2>external dataset:</h2>\n<p><a href=\"https://www.kaggle.com/seefun/hpa2021-extrain\" target=\"_blank\">Here</a> is also some external dataset (from minority class, which helps balancing the data), 1024x1024, 16bit, with cell level mask. In this dataset, we download the data and process them using <code>externalData.ipynb</code> and <code>MakeDataset.ipynb</code>;  <code>train_all.csv</code> concat the public training data and this external data</p>\n<h2>result:</h2>\n<p>Combine these two dataset. I could get CV 0.88 mAP for multi-task classification using ResNeSt50d(512x512,concat pooling+multi-sample dropout), with <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/233479\" target=\"_blank\">strong augmentation</a> (with HPA mixup), EMA, and LabelSmoothing, which costs 80 epoches with RangerLARS (RAdam+Lookahead+LARS(+GC)) optimizer. Using bigger model and training for more epoches could help increase the multi-task classification mAP.</p>",
  "messages": [
    {
      "id": 1278027,
      "postDate": "2021-04-19T13:49:43.290Z",
      "content": "<h2>public dataset:</h2>\n<p><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839\" target=\"_blank\">Here</a> is public training data with 1024x1024 size ,16bit depth, with cell level mask.</p>\n<h2>external dataset:</h2>\n<p><a href=\"https://www.kaggle.com/seefun/hpa2021-extrain\" target=\"_blank\">Here</a> is also some external dataset (from minority class, which helps balancing the data), 1024x1024, 16bit, with cell level mask. In this dataset, we download the data and process them using <code>externalData.ipynb</code> and <code>MakeDataset.ipynb</code>;  <code>train_all.csv</code> concat the public training data and this external data</p>\n<h2>result:</h2>\n<p>Combine these two dataset. I could get CV 0.88 mAP for multi-task classification using ResNeSt50d(512x512,concat pooling+multi-sample dropout), with <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/233479\" target=\"_blank\">strong augmentation</a> (with HPA mixup), EMA, and LabelSmoothing, which costs 80 epoches with RangerLARS (RAdam+Lookahead+LARS(+GC)) optimizer. Using bigger model and training for more epoches could help increase the multi-task classification mAP.</p>",
      "rawMarkdown": "## public dataset:\n[Here](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839) is public training data with 1024x1024 size ,16bit depth, with cell level mask.\n\n## external dataset:\n[Here](https://www.kaggle.com/seefun/hpa2021-extrain) is also some external dataset (from minority class, which helps balancing the data), 1024x1024, 16bit, with cell level mask. In this dataset, we download the data and process them using `externalData.ipynb` and `MakeDataset.ipynb`;  `train_all.csv` concat the public training data and this external data\n\n## result:\nCombine these two dataset. I could get CV 0.88 mAP for multi-task classification using ResNeSt50d(512x512,concat pooling+multi-sample dropout), with [strong augmentation](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/233479) (with HPA mixup), EMA, and LabelSmoothing, which costs 80 epoches with RangerLARS (RAdam+Lookahead+LARS(+GC)) optimizer. Using bigger model and training for more epoches could help increase the multi-task classification mAP.",
      "votes": 11
    },
    {
      "id": 1286492,
      "postDate": "2021-04-28T04:16:28.693Z",
      "content": "<p>Great work!~</p>",
      "rawMarkdown": "Great work!~"
    }
  ],
  "comments": [
    {
      "id": 1286492,
      "author_name": "howard.G",
      "author_url": "",
      "post_date": "2021-04-28T04:16:28.693000",
      "content": "<p>Great work!~</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1278027": "## public dataset:\n[Here](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839) is public training data with 1024x1024 size ,16bit depth, with cell level mask.\n\n## external dataset:\n[Here](https://www.kaggle.com/seefun/hpa2021-extrain) is also some external dataset (from minority class, which helps balancing the data), 1024x1024, 16bit, with cell level mask. In this dataset, we download the data and process them using `externalData.ipynb` and `MakeDataset.ipynb`;  `train_all.csv` concat the public training data and this external data\n\n## result:\nCombine these two dataset. I could get CV 0.88 mAP for multi-task classification using ResNeSt50d(512x512,concat pooling+multi-sample dropout), with [strong augmentation](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/233479) (with HPA mixup), EMA, and LabelSmoothing, which costs 80 epoches with RangerLARS (RAdam+Lookahead+LARS(+GC)) optimizer. Using bigger model and training for more epoches could help increase the multi-task classification mAP.",
    "1286492": "Great work!~"
  }
}