{
  "id": 102078,
  "title": "Single channel training & class weight calculate for multilabel ",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102078",
  "author_name": "",
  "post_date": "2019-07-30T20:49:07.343599400Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<h2>compute class weights</h2>\n\n<p>To solve the imbalanced classes problem, we can use <code>compute_class_weight()</code> to calculate the class weights of multiclass, but it doesn't work for multilabel.\nIn this case,  'compute_sample_weight()' is a good choice.\n<code>\nfrom sklearn.utils.class_weight import compute_sample_weight\ny_train_weight = compute_sample_weight(\"balanced\", y_train)\n</code></p>\n\n<h1>Single channel training</h1>\n\n<p>The blood vessels usually have high contrast in the green channel, anyone tried to train the model by the single green channel?</p>",
  "messages": [
    {
      "id": "588630",
      "postDate": "07/30/2019 20:49:07",
      "content": "<h2>compute class weights</h2>\n\n<p>To solve the imbalanced classes problem, we can use <code>compute_class_weight()</code> to calculate the class weights of multiclass, but it doesn't work for multilabel.\nIn this case,  'compute_sample_weight()' is a good choice.\n<code>\nfrom sklearn.utils.class_weight import compute_sample_weight\ny_train_weight = compute_sample_weight(\"balanced\", y_train)\n</code></p>\n\n<h1>Single channel training</h1>\n\n<p>The blood vessels usually have high contrast in the green channel, anyone tried to train the model by the single green channel?</p>",
      "rawMarkdown": "## compute class weights\nTo solve the imbalanced classes problem, we can use `compute_class_weight()` to calculate the class weights of multiclass, but it doesn't work for multilabel.\nIn this case,  'compute_sample_weight()' is a good choice.\n```\nfrom sklearn.utils.class_weight import compute_sample_weight\ny_train_weight = compute_sample_weight(\"balanced\", y_train)\n```\n\n# Single channel training\nThe blood vessels usually have high contrast in the green channel, anyone tried to train the model by the single green channel?",
      "votes": null
    },
    {
      "id": "588890",
      "postDate": "07/31/2019 07:45:19",
      "content": "<p>Yes, I am training on single green channel, it \n- reduces memory consumption a lot\n- permits to load all of test data in kernel memory while submission, eradicating data generators and speeding kernel rerun time\n- does not lose any vital information in data\n- performs as well as 3 channel if not well. (Hint: it can perform better than 3 channel with proper preprocessing, provided all other factors remain same)\n- Allows for comparatively less parameters in model ( my current score is with model having ~1.1 million params, with no TTA, no optimized threshold and without pre-trained model.)</p>",
      "rawMarkdown": "Yes, I am training on single green channel, it \n- reduces memory consumption a lot\n- permits to load all of test data in kernel memory while submission, eradicating data generators and speeding kernel rerun time\n- does not lose any vital information in data\n- performs as well as 3 channel if not well. (Hint: it can perform better than 3 channel with proper preprocessing, provided all other factors remain same)\n- Allows for comparatively less parameters in model ( my current score is with model having ~1.1 million params, with no TTA, no optimized threshold and without pre-trained model.)",
      "votes": null
    },
    {
      "id": "588971",
      "postDate": "07/31/2019 09:38:07",
      "content": "<p>Agree! But, the loss reduced slowly. So I tried to histequalize green channel first and then copyed it to three channels, it was a little bit faster. At last, I got CV:0.9126 LB:0.757 (The efficientNetB3 with ImageNet weights)</p>",
      "rawMarkdown": "Agree! But, the loss reduced slowly. So I tried to histequalize green channel first and then copyed it to three channels, it was a little bit faster. At last, I got CV:0.9126 LB:0.757 (The efficientNetB3 with ImageNet weights)",
      "votes": null
    },
    {
      "id": "588983",
      "postDate": "07/31/2019 10:07:09",
      "content": "<p>I am also doing histequalization of green channel. But I think that copying is bit redundant, also it makes image bit darker, making it difficult to detect small microaneurysms, which are present in mild DR. I strongly believe that we can get &gt; 0.75(LB) with much smaller single model (in terms of parameters) and some training tricks. </p>",
      "rawMarkdown": "I am also doing histequalization of green channel. But I think that copying is bit redundant, also it makes image bit darker, making it difficult to detect small microaneurysms, which are present in mild DR. I strongly believe that we can get &gt; 0.75(LB) with much smaller single model (in terms of parameters) and some training tricks.",
      "votes": null
    },
    {
      "id": "588992",
      "postDate": "07/31/2019 10:19:06",
      "content": "<p>It makes sense. I will try to train the single-channel and then we can discuss it. Currently, I am using a part of the old dataset (2015) to pre-train the model, it may help you as well.\nDataset Link:\n<a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a></p>",
      "rawMarkdown": "It makes sense. I will try to train the single-channel and then we can discuss it. Currently, I am using a part of the old dataset (2015) to pre-train the model, it may help you as well.\nDataset Link:\nhttps://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized",
      "votes": null
    },
    {
      "id": "589003",
      "postDate": "07/31/2019 10:32:24",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "594925",
      "postDate": "08/08/2019 16:36:15",
      "content": "<p>Hey can you explain this a bit more. the 5 values this function outputs,how are we supposed to use this?Insert it in our loss function?</p>",
      "rawMarkdown": "Hey can you explain this a bit more. the 5 values this function outputs,how are we supposed to use this?Insert it in our loss function?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 588890,
      "author_name": "ashwan1",
      "author_url": "",
      "post_date": "07/31/2019 07:45:19",
      "content": "<p>Yes, I am training on single green channel, it \n- reduces memory consumption a lot\n- permits to load all of test data in kernel memory while submission, eradicating data generators and speeding kernel rerun time\n- does not lose any vital information in data\n- performs as well as 3 channel if not well. (Hint: it can perform better than 3 channel with proper preprocessing, provided all other factors remain same)\n- Allows for comparatively less parameters in model ( my current score is with model having ~1.1 million params, with no TTA, no optimized threshold and without pre-trained model.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 588971,
          "author_name": "xiewen112",
          "author_url": "",
          "post_date": "07/31/2019 09:38:07",
          "content": "<p>Agree! But, the loss reduced slowly. So I tried to histequalize green channel first and then copyed it to three channels, it was a little bit faster. At last, I got CV:0.9126 LB:0.757 (The efficientNetB3 with ImageNet weights)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 588983,
          "author_name": "ashwan1",
          "author_url": "",
          "post_date": "07/31/2019 10:07:09",
          "content": "<p>I am also doing histequalization of green channel. But I think that copying is bit redundant, also it makes image bit darker, making it difficult to detect small microaneurysms, which are present in mild DR. I strongly believe that we can get &gt; 0.75(LB) with much smaller single model (in terms of parameters) and some training tricks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 588992,
          "author_name": "xiewen112",
          "author_url": "",
          "post_date": "07/31/2019 10:19:06",
          "content": "<p>It makes sense. I will try to train the single-channel and then we can discuss it. Currently, I am using a part of the old dataset (2015) to pre-train the model, it may help you as well.\nDataset Link:\n<a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589003,
          "author_name": "ashwan1",
          "author_url": "",
          "post_date": "07/31/2019 10:32:24",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 594925,
      "author_name": "decentmakeover",
      "author_url": "",
      "post_date": "08/08/2019 16:36:15",
      "content": "<p>Hey can you explain this a bit more. the 5 values this function outputs,how are we supposed to use this?Insert it in our loss function?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "588630": "## compute class weights\nTo solve the imbalanced classes problem, we can use `compute_class_weight()` to calculate the class weights of multiclass, but it doesn't work for multilabel.\nIn this case,  'compute_sample_weight()' is a good choice.\n```\nfrom sklearn.utils.class_weight import compute_sample_weight\ny_train_weight = compute_sample_weight(\"balanced\", y_train)\n```\n\n# Single channel training\nThe blood vessels usually have high contrast in the green channel, anyone tried to train the model by the single green channel?",
    "588890": "Yes, I am training on single green channel, it \n- reduces memory consumption a lot\n- permits to load all of test data in kernel memory while submission, eradicating data generators and speeding kernel rerun time\n- does not lose any vital information in data\n- performs as well as 3 channel if not well. (Hint: it can perform better than 3 channel with proper preprocessing, provided all other factors remain same)\n- Allows for comparatively less parameters in model ( my current score is with model having ~1.1 million params, with no TTA, no optimized threshold and without pre-trained model.)",
    "588971": "Agree! But, the loss reduced slowly. So I tried to histequalize green channel first and then copyed it to three channels, it was a little bit faster. At last, I got CV:0.9126 LB:0.757 (The efficientNetB3 with ImageNet weights)",
    "588983": "I am also doing histequalization of green channel. But I think that copying is bit redundant, also it makes image bit darker, making it difficult to detect small microaneurysms, which are present in mild DR. I strongly believe that we can get &gt; 0.75(LB) with much smaller single model (in terms of parameters) and some training tricks.",
    "588992": "It makes sense. I will try to train the single-channel and then we can discuss it. Currently, I am using a part of the old dataset (2015) to pre-train the model, it may help you as well.\nDataset Link:\nhttps://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized",
    "589003": "Thanks!",
    "594925": "Hey can you explain this a bit more. the 5 values this function outputs,how are we supposed to use this?Insert it in our loss function?"
  },
  "source": "meta"
}