{
  "id": 242606,
  "title": "Highly Imbalanced Class Samples",
  "url": "/competitions/siim-covid19-detection/discussion/242606",
  "author_name": "",
  "post_date": "2021-05-29T20:46:09.942147100Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<pre><code>typical          3007\nnegative         1736\nindeterminate    1108\natypical          483\n</code></pre>\n<p>Whoa!! That's pretty imbalanced. How does a professional data scientist handle it?</p>\n<p>Primary Ideas :</p>\n<ul>\n<li>Use of Class Weights</li>\n<li>Apply Heavy Augmentations</li>\n<li>Upsampling &amp; Downsampling</li>\n<li>Focal Loss for Training</li>\n</ul>",
  "messages": [
    {
      "id": "1328004",
      "postDate": "05/29/2021 20:46:09",
      "content": "<pre><code>typical          3007\nnegative         1736\nindeterminate    1108\natypical          483\n</code></pre>\n<p>Whoa!! That's pretty imbalanced. How does a professional data scientist handle it?</p>\n<p>Primary Ideas :</p>\n<ul>\n<li>Use of Class Weights</li>\n<li>Apply Heavy Augmentations</li>\n<li>Upsampling &amp; Downsampling</li>\n<li>Focal Loss for Training</li>\n</ul>",
      "rawMarkdown": "```\ntypical          3007\nnegative         1736\nindeterminate    1108\natypical          483\n```\n\nWhoa!! That's pretty imbalanced. How does a professional data scientist handle it?\n\nPrimary Ideas :\n- Use of Class Weights\n- Apply Heavy Augmentations\n- Upsampling & Downsampling\n- Focal Loss for Training",
      "votes": null
    },
    {
      "id": "1333082",
      "postDate": "06/02/2021 13:30:05",
      "content": "<p>I am not a professional data scientist yet 😁, but here is how I handle class imbalance -</p>\n<p>I prepare a common validation set out of my dataset and try out each and every method you have mentioned. But, I only didn't get how applying heavy augmentations will help in imbalanced dataset. It would be so amazing if you could explain it or sharer some resources regarding the same. Lastly, I would like to share one more approach I am aware of.</p>\n<p>If you have some classes that have very small number of instances, you can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called small_class for eg) so that your main classifier will classify small_class with all other big classes in the dataset. And if your main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. This technique can give you accuracy boosts are now main classifier does not need to deal with very small classes, and insted small_classifier will be looking just at these small classes.</p>\n<p>In this competition we will need to train an object detector, so perhaps afforementioned method will need to be done with object detectors. </p>",
      "rawMarkdown": "I am not a professional data scientist yet 😁, but here is how I handle class imbalance -\n\nI prepare a common validation set out of my dataset and try out each and every method you have mentioned. But, I only didn't get how applying heavy augmentations will help in imbalanced dataset. It would be so amazing if you could explain it or sharer some resources regarding the same. Lastly, I would like to share one more approach I am aware of.\n\nIf you have some classes that have very small number of instances, you can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called small_class for eg) so that your main classifier will classify small_class with all other big classes in the dataset. And if your main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. This technique can give you accuracy boosts are now main classifier does not need to deal with very small classes, and insted small_classifier will be looking just at these small classes.\n\nIn this competition we will need to train an object detector, so perhaps afforementioned method will need to be done with object detectors.",
      "votes": null
    },
    {
      "id": "1334321",
      "postDate": "06/03/2021 12:20:48",
      "content": "<p>Let's say if we define an augmentation pipeline <code>P(img)</code> of <code>RandomBrightnessContrast</code>,<code>Horizontal Flipping</code> (As Lungs are Symmetric) and <code>RandomRotation</code> (-15 to +15 degrees) then we can handle the class imbalances of the following distribution as follows:</p>\n<pre><code>typical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n</code></pre>\n<p>General Augmentations : </p>\n<pre><code>for each_img in original_img :  apply P(img)\n</code></pre>\n<p>Now, for specific augmentations:<br>\nTake the max sampled class : <code>class_max = A</code>. Now the objective is to bring the other classes to <code>n(class_max) = n(A)</code>.</p>\n<pre><code>for each_class in remaining classes : [B,C,D] -\nPerform P(original_img) for randomly chosen [ n(class_max) - n(each_class) ] images in each_class\n</code></pre>\n<p>in the end you find that <code>n(each_class) = n(class_max)</code>. Hope that sounds clear!</p>",
      "rawMarkdown": "Let's say if we define an augmentation pipeline `P(img)` of `RandomBrightnessContrast`,`Horizontal Flipping` (As Lungs are Symmetric) and `RandomRotation` (-15 to +15 degrees) then we can handle the class imbalances of the following distribution as follows:\n```\ntypical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n```\nGeneral Augmentations : \n```\nfor each_img in original_img :  apply P(img)\n```\nNow, for specific augmentations:\nTake the max sampled class : `class_max = A`. Now the objective is to bring the other classes to `n(class_max) = n(A)`.\n```\nfor each_class in remaining classes : [B,C,D] -\nPerform P(original_img) for randomly chosen [ n(class_max) - n(each_class) ] images in each_class\n```\n\nin the end you find that `n(each_class) = n(class_max)`. Hope that sounds clear!",
      "votes": null
    },
    {
      "id": "1334340",
      "postDate": "06/03/2021 12:36:15",
      "content": "<p>Summarizing your approach : </p>\n<pre><code>typical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n</code></pre>\n<p>Your idea is something like this if I understand it quite clearly.</p>\n<p>Let's assume we have <code>model (X)</code> (say Yolov5) which is SOTA ( State of the Art ) for a given ML Task.</p>\n<p>Approach to conquer data imbalance pbm : </p>\n<p>Take the minimum sampled class <code>class_min = D</code> after that :</p>\n<pre><code>Recursively combine the classes in ascending order of their samples in each step??\n</code></pre>\n<p>I don't exactly follow, if we follow te above algorithm,</p>\n<p>Step 1: Classifier for <code>indeterminate (C) 1108 vs. atypical (D) 483</code> need not be balanced itself. (approx 2.3:1 ratio).</p>\n<p>However I do understand the idea that as we gradually go up the ladder, it should be more balanced.</p>\n<p>For ex.</p>\n<p>Step 2  : Classifier for <code>CD[indeterminate (C) 1108 U atypical (D) 483] (1591) vs.  negative (B)  1736</code> is indeed more balanced. (approx 1:1.09 ratio).</p>\n<p>Step 3 : Classifier for <code>BCD (3327) vs. typical (A) 3007</code> is also more balanced. (approx 1.11:1 ratio).</p>",
      "rawMarkdown": "Summarizing your approach : \n```\ntypical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n```\n\nYour idea is something like this if I understand it quite clearly.\n\nLet's assume we have `model (X)` (say Yolov5) which is SOTA ( State of the Art ) for a given ML Task.\n\nApproach to conquer data imbalance pbm : \n\nTake the minimum sampled class `class_min = D` after that :\n\n```\nRecursively combine the classes in ascending order of their samples in each step??\n```\n\nI don't exactly follow, if we follow te above algorithm,\n\nStep 1: Classifier for `indeterminate (C) 1108 vs. atypical (D) 483` need not be balanced itself. (approx 2.3:1 ratio).\n\nHowever I do understand the idea that as we gradually go up the ladder, it should be more balanced.\n\nFor ex.\n\nStep 2  : Classifier for `CD[indeterminate (C) 1108 U atypical (D) 483] (1591) vs.  negative (B)  1736` is indeed more balanced. (approx 1:1.09 ratio).\n\nStep 3 : Classifier for `BCD (3327) vs. typical (A) 3007` is also more balanced. (approx 1.11:1 ratio).",
      "votes": null
    },
    {
      "id": "1334519",
      "postDate": "06/03/2021 15:15:28",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/farhanhaikhan\" target=\"_blank\">@farhanhaikhan</a> for explaining it. I got it very clearly now 👍</p>",
      "rawMarkdown": "Thank you @farhanhaikhan for explaining it. I got it very clearly now 👍",
      "votes": null
    },
    {
      "id": "1334565",
      "postDate": "06/03/2021 15:42:32",
      "content": "<p>Yes I think you got the idea, and just to make sure, I will explain it in more depth.</p>\n<p>Regarding my approach, it is as follows -</p>\n<p>For eg we have a dataset with following class distribution -</p>\n<p>A:120  B:100  C:90  D:80  E:20  F:15  G:10  H:8  I:5</p>\n<p>Now here </p>\n<blockquote>\n  <p>Some classes that have very small number of instances are = E F G H I</p>\n</blockquote>\n<p>After that</p>\n<blockquote>\n  <p>You can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called X for eg) so that your main classifier will classify X with all other big classes (A, B, C, D).</p>\n</blockquote>\n<p>So the new distribution for your main classifier becomes, A:120  B:100  C:90  D:80  X:58 (better class balance). And distribution of small_classifier becomes, E:20  F:15  G:10  H:8  I:5 (better class balance)</p>\n<blockquote>\n  <p>If your main classifier encounters any instance of X class, it will pass it to small_classifier, which will predict the actual class (E, F, G, H, I) for the small_class instance. Now main classifier does not need to deal with very small classes, and instead small_classifier will be looking just at these small classes.</p>\n</blockquote>\n<p>Now, in this competition -</p>\n<blockquote>\n  <p>typical (A)         3007<br>\n  negative (B)       1736<br>\n  indeterminate (C)    1108<br>\n  atypical  (D)        483</p>\n</blockquote>\n<p>we might try something like as you mentioned -</p>\n<blockquote>\n  <p>Main Classifier for BCD (3327) vs. A (3007) (approx 1.11:1 ratio). And small_classifier for B (1736) vs C (1108) vs D (483) (approx 2.3:2:1 ratio).</p>\n</blockquote>\n<p>OR</p>\n<blockquote>\n  <p>Main Classifier for CD (3327) vs B (1736) vs A (3007) (approx 2:1:1.8 ratio). And small_classifier for C (1108) vs D (483) (approx 2.3:1 ratio).</p>\n</blockquote>",
      "rawMarkdown": "Yes I think you got the idea, and just to make sure, I will explain it in more depth.\n\nRegarding my approach, it is as follows -\n\nFor eg we have a dataset with following class distribution -\n\nA:120  B:100  C:90  D:80  E:20  F:15  G:10  H:8  I:5\n\nNow here \n\n> Some classes that have very small number of instances are = E F G H I\n\nAfter that\n\n> You can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called X for eg) so that your main classifier will classify X with all other big classes (A, B, C, D).\n\nSo the new distribution for your main classifier becomes, A:120  B:100  C:90  D:80  X:58 (better class balance). And distribution of small_classifier becomes, E:20  F:15  G:10  H:8  I:5 (better class balance)\n\n> If your main classifier encounters any instance of X class, it will pass it to small_classifier, which will predict the actual class (E, F, G, H, I) for the small_class instance. Now main classifier does not need to deal with very small classes, and instead small_classifier will be looking just at these small classes.\n\nNow, in this competition -\n> typical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n\nwe might try something like as you mentioned -\n\n> Main Classifier for BCD (3327) vs. A (3007) (approx 1.11:1 ratio). And small_classifier for B (1736) vs C (1108) vs D (483) (approx 2.3:2:1 ratio).\n\nOR\n\n> Main Classifier for CD (3327) vs B (1736) vs A (3007) (approx 2:1:1.8 ratio). And small_classifier for C (1108) vs D (483) (approx 2.3:1 ratio).",
      "votes": null
    },
    {
      "id": "1334675",
      "postDate": "06/03/2021 17:20:36",
      "content": "<p>I'm sure that was very intuitive for everyone arriving at this discussion. Thanks for the fruitful explanation! Glad to hear that I could help too! :) Happy Kaggling!</p>",
      "rawMarkdown": "I'm sure that was very intuitive for everyone arriving at this discussion. Thanks for the fruitful explanation! Glad to hear that I could help too! :) Happy Kaggling!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1333082,
      "author_name": "devashishprasad",
      "author_url": "",
      "post_date": "06/02/2021 13:30:05",
      "content": "<p>I am not a professional data scientist yet 😁, but here is how I handle class imbalance -</p>\n<p>I prepare a common validation set out of my dataset and try out each and every method you have mentioned. But, I only didn't get how applying heavy augmentations will help in imbalanced dataset. It would be so amazing if you could explain it or sharer some resources regarding the same. Lastly, I would like to share one more approach I am aware of.</p>\n<p>If you have some classes that have very small number of instances, you can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called small_class for eg) so that your main classifier will classify small_class with all other big classes in the dataset. And if your main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. This technique can give you accuracy boosts are now main classifier does not need to deal with very small classes, and insted small_classifier will be looking just at these small classes.</p>\n<p>In this competition we will need to train an object detector, so perhaps afforementioned method will need to be done with object detectors. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1334321,
          "author_name": "farhanhaikhan",
          "author_url": "",
          "post_date": "06/03/2021 12:20:48",
          "content": "<p>Let's say if we define an augmentation pipeline <code>P(img)</code> of <code>RandomBrightnessContrast</code>,<code>Horizontal Flipping</code> (As Lungs are Symmetric) and <code>RandomRotation</code> (-15 to +15 degrees) then we can handle the class imbalances of the following distribution as follows:</p>\n<pre><code>typical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n</code></pre>\n<p>General Augmentations : </p>\n<pre><code>for each_img in original_img :  apply P(img)\n</code></pre>\n<p>Now, for specific augmentations:<br>\nTake the max sampled class : <code>class_max = A</code>. Now the objective is to bring the other classes to <code>n(class_max) = n(A)</code>.</p>\n<pre><code>for each_class in remaining classes : [B,C,D] -\nPerform P(original_img) for randomly chosen [ n(class_max) - n(each_class) ] images in each_class\n</code></pre>\n<p>in the end you find that <code>n(each_class) = n(class_max)</code>. Hope that sounds clear!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334340,
          "author_name": "farhanhaikhan",
          "author_url": "",
          "post_date": "06/03/2021 12:36:15",
          "content": "<p>Summarizing your approach : </p>\n<pre><code>typical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n</code></pre>\n<p>Your idea is something like this if I understand it quite clearly.</p>\n<p>Let's assume we have <code>model (X)</code> (say Yolov5) which is SOTA ( State of the Art ) for a given ML Task.</p>\n<p>Approach to conquer data imbalance pbm : </p>\n<p>Take the minimum sampled class <code>class_min = D</code> after that :</p>\n<pre><code>Recursively combine the classes in ascending order of their samples in each step??\n</code></pre>\n<p>I don't exactly follow, if we follow te above algorithm,</p>\n<p>Step 1: Classifier for <code>indeterminate (C) 1108 vs. atypical (D) 483</code> need not be balanced itself. (approx 2.3:1 ratio).</p>\n<p>However I do understand the idea that as we gradually go up the ladder, it should be more balanced.</p>\n<p>For ex.</p>\n<p>Step 2  : Classifier for <code>CD[indeterminate (C) 1108 U atypical (D) 483] (1591) vs.  negative (B)  1736</code> is indeed more balanced. (approx 1:1.09 ratio).</p>\n<p>Step 3 : Classifier for <code>BCD (3327) vs. typical (A) 3007</code> is also more balanced. (approx 1.11:1 ratio).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334519,
          "author_name": "devashishprasad",
          "author_url": "",
          "post_date": "06/03/2021 15:15:28",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/farhanhaikhan\" target=\"_blank\">@farhanhaikhan</a> for explaining it. I got it very clearly now 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334565,
          "author_name": "devashishprasad",
          "author_url": "",
          "post_date": "06/03/2021 15:42:32",
          "content": "<p>Yes I think you got the idea, and just to make sure, I will explain it in more depth.</p>\n<p>Regarding my approach, it is as follows -</p>\n<p>For eg we have a dataset with following class distribution -</p>\n<p>A:120  B:100  C:90  D:80  E:20  F:15  G:10  H:8  I:5</p>\n<p>Now here </p>\n<blockquote>\n  <p>Some classes that have very small number of instances are = E F G H I</p>\n</blockquote>\n<p>After that</p>\n<blockquote>\n  <p>You can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called X for eg) so that your main classifier will classify X with all other big classes (A, B, C, D).</p>\n</blockquote>\n<p>So the new distribution for your main classifier becomes, A:120  B:100  C:90  D:80  X:58 (better class balance). And distribution of small_classifier becomes, E:20  F:15  G:10  H:8  I:5 (better class balance)</p>\n<blockquote>\n  <p>If your main classifier encounters any instance of X class, it will pass it to small_classifier, which will predict the actual class (E, F, G, H, I) for the small_class instance. Now main classifier does not need to deal with very small classes, and instead small_classifier will be looking just at these small classes.</p>\n</blockquote>\n<p>Now, in this competition -</p>\n<blockquote>\n  <p>typical (A)         3007<br>\n  negative (B)       1736<br>\n  indeterminate (C)    1108<br>\n  atypical  (D)        483</p>\n</blockquote>\n<p>we might try something like as you mentioned -</p>\n<blockquote>\n  <p>Main Classifier for BCD (3327) vs. A (3007) (approx 1.11:1 ratio). And small_classifier for B (1736) vs C (1108) vs D (483) (approx 2.3:2:1 ratio).</p>\n</blockquote>\n<p>OR</p>\n<blockquote>\n  <p>Main Classifier for CD (3327) vs B (1736) vs A (3007) (approx 2:1:1.8 ratio). And small_classifier for C (1108) vs D (483) (approx 2.3:1 ratio).</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1334675,
          "author_name": "farhanhaikhan",
          "author_url": "",
          "post_date": "06/03/2021 17:20:36",
          "content": "<p>I'm sure that was very intuitive for everyone arriving at this discussion. Thanks for the fruitful explanation! Glad to hear that I could help too! :) Happy Kaggling!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1328004": "```\ntypical          3007\nnegative         1736\nindeterminate    1108\natypical          483\n```\n\nWhoa!! That's pretty imbalanced. How does a professional data scientist handle it?\n\nPrimary Ideas :\n- Use of Class Weights\n- Apply Heavy Augmentations\n- Upsampling & Downsampling\n- Focal Loss for Training",
    "1333082": "I am not a professional data scientist yet 😁, but here is how I handle class imbalance -\n\nI prepare a common validation set out of my dataset and try out each and every method you have mentioned. But, I only didn't get how applying heavy augmentations will help in imbalanced dataset. It would be so amazing if you could explain it or sharer some resources regarding the same. Lastly, I would like to share one more approach I am aware of.\n\nIf you have some classes that have very small number of instances, you can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called small_class for eg) so that your main classifier will classify small_class with all other big classes in the dataset. And if your main classifier encounters any instance of small_class, it will pass it to small_classifier, which will predict the actual class for the small_class instance. This technique can give you accuracy boosts are now main classifier does not need to deal with very small classes, and insted small_classifier will be looking just at these small classes.\n\nIn this competition we will need to train an object detector, so perhaps afforementioned method will need to be done with object detectors.",
    "1334321": "Let's say if we define an augmentation pipeline `P(img)` of `RandomBrightnessContrast`,`Horizontal Flipping` (As Lungs are Symmetric) and `RandomRotation` (-15 to +15 degrees) then we can handle the class imbalances of the following distribution as follows:\n```\ntypical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n```\nGeneral Augmentations : \n```\nfor each_img in original_img :  apply P(img)\n```\nNow, for specific augmentations:\nTake the max sampled class : `class_max = A`. Now the objective is to bring the other classes to `n(class_max) = n(A)`.\n```\nfor each_class in remaining classes : [B,C,D] -\nPerform P(original_img) for randomly chosen [ n(class_max) - n(each_class) ] images in each_class\n```\n\nin the end you find that `n(each_class) = n(class_max)`. Hope that sounds clear!",
    "1334340": "Summarizing your approach : \n```\ntypical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n```\n\nYour idea is something like this if I understand it quite clearly.\n\nLet's assume we have `model (X)` (say Yolov5) which is SOTA ( State of the Art ) for a given ML Task.\n\nApproach to conquer data imbalance pbm : \n\nTake the minimum sampled class `class_min = D` after that :\n\n```\nRecursively combine the classes in ascending order of their samples in each step??\n```\n\nI don't exactly follow, if we follow te above algorithm,\n\nStep 1: Classifier for `indeterminate (C) 1108 vs. atypical (D) 483` need not be balanced itself. (approx 2.3:1 ratio).\n\nHowever I do understand the idea that as we gradually go up the ladder, it should be more balanced.\n\nFor ex.\n\nStep 2  : Classifier for `CD[indeterminate (C) 1108 U atypical (D) 483] (1591) vs.  negative (B)  1736` is indeed more balanced. (approx 1:1.09 ratio).\n\nStep 3 : Classifier for `BCD (3327) vs. typical (A) 3007` is also more balanced. (approx 1.11:1 ratio).",
    "1334519": "Thank you @farhanhaikhan for explaining it. I got it very clearly now 👍",
    "1334565": "Yes I think you got the idea, and just to make sure, I will explain it in more depth.\n\nRegarding my approach, it is as follows -\n\nFor eg we have a dataset with following class distribution -\n\nA:120  B:100  C:90  D:80  E:20  F:15  G:10  H:8  I:5\n\nNow here \n\n> Some classes that have very small number of instances are = E F G H I\n\nAfter that\n\n> You can consider creating a separate classifier for these small classes (called small_classifier for eg). You can group together these small clases under a single class (called X for eg) so that your main classifier will classify X with all other big classes (A, B, C, D).\n\nSo the new distribution for your main classifier becomes, A:120  B:100  C:90  D:80  X:58 (better class balance). And distribution of small_classifier becomes, E:20  F:15  G:10  H:8  I:5 (better class balance)\n\n> If your main classifier encounters any instance of X class, it will pass it to small_classifier, which will predict the actual class (E, F, G, H, I) for the small_class instance. Now main classifier does not need to deal with very small classes, and instead small_classifier will be looking just at these small classes.\n\nNow, in this competition -\n> typical (A)         3007\nnegative (B)       1736\nindeterminate (C)    1108\natypical  (D)        483\n\nwe might try something like as you mentioned -\n\n> Main Classifier for BCD (3327) vs. A (3007) (approx 1.11:1 ratio). And small_classifier for B (1736) vs C (1108) vs D (483) (approx 2.3:2:1 ratio).\n\nOR\n\n> Main Classifier for CD (3327) vs B (1736) vs A (3007) (approx 2:1:1.8 ratio). And small_classifier for C (1108) vs D (483) (approx 2.3:1 ratio).",
    "1334675": "I'm sure that was very intuitive for everyone arriving at this discussion. Thanks for the fruitful explanation! Glad to hear that I could help too! :) Happy Kaggling!"
  },
  "source": "meta"
}