{
  "id": 136021,
  "title": "14th Place - Hacking Macro Recall - Chris Writeup",
  "url": "/competitions/bengaliai-cv19/discussion/136021",
  "author_name": "Chris Deotte",
  "post_date": "2020-03-17T04:35:01.784000",
  "votes": 157,
  "comment_count": 70,
  "views": 0,
  "content": "<p>This is my first computer vision gold medal. I'm very excited!</p>\n\n<p>Thank you Kaggle and Bengali.AI for hosting a fun comp. Thank you teammates Bojan, Shai, Yasin, Jahmed ( <a href=\"/tunguz\">@tunguz</a> <a href=\"/sgalib\">@sgalib</a> <a href=\"/mykttu\">@mykttu</a> <a href=\"/jasemahmed\">@jasemahmed</a> ). I had a blast working with you guys! Thank you Nvidia for providing GPU compute. Below are my contributions to our team's solution. The rest of the team will share more.</p>\n\n<h1>Competition Metric Explained</h1>\n\n<p>This competition's metric is macro recall. That means you compute the recall of each class individually and average them. Most importantly <code>recall = found / exist</code>. There is no penalty for making a false positive! Making more positive predictions for one class can only increase that class' recall never decrease!</p>\n\n<p>Below is a paradoxical example. Imagine that your are predicting one of seven consonant diacritic classes and your CNN outputs the following probabilities:</p>\n\n<h2>Softmax = [0.0, 0.7, 0.3, 0.0, 0.0, 0.0, 0.0]</h2>\n\n<h2>Question: Do you predict class 1 or class 2?</h2>\n\n<p>Let's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.</p>\n\n<p>Therefore your expected macro recall increase if you predict class 1 is <code>1e-4 = 0.70 * 1/1000 * 1/7</code>. And your expected macro recall increase if you predict class 2 is <code>4e-4 = 0.30 * 1/100 * 1/7</code>. Therefore you predict class 2.</p>\n\n<h2>Answer: You predict class 2 not class 1</h2>\n\n<p>By adjusting your predictions in this fashion you can gain a massive 0.0027 public LB increase and 0.0241 private LB increase!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7d138a76f20b20e65a83e2d8c2b88275%2Fsub.jpg?generation=1584419035681165&amp;alt=media\" alt=\"\"></p>\n\n<h3>Code</h3>\n\n<p>Try the following post process on your model to see how much it increases your public and private LB. </p>\n\n<pre><code>preds = model.predict(X_test)\np0 = np.argmax(preds[0],axis=1)\np1 = np.argmax(preds[1],axis=1)\np2 = np.argmax(preds[2],axis=1)\n\nEXP = -1.2\n\ns = pd.Series(p0)\nvc = s.value_counts().sort_index()\ndf = pd.DataFrame({'a':np.arange(168),'b':np.ones(168)})\ndf.b = df.a.map(vc)\ndf.fillna(df.b.min(),inplace=True)\nmat1 = np.diag(df.b.astype('float32')**EXP)\n\ns = pd.Series(p1)\nvc = s.value_counts().sort_index()\ndf = pd.DataFrame({'a':np.arange(11),'b':np.ones(11)})\ndf.b = df.a.map(vc)\ndf.fillna(df.b.min(),inplace=True)\nmat2 = np.diag(df.b.astype('float32')**EXP)\n\ns = pd.Series(p2)\nvc = s.value_counts().sort_index()\ndf = pd.DataFrame({'a':np.arange(7),'b':np.ones(7)})\ndf.b = df.a.map(vc)\ndf.fillna(df.b.min(),inplace=True)\nmat3 = np.diag(df.b.astype('float32')**EXP)\n\np0 = np.argmax( preds[0].dot(mat1), axis=1)\np1 = np.argmax( preds[1].dot(mat2), axis=1)\np2 = np.argmax( preds[2].dot(mat3), axis=1)\n</code></pre>\n\n<h1>Chris Model</h1>\n\n<p>Our team's solution is an ensemble. For my model I trained a 128x256 efficientNetB6 150 epochs with CAM CutMix in addition to basic rotation, scale, shift, cutout, and cutmix. Training took 24 hours on four Nvidia V100 GPUs. It's CV is 0.9967, public LB 0.9916, and private LB 0.9544. </p>\n\n<p>CAM CutMix is where you find the class activation maps of the images and then remove the most important parts of the image and replace them with another image. This challenges your CNN and helps it generalize.</p>\n\n<h2>CAM Maps</h2>\n\n<p>First you train one model to produce CAM maps. (First model was 256x256 efficientNetB6). Next you use CAM maps to train a second model. (Second model was 128x256 efficientNetB6). Here are some CAM maps: (More pictures <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136025\">here</a>).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb3a043c661bc766967d9369f51bb9059%2FScreen%20Shot%202020-03-16%20at%207.44.04%20PM.png?generation=1584414144102358&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F129399f9d691c9cf3323d73f82c7a262%2FScreen%20Shot%202020-03-16%20at%207.42.46%20PM.png?generation=1584414156528553&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2c9e3348089070dadf8ecf422c3f0a23%2FScreen%20Shot%202020-03-16%20at%207.43.27%20PM.png?generation=1584414169952396&amp;alt=media\" alt=\"\"></p>\n\n<h2>CAM CutMix</h2>\n\n<p>CutMix is a combination of two images. The first image is displayed as yellow below to help us visualize it. First one component either root, vowel, consonant is randomly selected. Next a random percentage from 15% to 25% is selected. Next that percentage of the first image is removed using the chosen component type's CAM map. Finally the same region from a second randomly selected image is inserted. The second image is displayed as blue below. 50% of images use CAM CutMix and 50% use regular CutMix. (CAM CutMix increased CV and LB by 0.001 over regular CutMix).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F253a777c731827a5bb5ced8cdf85558a%2FScreen%20Shot%202020-03-16%20at%207.44.42%20PM.png?generation=1584414301224112&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 776008,
      "postDate": "2020-03-17T04:35:01.783Z",
      "content": "<p>This is my first computer vision gold medal. I'm very excited!</p>\n\n<p>Thank you Kaggle and Bengali.AI for hosting a fun comp. Thank you teammates Bojan, Shai, Yasin, Jahmed ( <a href=\"/tunguz\">@tunguz</a> <a href=\"/sgalib\">@sgalib</a> <a href=\"/mykttu\">@mykttu</a> <a href=\"/jasemahmed\">@jasemahmed</a> ). I had a blast working with you guys! Thank you Nvidia for providing GPU compute. Below are my contributions to our team's solution. The rest of the team will share more.</p>\n\n<h1>Competition Metric Explained</h1>\n\n<p>This competition's metric is macro recall. That means you compute the recall of each class individually and average them. Most importantly <code>recall = found / exist</code>. There is no penalty for making a false positive! Making more positive predictions for one class can only increase that class' recall never decrease!</p>\n\n<p>Below is a paradoxical example. Imagine that your are predicting one of seven consonant diacritic classes and your CNN outputs the following probabilities:</p>\n\n<h2>Softmax = [0.0, 0.7, 0.3, 0.0, 0.0, 0.0, 0.0]</h2>\n\n<h2>Question: Do you predict class 1 or class 2?</h2>\n\n<p>Let's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.</p>\n\n<p>Therefore your expected macro recall increase if you predict class 1 is <code>1e-4 = 0.70 * 1/1000 * 1/7</code>. And your expected macro recall increase if you predict class 2 is <code>4e-4 = 0.30 * 1/100 * 1/7</code>. Therefore you predict class 2.</p>\n\n<h2>Answer: You predict class 2 not class 1</h2>\n\n<p>By adjusting your predictions in this fashion you can gain a massive 0.0027 public LB increase and 0.0241 private LB increase!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7d138a76f20b20e65a83e2d8c2b88275%2Fsub.jpg?generation=1584419035681165&amp;alt=media\" alt=\"\"></p>\n\n<h3>Code</h3>\n\n<p>Try the following post process on your model to see how much it increases your public and private LB. </p>\n\n<pre><code>preds = model.predict(X_test)\np0 = np.argmax(preds[0],axis=1)\np1 = np.argmax(preds[1],axis=1)\np2 = np.argmax(preds[2],axis=1)\n\nEXP = -1.2\n\ns = pd.Series(p0)\nvc = s.value_counts().sort_index()\ndf = pd.DataFrame({'a':np.arange(168),'b':np.ones(168)})\ndf.b = df.a.map(vc)\ndf.fillna(df.b.min(),inplace=True)\nmat1 = np.diag(df.b.astype('float32')**EXP)\n\ns = pd.Series(p1)\nvc = s.value_counts().sort_index()\ndf = pd.DataFrame({'a':np.arange(11),'b':np.ones(11)})\ndf.b = df.a.map(vc)\ndf.fillna(df.b.min(),inplace=True)\nmat2 = np.diag(df.b.astype('float32')**EXP)\n\ns = pd.Series(p2)\nvc = s.value_counts().sort_index()\ndf = pd.DataFrame({'a':np.arange(7),'b':np.ones(7)})\ndf.b = df.a.map(vc)\ndf.fillna(df.b.min(),inplace=True)\nmat3 = np.diag(df.b.astype('float32')**EXP)\n\np0 = np.argmax( preds[0].dot(mat1), axis=1)\np1 = np.argmax( preds[1].dot(mat2), axis=1)\np2 = np.argmax( preds[2].dot(mat3), axis=1)\n</code></pre>\n\n<h1>Chris Model</h1>\n\n<p>Our team's solution is an ensemble. For my model I trained a 128x256 efficientNetB6 150 epochs with CAM CutMix in addition to basic rotation, scale, shift, cutout, and cutmix. Training took 24 hours on four Nvidia V100 GPUs. It's CV is 0.9967, public LB 0.9916, and private LB 0.9544. </p>\n\n<p>CAM CutMix is where you find the class activation maps of the images and then remove the most important parts of the image and replace them with another image. This challenges your CNN and helps it generalize.</p>\n\n<h2>CAM Maps</h2>\n\n<p>First you train one model to produce CAM maps. (First model was 256x256 efficientNetB6). Next you use CAM maps to train a second model. (Second model was 128x256 efficientNetB6). Here are some CAM maps: (More pictures <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136025\">here</a>).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb3a043c661bc766967d9369f51bb9059%2FScreen%20Shot%202020-03-16%20at%207.44.04%20PM.png?generation=1584414144102358&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F129399f9d691c9cf3323d73f82c7a262%2FScreen%20Shot%202020-03-16%20at%207.42.46%20PM.png?generation=1584414156528553&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2c9e3348089070dadf8ecf422c3f0a23%2FScreen%20Shot%202020-03-16%20at%207.43.27%20PM.png?generation=1584414169952396&amp;alt=media\" alt=\"\"></p>\n\n<h2>CAM CutMix</h2>\n\n<p>CutMix is a combination of two images. The first image is displayed as yellow below to help us visualize it. First one component either root, vowel, consonant is randomly selected. Next a random percentage from 15% to 25% is selected. Next that percentage of the first image is removed using the chosen component type's CAM map. Finally the same region from a second randomly selected image is inserted. The second image is displayed as blue below. 50% of images use CAM CutMix and 50% use regular CutMix. (CAM CutMix increased CV and LB by 0.001 over regular CutMix).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F253a777c731827a5bb5ced8cdf85558a%2FScreen%20Shot%202020-03-16%20at%207.44.42%20PM.png?generation=1584414301224112&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This is my first computer vision gold medal. I'm very excited!\n\nThank you Kaggle and Bengali.AI for hosting a fun comp. Thank you teammates Bojan, Shai, Yasin, Jahmed ( @tunguz @sgalib @mykttu @jasemahmed ). I had a blast working with you guys! Thank you Nvidia for providing GPU compute. Below are my contributions to our team's solution. The rest of the team will share more.\n\n# Competition Metric Explained\n\nThis competition's metric is macro recall. That means you compute the recall of each class individually and average them. Most importantly `recall = found / exist`. There is no penalty for making a false positive! Making more positive predictions for one class can only increase that class' recall never decrease!\n\nBelow is a paradoxical example. Imagine that your are predicting one of seven consonant diacritic classes and your CNN outputs the following probabilities:\n\n## Softmax = [0.0, 0.7, 0.3, 0.0, 0.0, 0.0, 0.0]\n\n## Question: Do you predict class 1 or class 2?\n\nLet's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.\n\nTherefore your expected macro recall increase if you predict class 1 is `1e-4 = 0.70 * 1/1000 * 1/7`. And your expected macro recall increase if you predict class 2 is `4e-4 = 0.30 * 1/100 * 1/7`. Therefore you predict class 2.\n\n## Answer: You predict class 2 not class 1\nBy adjusting your predictions in this fashion you can gain a massive 0.0027 public LB increase and 0.0241 private LB increase!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7d138a76f20b20e65a83e2d8c2b88275%2Fsub.jpg?generation=1584419035681165&amp;alt=media)\n\n\n### Code\nTry the following post process on your model to see how much it increases your public and private LB. \n    \n    preds = model.predict(X_test)\n    p0 = np.argmax(preds[0],axis=1)\n    p1 = np.argmax(preds[1],axis=1)\n    p2 = np.argmax(preds[2],axis=1)\n\n    EXP = -1.2\n\n    s = pd.Series(p0)\n    vc = s.value_counts().sort_index()\n    df = pd.DataFrame({'a':np.arange(168),'b':np.ones(168)})\n    df.b = df.a.map(vc)\n    df.fillna(df.b.min(),inplace=True)\n    mat1 = np.diag(df.b.astype('float32')**EXP)\n\n    s = pd.Series(p1)\n    vc = s.value_counts().sort_index()\n    df = pd.DataFrame({'a':np.arange(11),'b':np.ones(11)})\n    df.b = df.a.map(vc)\n    df.fillna(df.b.min(),inplace=True)\n    mat2 = np.diag(df.b.astype('float32')**EXP)\n\n    s = pd.Series(p2)\n    vc = s.value_counts().sort_index()\n    df = pd.DataFrame({'a':np.arange(7),'b':np.ones(7)})\n    df.b = df.a.map(vc)\n    df.fillna(df.b.min(),inplace=True)\n    mat3 = np.diag(df.b.astype('float32')**EXP)\n\n    p0 = np.argmax( preds[0].dot(mat1), axis=1)\n    p1 = np.argmax( preds[1].dot(mat2), axis=1)\n    p2 = np.argmax( preds[2].dot(mat3), axis=1)\n\n# Chris Model\nOur team's solution is an ensemble. For my model I trained a 128x256 efficientNetB6 150 epochs with CAM CutMix in addition to basic rotation, scale, shift, cutout, and cutmix. Training took 24 hours on four Nvidia V100 GPUs. It's CV is 0.9967, public LB 0.9916, and private LB 0.9544. \n  \nCAM CutMix is where you find the class activation maps of the images and then remove the most important parts of the image and replace them with another image. This challenges your CNN and helps it generalize.\n\n## CAM Maps\nFirst you train one model to produce CAM maps. (First model was 256x256 efficientNetB6). Next you use CAM maps to train a second model. (Second model was 128x256 efficientNetB6). Here are some CAM maps: (More pictures [here][1]).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb3a043c661bc766967d9369f51bb9059%2FScreen%20Shot%202020-03-16%20at%207.44.04%20PM.png?generation=1584414144102358&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F129399f9d691c9cf3323d73f82c7a262%2FScreen%20Shot%202020-03-16%20at%207.42.46%20PM.png?generation=1584414156528553&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2c9e3348089070dadf8ecf422c3f0a23%2FScreen%20Shot%202020-03-16%20at%207.43.27%20PM.png?generation=1584414169952396&amp;alt=media)\n\n## CAM CutMix\n\nCutMix is a combination of two images. The first image is displayed as yellow below to help us visualize it. First one component either root, vowel, consonant is randomly selected. Next a random percentage from 15% to 25% is selected. Next that percentage of the first image is removed using the chosen component type's CAM map. Finally the same region from a second randomly selected image is inserted. The second image is displayed as blue below. 50% of images use CAM CutMix and 50% use regular CutMix. (CAM CutMix increased CV and LB by 0.001 over regular CutMix).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F253a777c731827a5bb5ced8cdf85558a%2FScreen%20Shot%202020-03-16%20at%207.44.42%20PM.png?generation=1584414301224112&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/bengaliai-cv19/discussion/136025",
      "votes": 156
    },
    {
      "id": 776316,
      "postDate": "2020-03-17T09:52:23.180Z",
      "content": "<p>Just amazing 🤓 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2726150%2F80f84f0436d148bd2ea5e5ee8e7a6eef%2Fmagic.JPG?generation=1584438727758691&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Just amazing 🤓 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2726150%2F80f84f0436d148bd2ea5e5ee8e7a6eef%2Fmagic.JPG?generation=1584438727758691&amp;alt=media)\n",
      "votes": 6,
      "replies": [
        {
          "id": 776503,
          "postDate": "2020-03-17T12:58:07.530Z",
          "content": "<p>You should definitely write a summary of your approach. Also mention about the <code>0.9643</code> private LB approach. </p>",
          "rawMarkdown": "You should definitely write a summary of your approach. Also mention about the `0.9643` private LB approach. ",
          "votes": 2
        },
        {
          "id": 776508,
          "postDate": "2020-03-17T13:00:13.003Z",
          "content": "<p>wow, what a jump by PP</p>",
          "rawMarkdown": "wow, what a jump by PP",
          "votes": 2
        },
        {
          "id": 776625,
          "postDate": "2020-03-17T14:16:27.227Z",
          "content": "<p>Wow private LB 0.9643 incredible. Too bad we didn't know that public LB in the 0.98 could place so high in private. I also had private LB over 0.96 on some of my public LB 0.98 solutions.</p>",
          "rawMarkdown": "Wow private LB 0.9643 incredible. Too bad we didn't know that public LB in the 0.98 could place so high in private. I also had private LB over 0.96 on some of my public LB 0.98 solutions.",
          "votes": 2
        }
      ]
    },
    {
      "id": 781110,
      "postDate": "2020-03-20T23:37:39.267Z",
      "content": "<p>I just applied PP to the highest scoring public notebook. It turns the 8th place solution into the 2nd place solution. The private LB becomes 0.9704! Notebook <a href=\"https://www.kaggle.com/bamps53/private0-9552-tpu-keras-metric-learning\">here</a>. Use <code>EXP = [-1.2, -1.2, -0.5]</code>.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc5ac8746ceed9dad3100932b825df2d6%2FScreen%20Shot%202020-03-20%20at%204.26.04%20PM.png?generation=1584747191411505&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I just applied PP to the highest scoring public notebook. It turns the 8th place solution into the 2nd place solution. The private LB becomes 0.9704! Notebook [here][1]. Use `EXP = [-1.2, -1.2, -0.5]`.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc5ac8746ceed9dad3100932b825df2d6%2FScreen%20Shot%202020-03-20%20at%204.26.04%20PM.png?generation=1584747191411505&amp;alt=media)\n\n[1]: https://www.kaggle.com/bamps53/private0-9552-tpu-keras-metric-learning\n",
      "votes": 3,
      "replies": [
        {
          "id": 781200,
          "postDate": "2020-03-21T02:59:47.833Z",
          "content": "<p>Thanks!! I update mine with your postprocess:)</p>",
          "rawMarkdown": "Thanks!! I update mine with your postprocess:)",
          "votes": 1
        },
        {
          "id": 782211,
          "postDate": "2020-03-22T03:29:11.693Z",
          "content": "<p>Your PP method just too powerful.. I simply trained 3 model(efficientnetB3) with cutmix, mixup and cutout augmentations for each class(root, vowel and consonant). And I got 0.9560 private score and 0.9759 public score..This method definitely the best technique I learned from this competition. Thanks again, Chris!!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2F6d96a469346e03e09daa4d16a1bd9af2%2F2020-03-22%2011.26.18.png?generation=1584847720533229&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Your PP method just too powerful.. I simply trained 3 model(efficientnetB3) with cutmix, mixup and cutout augmentations for each class(root, vowel and consonant). And I got 0.9560 private score and 0.9759 public score..This method definitely the best technique I learned from this competition. Thanks again, Chris!!![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2F6d96a469346e03e09daa4d16a1bd9af2%2F2020-03-22%2011.26.18.png?generation=1584847720533229&amp;alt=media)\n ",
          "votes": 1
        }
      ]
    },
    {
      "id": 776498,
      "postDate": "2020-03-17T12:53:23.970Z",
      "content": "<p>I love the post processing. I am not a fan of calling PP hacking though as I believe that optimizing the predictions towards the metric is an integral part of machine learning. If you adapt your loss functions to better capture the metric, you are also not calling it hacking. </p>\n\n<p>If the hosts choose a metric that overweights rare cases, then I would expect them to care about rare cases. So being more precise on them seems important. </p>",
      "rawMarkdown": "I love the post processing. I am not a fan of calling PP hacking though as I believe that optimizing the predictions towards the metric is an integral part of machine learning. If you adapt your loss functions to better capture the metric, you are also not calling it hacking. \n\nIf the hosts choose a metric that overweights rare cases, then I would expect them to care about rare cases. So being more precise on them seems important. ",
      "votes": 3,
      "replies": [
        {
          "id": 776614,
          "postDate": "2020-03-17T14:04:48.970Z",
          "content": "<p>Good point I agree. I chose the title \"Hacking Macro Recall\" to catch more attention. This post process is just mathematical optimization which should be done in every comp. </p>\n\n<p>I was sure that your team was using this PP when you asked for teammates (14 days ago <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#762152\">here</a>. your CV was 0.9890 and LB 0.9845) because your CV and LB gap was so small. Everyone else's gap was around 0.010 while yours and mine (with PP) were around 0.05.</p>\n\n<p>Understanding metrics is overlooked at Kaggle. In half of my past competitions, you could take the top public kernel and post process the predictions by optimizing to the metric and get silver medal or above. Google Quest Q&amp;A with spearman rank correlation coefficient is great example. And Cloud comp and Steel comp are two great examples. Both use <strong>image-wise</strong> dice instead of the common <strong>batch-wise</strong> dice.</p>\n\n<p>Congrats to you and your team on another amazing finish. I am so impressed how The Zoo does so well in such a wide range of competition types. I would love to team up some day and learn from you guys!</p>",
          "rawMarkdown": "Good point I agree. I chose the title \"Hacking Macro Recall\" to catch more attention. This post process is just mathematical optimization which should be done in every comp. \n\nI was sure that your team was using this PP when you asked for teammates (14 days ago [here][1]. your CV was 0.9890 and LB 0.9845) because your CV and LB gap was so small. Everyone else's gap was around 0.010 while yours and mine (with PP) were around 0.05.\n\nUnderstanding metrics is overlooked at Kaggle. In half of my past competitions, you could take the top public kernel and post process the predictions by optimizing to the metric and get silver medal or above. Google Quest Q&amp;A with spearman rank correlation coefficient is great example. And Cloud comp and Steel comp are two great examples. Both use **image-wise** dice instead of the common **batch-wise** dice.\n\nCongrats to you and your team on another amazing finish. I am so impressed how The Zoo does so well in such a wide range of competition types. I would love to team up some day and learn from you guys!\n\n[1]: https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#762152",
          "votes": 6
        },
        {
          "id": 776641,
          "postDate": "2020-03-17T14:27:35.780Z",
          "content": "<p>Thanks a lot Chris! As always you also did an amazing job and you are totally right that understanding the metric and adjusting to it still is sometimes an underappreciated aspect, but one of the most important ones to do well on LB.</p>\n\n<p>Regarding post processing: actually at that point, we only hardcoded a few extra C=3 and C=6 and that already gave us a boost of 30-40 points on LB. </p>",
          "rawMarkdown": "Thanks a lot Chris! As always you also did an amazing job and you are totally right that understanding the metric and adjusting to it still is sometimes an underappreciated aspect, but one of the most important ones to do well on LB.\n\nRegarding post processing: actually at that point, we only hardcoded a few extra C=3 and C=6 and that already gave us a boost of 30-40 points on LB. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 776249,
      "postDate": "2020-03-17T08:31:05.433Z",
      "content": "<p>Congrats, Chris!. One question.  You mentioned  'Therefore your expected macro recall increase if you predict class 1 is 2.3e-4 = 0.70 * 1/1000 * 1/3. And your expected macro recall increase if you predict class 2 is 1e-3. Therefore you predict class 2'. What is 1/3 in the equation? </p>",
      "rawMarkdown": "Congrats, Chris!. One question.  You mentioned  'Therefore your expected macro recall increase if you predict class 1 is 2.3e-4 = 0.70 * 1/1000 * 1/3. And your expected macro recall increase if you predict class 2 is 1e-3. Therefore you predict class 2'. What is 1/3 in the equation? ",
      "votes": 3,
      "replies": [
        {
          "id": 776461,
          "postDate": "2020-03-17T12:11:09.860Z",
          "content": "<p>I was curious too,have you figured it out??</p>",
          "rawMarkdown": "I was curious too,have you figured it out??",
          "votes": 1
        },
        {
          "id": 776595,
          "postDate": "2020-03-17T13:50:39.220Z",
          "content": "<p>That was a typo. I corrected it. Thanks. It should be class 1 is <code>1e-4 = 0.70 * 1/1000 * 1/7</code> and class 2 is <code>4e-4 =  0.30 * 1/100 * 1/7</code>. It is <code>1/7</code> because there are 7 consonant diacritics in the macro average.</p>",
          "rawMarkdown": "That was a typo. I corrected it. Thanks. It should be class 1 is `1e-4 = 0.70 * 1/1000 * 1/7` and class 2 is `4e-4 =  0.30 * 1/100 * 1/7`. It is `1/7` because there are 7 consonant diacritics in the macro average."
        }
      ]
    },
    {
      "id": 776037,
      "postDate": "2020-03-17T05:01:53.377Z",
      "content": "<p>Cool! This CAM CutMix and the refined prediction blow my mind! 👍  Congratulation! </p>",
      "rawMarkdown": "Cool! This CAM CutMix and the refined prediction blow my mind! 👍  Congratulation! ",
      "votes": 3
    },
    {
      "id": 776035,
      "postDate": "2020-03-17T04:58:06.413Z",
      "content": "<p>\"Let's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.\"</p>\n\n<p>how do you know the underlying prior probability of each class (e.g. class1 has 1000 samples, while class2 has 100)?\nwe can't probe the private test data in this case?</p>",
      "rawMarkdown": "\"Let's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.\"\n\nhow do you know the underlying prior probability of each class (e.g. class1 has 1000 samples, while class2 has 100)?\nwe can't probe the private test data in this case?",
      "votes": 3,
      "replies": [
        {
          "id": 776045,
          "postDate": "2020-03-17T05:08:32.293Z",
          "content": "<p>Yes we can. First you make predictions on the test data. Then you use those predictions as the prior probability and compute a second set or predictions by modifying the first set. (See code above).</p>",
          "rawMarkdown": "Yes we can. First you make predictions on the test data. Then you use those predictions as the prior probability and compute a second set or predictions by modifying the first set. (See code above).",
          "votes": 1
        },
        {
          "id": 776054,
          "postDate": "2020-03-17T05:13:30.537Z",
          "content": "<p>His distribution comes from the predicted results.</p>",
          "rawMarkdown": "His distribution comes from the predicted results.",
          "votes": 1
        },
        {
          "id": 776055,
          "postDate": "2020-03-17T05:14:40.377Z",
          "content": "<p>thanks for the reply.</p>\n\n<p>we can test only public test data and not the private one. </p>\n\n<p>i suppose you have made predictions on the test data, so does it reveal that the private and public test data are similar distribution at your test? e.g. do have results like</p>\n\n<p>```\nno post processing:\nprivate lb = xxx\npublic lb = xxx</p>\n\n<p>post processing of fixing 1 class:\nprivate lb = xxx\npublic lb = xxx</p>\n\n<p>post processing of fixing 10 class:\nprivate lb = xxx\npublic lb = xxx</p>\n\n<p>etc</p>\n\n<p>```</p>",
          "rawMarkdown": "thanks for the reply.\n\nwe can test only public test data and not the private one. \n\ni suppose you have made predictions on the test data, so does it reveal that the private and public test data are similar distribution at your test? e.g. do have results like\n\n```\nno post processing:\nprivate lb = xxx\npublic lb = xxx\n\npost processing of fixing 1 class:\nprivate lb = xxx\npublic lb = xxx\n\npost processing of fixing 10 class:\nprivate lb = xxx\npublic lb = xxx\n\netc\n\n\n```"
        },
        {
          "id": 776065,
          "postDate": "2020-03-17T05:20:57.230Z",
          "content": "<p>We predict <strong>both</strong> public and private at the same time with our submitted notebooks. When your code makes test predictions it makes all 200,000 at once. That is both public and private. The post process uses all these 200,000 predictions to compute distribution. Afterward, we only see public LB score but submitted notebook computed public LB and private LB together.</p>",
          "rawMarkdown": "We predict **both** public and private at the same time with our submitted notebooks. When your code makes test predictions it makes all 200,000 at once. That is both public and private. The post process uses all these 200,000 predictions to compute distribution. Afterward, we only see public LB score but submitted notebook computed public LB and private LB together.",
          "votes": 1
        },
        {
          "id": 776083,
          "postDate": "2020-03-17T05:40:11.813Z",
          "content": "<p>Here are results for fixing all 168 roots, 11 vowels and 7 consonants:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7d138a76f20b20e65a83e2d8c2b88275%2Fsub.jpg?generation=1584419035681165&amp;alt=media\" alt=\"image\"></p>\n\n<p>We do have some submissions that only fix all root or only fix all vowel or only fix all consonant. Basically each helps an equal amount. We don't have results for just fixing individual classes (within root or vowel or consonant).</p>\n\n<p>You sort of need to fix them all at once. If you make your model more likely to predict a particular rare class, you must also tell it what the other rare classes are. Otherwise it steals predictions from a rare class to fill the other rare classes. Whereas when you do it all at once, rare classes don't steal from each other but rather steal from common classes.</p>",
          "rawMarkdown": "Here are results for fixing all 168 roots, 11 vowels and 7 consonants:\n![image](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7d138a76f20b20e65a83e2d8c2b88275%2Fsub.jpg?generation=1584419035681165&amp;alt=media)\n\nWe do have some submissions that only fix all root or only fix all vowel or only fix all consonant. Basically each helps an equal amount. We don't have results for just fixing individual classes (within root or vowel or consonant).\n\nYou sort of need to fix them all at once. If you make your model more likely to predict a particular rare class, you must also tell it what the other rare classes are. Otherwise it steals predictions from a rare class to fill the other rare classes. Whereas when you do it all at once, rare classes don't steal from each other but rather steal from common classes.",
          "votes": 1
        }
      ]
    },
    {
      "id": 776057,
      "postDate": "2020-03-17T05:16:02.793Z",
      "content": "<p>Congrats Chris, thanks for sharing.  I thought about the macro recall thing a few days ago but didn't figure out how to do a correct post processing 😂 . Your trick is pretty good 👍</p>",
      "rawMarkdown": "Congrats Chris, thanks for sharing.  I thought about the macro recall thing a few days ago but didn't figure out how to do a correct post processing 😂 . Your trick is pretty good 👍",
      "votes": 4,
      "replies": [
        {
          "id": 776061,
          "postDate": "2020-03-17T05:18:32.073Z",
          "content": "<p>Thanks Venn. Sorry about your shakup. Your team built a great model.</p>",
          "rawMarkdown": "Thanks Venn. Sorry about your shakup. Your team built a great model.",
          "votes": 1
        }
      ]
    },
    {
      "id": 782038,
      "postDate": "2020-03-21T21:40:12.160Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> and congratulation. CAM CutMix is interesting. </p>",
      "rawMarkdown": "Thanks @cdeotte and congratulation. CAM CutMix is interesting. ",
      "votes": 1
    },
    {
      "id": 778975,
      "postDate": "2020-03-18T22:45:10.703Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you for sharing, it's really magic! I'm now trying to reproduce it, this post process is supposed to be used with all 200,000 prediction, not with every batch, right?\nI'm encountering too much memory issue, was it possible to store (200,000, 168) numpy array for your case?</p>\n\n<p>Thanks, I love your cam cutmix idea too!!</p>",
      "rawMarkdown": "@cdeotte Thank you for sharing, it's really magic! I'm now trying to reproduce it, this post process is supposed to be used with all 200,000 prediction, not with every batch, right?\nI'm encountering too much memory issue, was it possible to store (200,000, 168) numpy array for your case?\n\nThanks, I love your cam cutmix idea too!!",
      "votes": 1,
      "replies": [
        {
          "id": 778980,
          "postDate": "2020-03-18T22:56:08.333Z",
          "content": "<p>Thanks. Yes the entire 200,000 predictions. We predict in batches of 512, so <code>batch_preds</code> are softmax output with size <code>(512, 168 + 11 + 7)</code>. Then we add these to a list <code>preds.append(batch_preds)</code>. And afterwards convert to numpy array with <code>preds = np.vstack(preds)</code>. That final numpy array is 200,000 rows by 168+11+7 columns. Since it is <code>float32</code>, that is only <code>148MB = 200000 * (168+11+7) * 4</code>. Not much memory at all.</p>\n\n<p>Make sure you are predicting in batches and not predicting all 200,000 at once.</p>",
          "rawMarkdown": "Thanks. Yes the entire 200,000 predictions. We predict in batches of 512, so `batch_preds` are softmax output with size `(512, 168 + 11 + 7)`. Then we add these to a list `preds.append(batch_preds)`. And afterwards convert to numpy array with `preds = np.vstack(preds)`. That final numpy array is 200,000 rows by 168+11+7 columns. Since it is `float32`, that is only `148MB = 200000 * (168+11+7) * 4`. Not much memory at all.\n\nMake sure you are predicting in batches and not predicting all 200,000 at once."
        },
        {
          "id": 779138,
          "postDate": "2020-03-19T03:45:47.860Z",
          "content": "<p>I finally figured out the problem was in my image loading rather than preds numpy array.\nBut still struggling to make it works. It seems EXP = -1.2 is not suitable for my model, I might need to tune it by validation.\nAnyway, thanks for sharing!!</p>",
          "rawMarkdown": "I finally figured out the problem was in my image loading rather than preds numpy array.\nBut still struggling to make it works. It seems EXP = -1.2 is not suitable for my model, I might need to tune it by validation.\nAnyway, thanks for sharing!!"
        },
        {
          "id": 779148,
          "postDate": "2020-03-19T03:59:12.540Z",
          "content": "<p>Some models don't like -1.2. Try setting root EXP to -0.9, vowel EXP to -1.1 and consonant EXP to -0.6. That worked best for our team ensemble.</p>",
          "rawMarkdown": "Some models don't like -1.2. Try setting root EXP to -0.9, vowel EXP to -1.1 and consonant EXP to -0.6. That worked best for our team ensemble."
        },
        {
          "id": 779150,
          "postDate": "2020-03-19T04:02:36.770Z",
          "content": "<p>Also make sure that your model is outputting probabilities. If is it outputting logits. First apply <code>scipy.special.softmax(preds, axis=1)</code></p>",
          "rawMarkdown": "Also make sure that your model is outputting probabilities. If is it outputting logits. First apply `scipy.special.softmax(preds, axis=1)`"
        }
      ]
    },
    {
      "id": 777856,
      "postDate": "2020-03-18T01:53:43.387Z",
      "content": "<p>Thank you for always sharing detailed, nice explanation write-up!</p>",
      "rawMarkdown": "Thank you for always sharing detailed, nice explanation write-up!",
      "votes": 1
    },
    {
      "id": 777842,
      "postDate": "2020-03-18T01:45:01.267Z",
      "content": "<p>Thanks for sharing! the ideas are wonderful and easy to implement! \nI didn't survive from the shake up 😂 and all I can do is keep learning from the best like you, \nCongrats to you and your team!</p>",
      "rawMarkdown": "Thanks for sharing! the ideas are wonderful and easy to implement! \nI didn't survive from the shake up 😂 and all I can do is keep learning from the best like you, \nCongrats to you and your team!",
      "votes": 1,
      "replies": [
        {
          "id": 777854,
          "postDate": "2020-03-18T01:53:15.937Z",
          "content": "<p>Thanks Geroko. I'm sorry about your shakeup. Your model is great with public LB 0.9909. I bet if you apply PP, you will have one of the best models on private LB.</p>",
          "rawMarkdown": "Thanks Geroko. I'm sorry about your shakeup. Your model is great with public LB 0.9909. I bet if you apply PP, you will have one of the best models on private LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 776603,
      "postDate": "2020-03-17T13:56:43.547Z",
      "content": "<p>Amazing postprocessing on the final prediction! </p>\n\n<p>Did you try to compute the prior distribution from train set true label and apply it on test set, instead of computing the distribution from test set prediction and applying on test set? Just curious about the robustness of this magic!</p>",
      "rawMarkdown": "Amazing postprocessing on the final prediction! \n\nDid you try to compute the prior distribution from train set true label and apply it on test set, instead of computing the distribution from test set prediction and applying on test set? Just curious about the robustness of this magic!",
      "votes": 1,
      "replies": [
        {
          "id": 776619,
          "postDate": "2020-03-17T14:07:24.120Z",
          "content": "<p>Yes. Using train distribution achieves similar result. This confirms that train and test distribution are similar. I think the most difference is between public test and private test. I believe some rare classes become even more rare making this PP even more important for private LB.</p>",
          "rawMarkdown": "Yes. Using train distribution achieves similar result. This confirms that train and test distribution are similar. I think the most difference is between public test and private test. I believe some rare classes become even more rare making this PP even more important for private LB.",
          "votes": 2
        }
      ]
    },
    {
      "id": 776491,
      "postDate": "2020-03-17T12:41:18.727Z",
      "content": "<p>Wow, congrats, both post processing and cam cutmix are great.  We where on the post processing track, but certainly didn't think of cam based stuff.</p>",
      "rawMarkdown": "Wow, congrats, both post processing and cam cutmix are great.  We where on the post processing track, but certainly didn't think of cam based stuff.",
      "votes": 1,
      "replies": [
        {
          "id": 776623,
          "postDate": "2020-03-17T14:14:02.720Z",
          "content": "<p>Thanks CPMP. Congrats to you and Kaz on a strong finish.</p>",
          "rawMarkdown": "Thanks CPMP. Congrats to you and Kaz on a strong finish.",
          "votes": 2
        }
      ]
    },
    {
      "id": 776299,
      "postDate": "2020-03-17T09:31:47.097Z",
      "content": "<p>Oh my..this just genius, so simple! so powerful!</p>",
      "rawMarkdown": "Oh my..this just genius, so simple! so powerful!",
      "votes": 1
    },
    {
      "id": 776241,
      "postDate": "2020-03-17T08:27:28.310Z",
      "content": "<p>Congrats your team and thanks for detailed solution!</p>",
      "rawMarkdown": "Congrats your team and thanks for detailed solution!",
      "votes": 1
    },
    {
      "id": 776171,
      "postDate": "2020-03-17T07:03:48.377Z",
      "content": "<p>Congrats your team and thanks for detailed solution <a href=\"/cdeotte\">@cdeotte</a>!</p>",
      "rawMarkdown": "Congrats your team and thanks for detailed solution @cdeotte!",
      "votes": 1
    },
    {
      "id": 776073,
      "postDate": "2020-03-17T05:29:46.443Z",
      "content": "<p>Congratulation. In CAM CutMix, you randomly remove 15% to 25% by CAM map of one of random selected component, have you also do some operation on target?  (e.g. reduce the weight of the selected component)</p>",
      "rawMarkdown": "Congratulation. In CAM CutMix, you randomly remove 15% to 25% by CAM map of one of random selected component, have you also do some operation on target?  (e.g. reduce the weight of the selected component)",
      "votes": 1,
      "replies": [
        {
          "id": 776080,
          "postDate": "2020-03-17T05:33:40.910Z",
          "content": "<p>I change target in the same way as CutMix. For example, if i replace 17% of image 1 with image 2. Then i change the label to <code>0.83 * label_1 + 0.17 * label_2</code> where labels are one hot encoded.</p>",
          "rawMarkdown": "I change target in the same way as CutMix. For example, if i replace 17% of image 1 with image 2. Then i change the label to `0.83 * label_1 + 0.17 * label_2` where labels are one hot encoded."
        }
      ]
    },
    {
      "id": 789015,
      "postDate": "2020-03-28T09:14:35.623Z",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for sharing and congrats on your 14th(gold) place!</p>\n\n<p>I tried your post-processing in my late submission (I added results in <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136815\">my solution</a>).</p>\n\n<p>By your magic, my score increased from 0.9536 (10th) to <strong>0.9653(3rd)!!!</strong>  Amaging!\nYour magic is a really magic!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2Fd416d7853979e1c35a4e6d66a8307283%2Fbest_private_score_in_late_sub.png?generation=1585386609111447&amp;alt=media\" alt=\"best private score in late submissions\"></p>",
      "rawMarkdown": "Thanks @cdeotte for sharing and congrats on your 14th(gold) place!\n\nI tried your post-processing in my late submission (I added results in [my solution](https://www.kaggle.com/c/bengaliai-cv19/discussion/136815)).\n\nBy your magic, my score increased from 0.9536 (10th) to **0.9653(3rd)!!!**  Amaging!\nYour magic is a really magic!\n\n![best private score in late submissions](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2Fd416d7853979e1c35a4e6d66a8307283%2Fbest_private_score_in_late_sub.png?generation=1585386609111447&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 789740,
          "postDate": "2020-03-29T00:05:10.063Z",
          "content": "<p>Awesome!</p>",
          "rawMarkdown": "Awesome!",
          "votes": 1
        }
      ]
    },
    {
      "id": 781677,
      "postDate": "2020-03-21T15:02:57.443Z",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> ,\nThank you for your sharing.\nAbout the PP method, I have some questions.</p>\n\n<ol>\n<li><p>Does it have to be applied after ensemble? I applied the PP for each model result and ensemble them, but private LB drop 0.1. Or maybe I need to fine tune EXP?</p></li>\n<li><p>Have you try this PP to other compete? It seems like works for other competes.</p></li>\n</ol>",
      "rawMarkdown": "Hi @cdeotte ,\nThank you for your sharing.\nAbout the PP method, I have some questions.\n\n1. Does it have to be applied after ensemble? I applied the PP for each model result and ensemble them, but private LB drop 0.1. Or maybe I need to fine tune EXP?\n\n2. Have you try this PP to other compete? It seems like works for other competes.",
      "votes": 2,
      "replies": [
        {
          "id": 781708,
          "postDate": "2020-03-21T15:18:57.247Z",
          "content": "<p>You can apply it before or after ensemble. If your ensemble is a linear combination then it is mathematically the same. If LB drops try adding <code>EXP = 0.2</code> first to see if that helps. Also make sure your model is outputting <code>softmax</code> which are probabilities and not <code>logits</code> which are log odds.</p>\n\n<p>After <code>EXP = 0.2</code> works, try -0.4, -0.6, etc. At some point the LB will drop. That is usually caused by the consonant EXP being too large. So at that point leave the consonant EXP at a good value and increase the others. </p>\n\n<p>For example if <code>EXP = -0.5</code> works for all and <code>EXP = -0.6</code> does not work for all. Then try <code>EXP = -0.6, -0.6, -0.5</code> and then keep increasing with <code>EXP = -0.8, -0.8, -0.5</code> etc. When i write three <code>EXP</code>, i mean assign the root, vowel, consonant <code>EXP</code> with different values.</p>\n\n<p>You can find optimal <code>EXP</code> using your local CV or the public LB. And if it increases those then it will increase private LB too. You can even make plots of the effects of <code>EXP</code> on CV for root, vowel, and consonant using your local CV.</p>",
          "rawMarkdown": "You can apply it before or after ensemble. If your ensemble is a linear combination then it is mathematically the same. If LB drops try adding `EXP = 0.2` first to see if that helps. Also make sure your model is outputting `softmax` which are probabilities and not `logits` which are log odds.\n\nAfter `EXP = 0.2` works, try -0.4, -0.6, etc. At some point the LB will drop. That is usually caused by the consonant EXP being too large. So at that point leave the consonant EXP at a good value and increase the others. \n\nFor example if `EXP = -0.5` works for all and `EXP = -0.6` does not work for all. Then try `EXP = -0.6, -0.6, -0.5` and then keep increasing with `EXP = -0.8, -0.8, -0.5` etc. When i write three `EXP`, i mean assign the root, vowel, consonant `EXP` with different values.\n\nYou can find optimal `EXP` using your local CV or the public LB. And if it increases those then it will increase private LB too. You can even make plots of the effects of `EXP` on CV for root, vowel, and consonant using your local CV.",
          "votes": 3
        },
        {
          "id": 781764,
          "postDate": "2020-03-21T16:01:49.080Z",
          "content": "<p>Thank you for your clear description.\nI will try to fine tune EXP.</p>",
          "rawMarkdown": "Thank you for your clear description.\nI will try to fine tune EXP.",
          "votes": 1
        }
      ]
    },
    {
      "id": 779411,
      "postDate": "2020-03-19T10:17:07.153Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> congratulations and thank you for sharing. \nOur team had thought about pp of Macro Recall, but we didn't have a good idea，so we suffered from shakeup.\nI tried you method, it is really impressived, EXP=-1.2 did not work for me, so I tried setting root EXP to -0.9, vowel EXP to -1.1 and consonant EXP to -0.6, and got public LB: 0.9920-&gt;0.9931, private LB: 0.9360-&gt;0.9528. </p>",
      "rawMarkdown": "@cdeotte congratulations and thank you for sharing. \nOur team had thought about pp of Macro Recall, but we didn't have a good idea，so we suffered from shakeup.\nI tried you method, it is really impressived, EXP=-1.2 did not work for me, so I tried setting root EXP to -0.9, vowel EXP to -1.1 and consonant EXP to -0.6, and got public LB: 0.9920-&gt;0.9931, private LB: 0.9360-&gt;0.9528. ",
      "votes": 2,
      "replies": [
        {
          "id": 781109,
          "postDate": "2020-03-20T23:30:16.973Z",
          "content": "<p>Thanks Gary</p>",
          "rawMarkdown": "Thanks Gary"
        }
      ]
    },
    {
      "id": 776760,
      "postDate": "2020-03-17T16:10:34.487Z",
      "content": "<p>Hey Chris nice job! Thanks for sharing. Can you please explain at a high level why you created matricies and what the matrix multiplication is doing for the less mathematically inclined. </p>",
      "rawMarkdown": "Hey Chris nice job! Thanks for sharing. Can you please explain at a high level why you created matricies and what the matrix multiplication is doing for the less mathematically inclined. ",
      "votes": 2,
      "replies": [
        {
          "id": 776800,
          "postDate": "2020-03-17T16:35:52.390Z",
          "content": "<p>The matrix is just a way to do the following. Say that your softmax output is <code>x = [0, 0.7, 0.3, 0]</code> and you want to multiply those 4 numbers by these 4 numbers <code>2, 5, 3, 1</code> respectively. Then you create a diagonal matrix with the 4 numbers as follows:</p>\n\n<pre><code>mat = [\n[2, 0, 0, 0]\n[0, 5, 0, 0]\n[0, 0, 3, 0]\n[0, 0, 0, 1]]\n</code></pre>\n\n<p>The matrix is created by calling <code>mat = np.diag([2, 5, 3, 1])</code> and then <code>x.dot(mat)</code> multiplies the softmax output by the 4 numbers. The final result is <code>[0*2, 0.7*5, 0.3*3, 0*1]</code>.</p>",
          "rawMarkdown": "The matrix is just a way to do the following. Say that your softmax output is `x = [0, 0.7, 0.3, 0]` and you want to multiply those 4 numbers by these 4 numbers `2, 5, 3, 1` respectively. Then you create a diagonal matrix with the 4 numbers as follows:\n\n    mat = [\n    [2, 0, 0, 0]\n    [0, 5, 0, 0]\n    [0, 0, 3, 0]\n    [0, 0, 0, 1]]\n\nThe matrix is created by calling `mat = np.diag([2, 5, 3, 1])` and then `x.dot(mat)` multiplies the softmax output by the 4 numbers. The final result is `[0*2, 0.7*5, 0.3*3, 0*1]`.",
          "votes": 3
        }
      ]
    },
    {
      "id": 776116,
      "postDate": "2020-03-17T06:08:15.007Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Congratulations! Such a cool method! I tried to leverage the recall information by designing custom loss but failed hahaha. I have a question: it seems that the expected value of recall increase has to be computed based on a ready distribution estimation, while the final distribution estimation is really computed after the recall hack...So using this trick is kind of like solving a problem with recursive solution dependence. How do you solve this? I take that you simply use first-time argmax predictions as estimation for distribution? </p>",
      "rawMarkdown": "@cdeotte Congratulations! Such a cool method! I tried to leverage the recall information by designing custom loss but failed hahaha. I have a question: it seems that the expected value of recall increase has to be computed based on a ready distribution estimation, while the final distribution estimation is really computed after the recall hack...So using this trick is kind of like solving a problem with recursive solution dependence. How do you solve this? I take that you simply use first-time argmax predictions as estimation for distribution? ",
      "votes": 2,
      "replies": [
        {
          "id": 776135,
          "postDate": "2020-03-17T06:22:57.487Z",
          "content": "<p>Yes. Use first-time argmax as estimation. Furthermore, we considered training a model by upsampling all classes to have equal distribution so that our first-time argmax would be more unbiased but we didn't have time to do that.</p>",
          "rawMarkdown": "Yes. Use first-time argmax as estimation. Furthermore, we considered training a model by upsampling all classes to have equal distribution so that our first-time argmax would be more unbiased but we didn't have time to do that.",
          "votes": 1
        },
        {
          "id": 776539,
          "postDate": "2020-03-17T13:15:39.723Z",
          "content": "<p>Thank you very much. I personaly think this hack is the most code/complexity-efficient hack for improving score in this competition hahaha.</p>",
          "rawMarkdown": "Thank you very much. I personaly think this hack is the most code/complexity-efficient hack for improving score in this competition hahaha."
        }
      ]
    },
    {
      "id": 776109,
      "postDate": "2020-03-17T06:00:56.257Z",
      "content": "<p>OMG! After the post-processing, my private LB and public LB massively increase by 0.022 and 0.048 respectively!! \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F8a63beec93df3988d639a9c258ce86ac%2FScreen%20Shot%202020-03-16%20at%2010.55.41%20PM.png?generation=1584424713221162&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "OMG! After the post-processing, my private LB and public LB massively increase by 0.022 and 0.048 respectively!! \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F8a63beec93df3988d639a9c258ce86ac%2FScreen%20Shot%202020-03-16%20at%2010.55.41%20PM.png?generation=1584424713221162&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 776133,
          "postDate": "2020-03-17T06:20:45.903Z",
          "content": "<p>Wow great increase. You can also try <code>EXP = -1.2</code> that works best for me. It will probably increase you even higher!</p>",
          "rawMarkdown": "Wow great increase. You can also try `EXP = -1.2` that works best for me. It will probably increase you even higher!",
          "votes": 1
        },
        {
          "id": 776179,
          "postDate": "2020-03-17T07:17:20.967Z",
          "content": "<p>This post-process is really impressive! 👍 👍\n I use exp=-1 to my previous top LB kernel, the private LB score increases by 0.0247, to 0.9561.! OMG!! I'll try exp=-1.2 tomorrow! Many thanks!!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F74bfa168ba8e34b7a96dfe808d12f51d%2FIMG_5372.PNG?generation=1584429338367971&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F0f81ab5a845b2127baa1ed60c8e28e41%2FIMG_5371.jpg?generation=1584429338405133&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "This post-process is really impressive! 👍 👍\n I use exp=-1 to my previous top LB kernel, the private LB score increases by 0.0247, to 0.9561.! OMG!! I'll try exp=-1.2 tomorrow! Many thanks!!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F74bfa168ba8e34b7a96dfe808d12f51d%2FIMG_5372.PNG?generation=1584429338367971&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F0f81ab5a845b2127baa1ed60c8e28e41%2FIMG_5371.jpg?generation=1584429338405133&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 776627,
          "postDate": "2020-03-17T14:17:57.123Z",
          "content": "<p>Wow, such a high private LB score. Awesome!</p>",
          "rawMarkdown": "Wow, such a high private LB score. Awesome!",
          "votes": 1
        },
        {
          "id": 776665,
          "postDate": "2020-03-17T14:41:56.340Z",
          "content": "<p>What is awesome is your post-processing! This is definitely one of the most effective methods what I've learned from this competition! </p>",
          "rawMarkdown": "What is awesome is your post-processing! This is definitely one of the most effective methods what I've learned from this competition! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 776015,
      "postDate": "2020-03-17T04:40:24.947Z",
      "content": "<p>WOW THIS IS REALLLLL\nMMMMMMAGICCCCCC!</p>\n\n<p>Congratulation! ;)</p>",
      "rawMarkdown": "WOW THIS IS REALLLLL\nMMMMMMAGICCCCCC!\n\nCongratulation! ;)",
      "votes": 2,
      "replies": [
        {
          "id": 776019,
          "postDate": "2020-03-17T04:43:44.020Z",
          "content": "<p>Thanks. Congrats to you Qishen. It was impressive how you held 1st place for so long. I discovered this trick a month ago and I thought this was why you were at 0.9930 while everyone else was at 0.9900. But after reading your solution, it appears that you didn't use this trick.</p>",
          "rawMarkdown": "Thanks. Congrats to you Qishen. It was impressive how you held 1st place for so long. I discovered this trick a month ago and I thought this was why you were at 0.9930 while everyone else was at 0.9900. But after reading your solution, it appears that you didn't use this trick.",
          "votes": 5
        },
        {
          "id": 776030,
          "postDate": "2020-03-17T04:54:20.083Z",
          "content": "<p>I thought that distribution postprocessing may work, but it's too dangerous, so I even didn't give a try.\nBut I'm wrong, it works so well -- If top 1 use it, he may have gone to 0.98+ or something 😂</p>",
          "rawMarkdown": "I thought that distribution postprocessing may work, but it's too dangerous, so I even didn't give a try.\nBut I'm wrong, it works so well -- If top 1 use it, he may have gone to 0.98+ or something 😂",
          "votes": 5
        },
        {
          "id": 776032,
          "postDate": "2020-03-17T04:57:30.100Z",
          "content": "<p>Yes it works well. Notice that the post processing computes the entire test distribution which includes both the private and public test data. Therefore it does not overfit the public LB.</p>",
          "rawMarkdown": "Yes it works well. Notice that the post processing computes the entire test distribution which includes both the private and public test data. Therefore it does not overfit the public LB.",
          "votes": 3
        },
        {
          "id": 776046,
          "postDate": "2020-03-17T05:10:02.677Z",
          "content": "<blockquote>\n  <p>computes the entire test distribution which includes both the private and public test data</p>\n</blockquote>\n\n<p>Didn't realized it, thanks! </p>",
          "rawMarkdown": "&gt; computes the entire test distribution which includes both the private and public test data\n\nDidn't realized it, thanks! ",
          "votes": 1
        },
        {
          "id": 778645,
          "postDate": "2020-03-18T15:59:28.990Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I've got a question. Have you tried using this postprocessing twice?</p>\n\n<blockquote>\n  <p>First you make predictions on the test data. Then you use those predictions as the prior probability and compute a second set or predictions by modifying the first set. (See code above).</p>\n</blockquote>\n\n<p>Then you use the second set or predictions as the prior probability and compute a third one?</p>",
          "rawMarkdown": "@cdeotte I've got a question. Have you tried using this postprocessing twice?\n\n\n&gt; First you make predictions on the test data. Then you use those predictions as the prior probability and compute a second set or predictions by modifying the first set. (See code above).\n\nThen you use the second set or predictions as the prior probability and compute a third one?"
        }
      ]
    },
    {
      "id": 945635,
      "postDate": "2020-07-26T03:35:57.870Z",
      "content": "<p>Thank you very much, I learned a lot from this.  Somehow and unbelievably to my very novice level, I was able to apply your notebook examples to my own jupyter notebook, after watching your presentation on the Kaggle grandmaster powerhour youtube video.  My learning journey for me was a significant leap JUST to get to the Train Model stage w/o throwing errors! Unfortunately, however, I wasn't able to use the GPU strategy, as this iMac w TF just won't allow it as I've learned.  I'll give it another run on my linux laptop tomorrow to compare time and accuracy stats just for fun.  I sure did appreciate your example, and will be sure to spend more time studying and unpacking everything going forward. </p>",
      "rawMarkdown": "Thank you very much, I learned a lot from this.  Somehow and unbelievably to my very novice level, I was able to apply your notebook examples to my own jupyter notebook, after watching your presentation on the Kaggle grandmaster powerhour youtube video.  My learning journey for me was a significant leap JUST to get to the Train Model stage w/o throwing errors! Unfortunately, however, I wasn't able to use the GPU strategy, as this iMac w TF just won't allow it as I've learned.  I'll give it another run on my linux laptop tomorrow to compare time and accuracy stats just for fun.  I sure did appreciate your example, and will be sure to spend more time studying and unpacking everything going forward. "
    },
    {
      "id": 901719,
      "postDate": "2020-06-25T16:39:19.257Z",
      "content": "<p>Hi, Chris. Thanks for this discussion. I'd a huge jump when compared to now. Thanks for sharing. I learned a lot from this kernel. Even top 1,2,3 ranks are amused by your logic. Again, Thank you for sharing. \nBest score:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2Ffeff055aa7817f8e008faf7edd387950%2FScreenshot%202020-06-25%20at%2010.05.29%20PM.png?generation=1593103206943415&amp;alt=media\" alt=\"\"></p>\n\n<p>Previous one:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2F5144f1dc8e6c29f34a95042e6918cb67%2FScreenshot%202020-06-25%20at%2010.04.38%20PM.png?generation=1593103103663912&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi, Chris. Thanks for this discussion. I'd a huge jump when compared to now. Thanks for sharing. I learned a lot from this kernel. Even top 1,2,3 ranks are amused by your logic. Again, Thank you for sharing. \nBest score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2Ffeff055aa7817f8e008faf7edd387950%2FScreenshot%202020-06-25%20at%2010.05.29%20PM.png?generation=1593103206943415&amp;alt=media)\n\nPrevious one:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2F5144f1dc8e6c29f34a95042e6918cb67%2FScreenshot%202020-06-25%20at%2010.04.38%20PM.png?generation=1593103103663912&amp;alt=media)\n"
    },
    {
      "id": 776722,
      "postDate": "2020-03-17T15:35:46.437Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> congrats! Can you share the code for cam-cutmix btw (sorry, if you already shared!). I actually wanted to implement that - but, ultimately couldn't make it to work during the competition!</p>",
      "rawMarkdown": "@cdeotte congrats! Can you share the code for cam-cutmix btw (sorry, if you already shared!). I actually wanted to implement that - but, ultimately couldn't make it to work during the competition!"
    },
    {
      "id": 776102,
      "postDate": "2020-03-17T05:51:26.167Z",
      "content": "<p>Your CAM CutMix is awesome. I'm curious about the implementation. Did you use cutmix in your first model? And how much boost did CAM Cutmix gain?</p>",
      "rawMarkdown": "Your CAM CutMix is awesome. I'm curious about the implementation. Did you use cutmix in your first model? And how much boost did CAM Cutmix gain?",
      "replies": [
        {
          "id": 776106,
          "postDate": "2020-03-17T05:54:18.760Z",
          "content": "<p>Yes, first model used CutMix. Second model gained 0.001 CV and LB from CAM CutMix (compared to second model using regular CutMix).</p>",
          "rawMarkdown": "Yes, first model used CutMix. Second model gained 0.001 CV and LB from CAM CutMix (compared to second model using regular CutMix)."
        }
      ]
    },
    {
      "id": 776098,
      "postDate": "2020-03-17T05:50:00.993Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 778178,
      "postDate": "2020-03-18T07:54:26.790Z",
      "content": "<p>thanks for the clear code</p>",
      "rawMarkdown": "thanks for the clear code",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 776316,
      "author_name": "Chulvi",
      "author_url": "",
      "post_date": "2020-03-17T09:52:23.180000",
      "content": "<p>Just amazing 🤓 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2726150%2F80f84f0436d148bd2ea5e5ee8e7a6eef%2Fmagic.JPG?generation=1584438727758691&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 776503,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-03-17T12:58:07.530000",
          "content": "<p>You should definitely write a summary of your approach. Also mention about the <code>0.9643</code> private LB approach. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 776508,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-17T13:00:13.003000",
          "content": "<p>wow, what a jump by PP</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 776625,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T14:16:27.227000",
          "content": "<p>Wow private LB 0.9643 incredible. Too bad we didn't know that public LB in the 0.98 could place so high in private. I also had private LB over 0.96 on some of my public LB 0.98 solutions.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 781110,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-03-20T23:37:39.267000",
      "content": "<p>I just applied PP to the highest scoring public notebook. It turns the 8th place solution into the 2nd place solution. The private LB becomes 0.9704! Notebook <a href=\"https://www.kaggle.com/bamps53/private0-9552-tpu-keras-metric-learning\">here</a>. Use <code>EXP = [-1.2, -1.2, -0.5]</code>.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc5ac8746ceed9dad3100932b825df2d6%2FScreen%20Shot%202020-03-20%20at%204.26.04%20PM.png?generation=1584747191411505&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 781200,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-21T02:59:47.833000",
          "content": "<p>Thanks!! I update mine with your postprocess:)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 782211,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-03-22T03:29:11.693000",
          "content": "<p>Your PP method just too powerful.. I simply trained 3 model(efficientnetB3) with cutmix, mixup and cutout augmentations for each class(root, vowel and consonant). And I got 0.9560 private score and 0.9759 public score..This method definitely the best technique I learned from this competition. Thanks again, Chris!!<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2F6d96a469346e03e09daa4d16a1bd9af2%2F2020-03-22%2011.26.18.png?generation=1584847720533229&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 776498,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-03-17T12:53:23.970000",
      "content": "<p>I love the post processing. I am not a fan of calling PP hacking though as I believe that optimizing the predictions towards the metric is an integral part of machine learning. If you adapt your loss functions to better capture the metric, you are also not calling it hacking. </p>\n\n<p>If the hosts choose a metric that overweights rare cases, then I would expect them to care about rare cases. So being more precise on them seems important. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 776614,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T14:04:48.970000",
          "content": "<p>Good point I agree. I chose the title \"Hacking Macro Recall\" to catch more attention. This post process is just mathematical optimization which should be done in every comp. </p>\n\n<p>I was sure that your team was using this PP when you asked for teammates (14 days ago <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#762152\">here</a>. your CV was 0.9890 and LB 0.9845) because your CV and LB gap was so small. Everyone else's gap was around 0.010 while yours and mine (with PP) were around 0.05.</p>\n\n<p>Understanding metrics is overlooked at Kaggle. In half of my past competitions, you could take the top public kernel and post process the predictions by optimizing to the metric and get silver medal or above. Google Quest Q&amp;A with spearman rank correlation coefficient is great example. And Cloud comp and Steel comp are two great examples. Both use <strong>image-wise</strong> dice instead of the common <strong>batch-wise</strong> dice.</p>\n\n<p>Congrats to you and your team on another amazing finish. I am so impressed how The Zoo does so well in such a wide range of competition types. I would love to team up some day and learn from you guys!</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 776641,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-17T14:27:35.780000",
          "content": "<p>Thanks a lot Chris! As always you also did an amazing job and you are totally right that understanding the metric and adjusting to it still is sometimes an underappreciated aspect, but one of the most important ones to do well on LB.</p>\n\n<p>Regarding post processing: actually at that point, we only hardcoded a few extra C=3 and C=6 and that already gave us a boost of 30-40 points on LB. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 776249,
      "author_name": "Morphy",
      "author_url": "",
      "post_date": "2020-03-17T08:31:05.433000",
      "content": "<p>Congrats, Chris!. One question.  You mentioned  'Therefore your expected macro recall increase if you predict class 1 is 2.3e-4 = 0.70 * 1/1000 * 1/3. And your expected macro recall increase if you predict class 2 is 1e-3. Therefore you predict class 2'. What is 1/3 in the equation? </p>",
      "votes": 3,
      "replies": [
        {
          "id": 776461,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "2020-03-17T12:11:09.860000",
          "content": "<p>I was curious too,have you figured it out??</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776595,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T13:50:39.220000",
          "content": "<p>That was a typo. I corrected it. Thanks. It should be class 1 is <code>1e-4 = 0.70 * 1/1000 * 1/7</code> and class 2 is <code>4e-4 =  0.30 * 1/100 * 1/7</code>. It is <code>1/7</code> because there are 7 consonant diacritics in the macro average.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776037,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-17T05:01:53.377000",
      "content": "<p>Cool! This CAM CutMix and the refined prediction blow my mind! 👍  Congratulation! </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 776035,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-17T04:58:06.413000",
      "content": "<p>\"Let's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.\"</p>\n\n<p>how do you know the underlying prior probability of each class (e.g. class1 has 1000 samples, while class2 has 100)?\nwe can't probe the private test data in this case?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 776045,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T05:08:32.293000",
          "content": "<p>Yes we can. First you make predictions on the test data. Then you use those predictions as the prior probability and compute a second set or predictions by modifying the first set. (See code above).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776054,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2020-03-17T05:13:30.537000",
          "content": "<p>His distribution comes from the predicted results.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776055,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-17T05:14:40.377000",
          "content": "<p>thanks for the reply.</p>\n\n<p>we can test only public test data and not the private one. </p>\n\n<p>i suppose you have made predictions on the test data, so does it reveal that the private and public test data are similar distribution at your test? e.g. do have results like</p>\n\n<p>```\nno post processing:\nprivate lb = xxx\npublic lb = xxx</p>\n\n<p>post processing of fixing 1 class:\nprivate lb = xxx\npublic lb = xxx</p>\n\n<p>post processing of fixing 10 class:\nprivate lb = xxx\npublic lb = xxx</p>\n\n<p>etc</p>\n\n<p>```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 776065,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T05:20:57.230000",
          "content": "<p>We predict <strong>both</strong> public and private at the same time with our submitted notebooks. When your code makes test predictions it makes all 200,000 at once. That is both public and private. The post process uses all these 200,000 predictions to compute distribution. Afterward, we only see public LB score but submitted notebook computed public LB and private LB together.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776083,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T05:40:11.813000",
          "content": "<p>Here are results for fixing all 168 roots, 11 vowels and 7 consonants:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7d138a76f20b20e65a83e2d8c2b88275%2Fsub.jpg?generation=1584419035681165&amp;alt=media\" alt=\"image\"></p>\n\n<p>We do have some submissions that only fix all root or only fix all vowel or only fix all consonant. Basically each helps an equal amount. We don't have results for just fixing individual classes (within root or vowel or consonant).</p>\n\n<p>You sort of need to fix them all at once. If you make your model more likely to predict a particular rare class, you must also tell it what the other rare classes are. Otherwise it steals predictions from a rare class to fill the other rare classes. Whereas when you do it all at once, rare classes don't steal from each other but rather steal from common classes.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 776057,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2020-03-17T05:16:02.793000",
      "content": "<p>Congrats Chris, thanks for sharing.  I thought about the macro recall thing a few days ago but didn't figure out how to do a correct post processing 😂 . Your trick is pretty good 👍</p>",
      "votes": 4,
      "replies": [
        {
          "id": 776061,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T05:18:32.073000",
          "content": "<p>Thanks Venn. Sorry about your shakup. Your team built a great model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 782038,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-03-21T21:40:12.160000",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> and congratulation. CAM CutMix is interesting. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 778975,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-03-18T22:45:10.703000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you for sharing, it's really magic! I'm now trying to reproduce it, this post process is supposed to be used with all 200,000 prediction, not with every batch, right?\nI'm encountering too much memory issue, was it possible to store (200,000, 168) numpy array for your case?</p>\n\n<p>Thanks, I love your cam cutmix idea too!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 778980,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-18T22:56:08.333000",
          "content": "<p>Thanks. Yes the entire 200,000 predictions. We predict in batches of 512, so <code>batch_preds</code> are softmax output with size <code>(512, 168 + 11 + 7)</code>. Then we add these to a list <code>preds.append(batch_preds)</code>. And afterwards convert to numpy array with <code>preds = np.vstack(preds)</code>. That final numpy array is 200,000 rows by 168+11+7 columns. Since it is <code>float32</code>, that is only <code>148MB = 200000 * (168+11+7) * 4</code>. Not much memory at all.</p>\n\n<p>Make sure you are predicting in batches and not predicting all 200,000 at once.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 779138,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-19T03:45:47.860000",
          "content": "<p>I finally figured out the problem was in my image loading rather than preds numpy array.\nBut still struggling to make it works. It seems EXP = -1.2 is not suitable for my model, I might need to tune it by validation.\nAnyway, thanks for sharing!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 779148,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-19T03:59:12.540000",
          "content": "<p>Some models don't like -1.2. Try setting root EXP to -0.9, vowel EXP to -1.1 and consonant EXP to -0.6. That worked best for our team ensemble.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 779150,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-19T04:02:36.770000",
          "content": "<p>Also make sure that your model is outputting probabilities. If is it outputting logits. First apply <code>scipy.special.softmax(preds, axis=1)</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 777856,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-03-18T01:53:43.387000",
      "content": "<p>Thank you for always sharing detailed, nice explanation write-up!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 777842,
      "author_name": "ccchang",
      "author_url": "",
      "post_date": "2020-03-18T01:45:01.267000",
      "content": "<p>Thanks for sharing! the ideas are wonderful and easy to implement! \nI didn't survive from the shake up 😂 and all I can do is keep learning from the best like you, \nCongrats to you and your team!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 777854,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-18T01:53:15.937000",
          "content": "<p>Thanks Geroko. I'm sorry about your shakeup. Your model is great with public LB 0.9909. I bet if you apply PP, you will have one of the best models on private LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 776603,
      "author_name": "Hugo Tong",
      "author_url": "",
      "post_date": "2020-03-17T13:56:43.547000",
      "content": "<p>Amazing postprocessing on the final prediction! </p>\n\n<p>Did you try to compute the prior distribution from train set true label and apply it on test set, instead of computing the distribution from test set prediction and applying on test set? Just curious about the robustness of this magic!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776619,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T14:07:24.120000",
          "content": "<p>Yes. Using train distribution achieves similar result. This confirms that train and test distribution are similar. I think the most difference is between public test and private test. I believe some rare classes become even more rare making this PP even more important for private LB.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 776491,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-03-17T12:41:18.727000",
      "content": "<p>Wow, congrats, both post processing and cam cutmix are great.  We where on the post processing track, but certainly didn't think of cam based stuff.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776623,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T14:14:02.720000",
          "content": "<p>Thanks CPMP. Congrats to you and Kaz on a strong finish.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 776299,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2020-03-17T09:31:47.097000",
      "content": "<p>Oh my..this just genius, so simple! so powerful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776241,
      "author_name": "Bryce1010",
      "author_url": "",
      "post_date": "2020-03-17T08:27:28.310000",
      "content": "<p>Congrats your team and thanks for detailed solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776171,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-03-17T07:03:48.377000",
      "content": "<p>Congrats your team and thanks for detailed solution <a href=\"/cdeotte\">@cdeotte</a>!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 776073,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "2020-03-17T05:29:46.443000",
      "content": "<p>Congratulation. In CAM CutMix, you randomly remove 15% to 25% by CAM map of one of random selected component, have you also do some operation on target?  (e.g. reduce the weight of the selected component)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 776080,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T05:33:40.910000",
          "content": "<p>I change target in the same way as CutMix. For example, if i replace 17% of image 1 with image 2. Then i change the label to <code>0.83 * label_1 + 0.17 * label_2</code> where labels are one hot encoded.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 789015,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2020-03-28T09:14:35.623000",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> for sharing and congrats on your 14th(gold) place!</p>\n\n<p>I tried your post-processing in my late submission (I added results in <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136815\">my solution</a>).</p>\n\n<p>By your magic, my score increased from 0.9536 (10th) to <strong>0.9653(3rd)!!!</strong>  Amaging!\nYour magic is a really magic!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2Fd416d7853979e1c35a4e6d66a8307283%2Fbest_private_score_in_late_sub.png?generation=1585386609111447&amp;alt=media\" alt=\"best private score in late submissions\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 789740,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-29T00:05:10.063000",
          "content": "<p>Awesome!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 781677,
      "author_name": "YS",
      "author_url": "",
      "post_date": "2020-03-21T15:02:57.443000",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> ,\nThank you for your sharing.\nAbout the PP method, I have some questions.</p>\n\n<ol>\n<li><p>Does it have to be applied after ensemble? I applied the PP for each model result and ensemble them, but private LB drop 0.1. Or maybe I need to fine tune EXP?</p></li>\n<li><p>Have you try this PP to other compete? It seems like works for other competes.</p></li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 781708,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-21T15:18:57.247000",
          "content": "<p>You can apply it before or after ensemble. If your ensemble is a linear combination then it is mathematically the same. If LB drops try adding <code>EXP = 0.2</code> first to see if that helps. Also make sure your model is outputting <code>softmax</code> which are probabilities and not <code>logits</code> which are log odds.</p>\n\n<p>After <code>EXP = 0.2</code> works, try -0.4, -0.6, etc. At some point the LB will drop. That is usually caused by the consonant EXP being too large. So at that point leave the consonant EXP at a good value and increase the others. </p>\n\n<p>For example if <code>EXP = -0.5</code> works for all and <code>EXP = -0.6</code> does not work for all. Then try <code>EXP = -0.6, -0.6, -0.5</code> and then keep increasing with <code>EXP = -0.8, -0.8, -0.5</code> etc. When i write three <code>EXP</code>, i mean assign the root, vowel, consonant <code>EXP</code> with different values.</p>\n\n<p>You can find optimal <code>EXP</code> using your local CV or the public LB. And if it increases those then it will increase private LB too. You can even make plots of the effects of <code>EXP</code> on CV for root, vowel, and consonant using your local CV.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 781764,
          "author_name": "YS",
          "author_url": "",
          "post_date": "2020-03-21T16:01:49.080000",
          "content": "<p>Thank you for your clear description.\nI will try to fine tune EXP.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 779411,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2020-03-19T10:17:07.153000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> congratulations and thank you for sharing. \nOur team had thought about pp of Macro Recall, but we didn't have a good idea，so we suffered from shakeup.\nI tried you method, it is really impressived, EXP=-1.2 did not work for me, so I tried setting root EXP to -0.9, vowel EXP to -1.1 and consonant EXP to -0.6, and got public LB: 0.9920-&gt;0.9931, private LB: 0.9360-&gt;0.9528. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 781109,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-20T23:30:16.973000",
          "content": "<p>Thanks Gary</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776760,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-03-17T16:10:34.487000",
      "content": "<p>Hey Chris nice job! Thanks for sharing. Can you please explain at a high level why you created matricies and what the matrix multiplication is doing for the less mathematically inclined. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 776800,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T16:35:52.390000",
          "content": "<p>The matrix is just a way to do the following. Say that your softmax output is <code>x = [0, 0.7, 0.3, 0]</code> and you want to multiply those 4 numbers by these 4 numbers <code>2, 5, 3, 1</code> respectively. Then you create a diagonal matrix with the 4 numbers as follows:</p>\n\n<pre><code>mat = [\n[2, 0, 0, 0]\n[0, 5, 0, 0]\n[0, 0, 3, 0]\n[0, 0, 0, 1]]\n</code></pre>\n\n<p>The matrix is created by calling <code>mat = np.diag([2, 5, 3, 1])</code> and then <code>x.dot(mat)</code> multiplies the softmax output by the 4 numbers. The final result is <code>[0*2, 0.7*5, 0.3*3, 0*1]</code>.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 776116,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-03-17T06:08:15.007000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Congratulations! Such a cool method! I tried to leverage the recall information by designing custom loss but failed hahaha. I have a question: it seems that the expected value of recall increase has to be computed based on a ready distribution estimation, while the final distribution estimation is really computed after the recall hack...So using this trick is kind of like solving a problem with recursive solution dependence. How do you solve this? I take that you simply use first-time argmax predictions as estimation for distribution? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 776135,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T06:22:57.487000",
          "content": "<p>Yes. Use first-time argmax as estimation. Furthermore, we considered training a model by upsampling all classes to have equal distribution so that our first-time argmax would be more unbiased but we didn't have time to do that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776539,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-17T13:15:39.723000",
          "content": "<p>Thank you very much. I personaly think this hack is the most code/complexity-efficient hack for improving score in this competition hahaha.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776109,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-17T06:00:56.257000",
      "content": "<p>OMG! After the post-processing, my private LB and public LB massively increase by 0.022 and 0.048 respectively!! \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F8a63beec93df3988d639a9c258ce86ac%2FScreen%20Shot%202020-03-16%20at%2010.55.41%20PM.png?generation=1584424713221162&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 776133,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T06:20:45.903000",
          "content": "<p>Wow great increase. You can also try <code>EXP = -1.2</code> that works best for me. It will probably increase you even higher!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776179,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-03-17T07:17:20.967000",
          "content": "<p>This post-process is really impressive! 👍 👍\n I use exp=-1 to my previous top LB kernel, the private LB score increases by 0.0247, to 0.9561.! OMG!! I'll try exp=-1.2 tomorrow! Many thanks!!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F74bfa168ba8e34b7a96dfe808d12f51d%2FIMG_5372.PNG?generation=1584429338367971&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F0f81ab5a845b2127baa1ed60c8e28e41%2FIMG_5371.jpg?generation=1584429338405133&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776627,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T14:17:57.123000",
          "content": "<p>Wow, such a high private LB score. Awesome!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 776665,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-03-17T14:41:56.340000",
          "content": "<p>What is awesome is your post-processing! This is definitely one of the most effective methods what I've learned from this competition! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 776015,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-17T04:40:24.947000",
      "content": "<p>WOW THIS IS REALLLLL\nMMMMMMAGICCCCCC!</p>\n\n<p>Congratulation! ;)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 776019,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T04:43:44.020000",
          "content": "<p>Thanks. Congrats to you Qishen. It was impressive how you held 1st place for so long. I discovered this trick a month ago and I thought this was why you were at 0.9930 while everyone else was at 0.9900. But after reading your solution, it appears that you didn't use this trick.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 776030,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-17T04:54:20.083000",
          "content": "<p>I thought that distribution postprocessing may work, but it's too dangerous, so I even didn't give a try.\nBut I'm wrong, it works so well -- If top 1 use it, he may have gone to 0.98+ or something 😂</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 776032,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T04:57:30.100000",
          "content": "<p>Yes it works well. Notice that the post processing computes the entire test distribution which includes both the private and public test data. Therefore it does not overfit the public LB.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 776046,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-17T05:10:02.677000",
          "content": "<blockquote>\n  <p>computes the entire test distribution which includes both the private and public test data</p>\n</blockquote>\n\n<p>Didn't realized it, thanks! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 778645,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-18T15:59:28.990000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I've got a question. Have you tried using this postprocessing twice?</p>\n\n<blockquote>\n  <p>First you make predictions on the test data. Then you use those predictions as the prior probability and compute a second set or predictions by modifying the first set. (See code above).</p>\n</blockquote>\n\n<p>Then you use the second set or predictions as the prior probability and compute a third one?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945635,
      "author_name": "Evan Harrison",
      "author_url": "",
      "post_date": "2020-07-26T03:35:57.870000",
      "content": "<p>Thank you very much, I learned a lot from this.  Somehow and unbelievably to my very novice level, I was able to apply your notebook examples to my own jupyter notebook, after watching your presentation on the Kaggle grandmaster powerhour youtube video.  My learning journey for me was a significant leap JUST to get to the Train Model stage w/o throwing errors! Unfortunately, however, I wasn't able to use the GPU strategy, as this iMac w TF just won't allow it as I've learned.  I'll give it another run on my linux laptop tomorrow to compare time and accuracy stats just for fun.  I sure did appreciate your example, and will be sure to spend more time studying and unpacking everything going forward. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 901719,
      "author_name": "karan",
      "author_url": "",
      "post_date": "2020-06-25T16:39:19.257000",
      "content": "<p>Hi, Chris. Thanks for this discussion. I'd a huge jump when compared to now. Thanks for sharing. I learned a lot from this kernel. Even top 1,2,3 ranks are amused by your logic. Again, Thank you for sharing. \nBest score:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2Ffeff055aa7817f8e008faf7edd387950%2FScreenshot%202020-06-25%20at%2010.05.29%20PM.png?generation=1593103206943415&amp;alt=media\" alt=\"\"></p>\n\n<p>Previous one:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2F5144f1dc8e6c29f34a95042e6918cb67%2FScreenshot%202020-06-25%20at%2010.04.38%20PM.png?generation=1593103103663912&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776722,
      "author_name": "Partha ",
      "author_url": "",
      "post_date": "2020-03-17T15:35:46.437000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> congrats! Can you share the code for cam-cutmix btw (sorry, if you already shared!). I actually wanted to implement that - but, ultimately couldn't make it to work during the competition!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776102,
      "author_name": "syoya",
      "author_url": "",
      "post_date": "2020-03-17T05:51:26.167000",
      "content": "<p>Your CAM CutMix is awesome. I'm curious about the implementation. Did you use cutmix in your first model? And how much boost did CAM Cutmix gain?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 776106,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-03-17T05:54:18.760000",
          "content": "<p>Yes, first model used CutMix. Second model gained 0.001 CV and LB from CAM CutMix (compared to second model using regular CutMix).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776098,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T05:50:00.993000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 778178,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-18T07:54:26.790000",
      "content": "",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "776008": "This is my first computer vision gold medal. I'm very excited!\n\nThank you Kaggle and Bengali.AI for hosting a fun comp. Thank you teammates Bojan, Shai, Yasin, Jahmed ( @tunguz @sgalib @mykttu @jasemahmed ). I had a blast working with you guys! Thank you Nvidia for providing GPU compute. Below are my contributions to our team's solution. The rest of the team will share more.\n\n# Competition Metric Explained\n\nThis competition's metric is macro recall. That means you compute the recall of each class individually and average them. Most importantly `recall = found / exist`. There is no penalty for making a false positive! Making more positive predictions for one class can only increase that class' recall never decrease!\n\nBelow is a paradoxical example. Imagine that your are predicting one of seven consonant diacritic classes and your CNN outputs the following probabilities:\n\n## Softmax = [0.0, 0.7, 0.3, 0.0, 0.0, 0.0, 0.0]\n\n## Question: Do you predict class 1 or class 2?\n\nLet's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.\n\nTherefore your expected macro recall increase if you predict class 1 is `1e-4 = 0.70 * 1/1000 * 1/7`. And your expected macro recall increase if you predict class 2 is `4e-4 = 0.30 * 1/100 * 1/7`. Therefore you predict class 2.\n\n## Answer: You predict class 2 not class 1\nBy adjusting your predictions in this fashion you can gain a massive 0.0027 public LB increase and 0.0241 private LB increase!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7d138a76f20b20e65a83e2d8c2b88275%2Fsub.jpg?generation=1584419035681165&amp;alt=media)\n\n\n### Code\nTry the following post process on your model to see how much it increases your public and private LB. \n    \n    preds = model.predict(X_test)\n    p0 = np.argmax(preds[0],axis=1)\n    p1 = np.argmax(preds[1],axis=1)\n    p2 = np.argmax(preds[2],axis=1)\n\n    EXP = -1.2\n\n    s = pd.Series(p0)\n    vc = s.value_counts().sort_index()\n    df = pd.DataFrame({'a':np.arange(168),'b':np.ones(168)})\n    df.b = df.a.map(vc)\n    df.fillna(df.b.min(),inplace=True)\n    mat1 = np.diag(df.b.astype('float32')**EXP)\n\n    s = pd.Series(p1)\n    vc = s.value_counts().sort_index()\n    df = pd.DataFrame({'a':np.arange(11),'b':np.ones(11)})\n    df.b = df.a.map(vc)\n    df.fillna(df.b.min(),inplace=True)\n    mat2 = np.diag(df.b.astype('float32')**EXP)\n\n    s = pd.Series(p2)\n    vc = s.value_counts().sort_index()\n    df = pd.DataFrame({'a':np.arange(7),'b':np.ones(7)})\n    df.b = df.a.map(vc)\n    df.fillna(df.b.min(),inplace=True)\n    mat3 = np.diag(df.b.astype('float32')**EXP)\n\n    p0 = np.argmax( preds[0].dot(mat1), axis=1)\n    p1 = np.argmax( preds[1].dot(mat2), axis=1)\n    p2 = np.argmax( preds[2].dot(mat3), axis=1)\n\n# Chris Model\nOur team's solution is an ensemble. For my model I trained a 128x256 efficientNetB6 150 epochs with CAM CutMix in addition to basic rotation, scale, shift, cutout, and cutmix. Training took 24 hours on four Nvidia V100 GPUs. It's CV is 0.9967, public LB 0.9916, and private LB 0.9544. \n  \nCAM CutMix is where you find the class activation maps of the images and then remove the most important parts of the image and replace them with another image. This challenges your CNN and helps it generalize.\n\n## CAM Maps\nFirst you train one model to produce CAM maps. (First model was 256x256 efficientNetB6). Next you use CAM maps to train a second model. (Second model was 128x256 efficientNetB6). Here are some CAM maps: (More pictures [here][1]).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb3a043c661bc766967d9369f51bb9059%2FScreen%20Shot%202020-03-16%20at%207.44.04%20PM.png?generation=1584414144102358&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F129399f9d691c9cf3323d73f82c7a262%2FScreen%20Shot%202020-03-16%20at%207.42.46%20PM.png?generation=1584414156528553&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F2c9e3348089070dadf8ecf422c3f0a23%2FScreen%20Shot%202020-03-16%20at%207.43.27%20PM.png?generation=1584414169952396&amp;alt=media)\n\n## CAM CutMix\n\nCutMix is a combination of two images. The first image is displayed as yellow below to help us visualize it. First one component either root, vowel, consonant is randomly selected. Next a random percentage from 15% to 25% is selected. Next that percentage of the first image is removed using the chosen component type's CAM map. Finally the same region from a second randomly selected image is inserted. The second image is displayed as blue below. 50% of images use CAM CutMix and 50% use regular CutMix. (CAM CutMix increased CV and LB by 0.001 over regular CutMix).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F253a777c731827a5bb5ced8cdf85558a%2FScreen%20Shot%202020-03-16%20at%207.44.42%20PM.png?generation=1584414301224112&amp;alt=media)\n\n[1]: https://www.kaggle.com/c/bengaliai-cv19/discussion/136025",
    "776316": "Just amazing 🤓 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2726150%2F80f84f0436d148bd2ea5e5ee8e7a6eef%2Fmagic.JPG?generation=1584438727758691&amp;alt=media)\n",
    "781110": "I just applied PP to the highest scoring public notebook. It turns the 8th place solution into the 2nd place solution. The private LB becomes 0.9704! Notebook [here][1]. Use `EXP = [-1.2, -1.2, -0.5]`.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc5ac8746ceed9dad3100932b825df2d6%2FScreen%20Shot%202020-03-20%20at%204.26.04%20PM.png?generation=1584747191411505&amp;alt=media)\n\n[1]: https://www.kaggle.com/bamps53/private0-9552-tpu-keras-metric-learning\n",
    "776498": "I love the post processing. I am not a fan of calling PP hacking though as I believe that optimizing the predictions towards the metric is an integral part of machine learning. If you adapt your loss functions to better capture the metric, you are also not calling it hacking. \n\nIf the hosts choose a metric that overweights rare cases, then I would expect them to care about rare cases. So being more precise on them seems important. ",
    "776249": "Congrats, Chris!. One question.  You mentioned  'Therefore your expected macro recall increase if you predict class 1 is 2.3e-4 = 0.70 * 1/1000 * 1/3. And your expected macro recall increase if you predict class 2 is 1e-3. Therefore you predict class 2'. What is 1/3 in the equation? ",
    "776037": "Cool! This CAM CutMix and the refined prediction blow my mind! 👍  Congratulation! ",
    "776035": "\"Let's say that there are 1000 samples in class 1 and 100 samples in class 2. Then if you predict class 1 you have a 70% chance of increasing your class 1 recall by 1/1000. If you predict class 2 then you have a 30% chance of increasing your class 2 recall by 1/100.\"\n\nhow do you know the underlying prior probability of each class (e.g. class1 has 1000 samples, while class2 has 100)?\nwe can't probe the private test data in this case?",
    "776057": "Congrats Chris, thanks for sharing.  I thought about the macro recall thing a few days ago but didn't figure out how to do a correct post processing 😂 . Your trick is pretty good 👍",
    "782038": "Thanks @cdeotte and congratulation. CAM CutMix is interesting. ",
    "778975": "@cdeotte Thank you for sharing, it's really magic! I'm now trying to reproduce it, this post process is supposed to be used with all 200,000 prediction, not with every batch, right?\nI'm encountering too much memory issue, was it possible to store (200,000, 168) numpy array for your case?\n\nThanks, I love your cam cutmix idea too!!",
    "777856": "Thank you for always sharing detailed, nice explanation write-up!",
    "777842": "Thanks for sharing! the ideas are wonderful and easy to implement! \nI didn't survive from the shake up 😂 and all I can do is keep learning from the best like you, \nCongrats to you and your team!",
    "776603": "Amazing postprocessing on the final prediction! \n\nDid you try to compute the prior distribution from train set true label and apply it on test set, instead of computing the distribution from test set prediction and applying on test set? Just curious about the robustness of this magic!",
    "776491": "Wow, congrats, both post processing and cam cutmix are great.  We where on the post processing track, but certainly didn't think of cam based stuff.",
    "776299": "Oh my..this just genius, so simple! so powerful!",
    "776241": "Congrats your team and thanks for detailed solution!",
    "776171": "Congrats your team and thanks for detailed solution @cdeotte!",
    "776073": "Congratulation. In CAM CutMix, you randomly remove 15% to 25% by CAM map of one of random selected component, have you also do some operation on target?  (e.g. reduce the weight of the selected component)",
    "789015": "Thanks @cdeotte for sharing and congrats on your 14th(gold) place!\n\nI tried your post-processing in my late submission (I added results in [my solution](https://www.kaggle.com/c/bengaliai-cv19/discussion/136815)).\n\nBy your magic, my score increased from 0.9536 (10th) to **0.9653(3rd)!!!**  Amaging!\nYour magic is a really magic!\n\n![best private score in late submissions](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F473234%2Fd416d7853979e1c35a4e6d66a8307283%2Fbest_private_score_in_late_sub.png?generation=1585386609111447&amp;alt=media)\n",
    "781677": "Hi @cdeotte ,\nThank you for your sharing.\nAbout the PP method, I have some questions.\n\n1. Does it have to be applied after ensemble? I applied the PP for each model result and ensemble them, but private LB drop 0.1. Or maybe I need to fine tune EXP?\n\n2. Have you try this PP to other compete? It seems like works for other competes.",
    "779411": "@cdeotte congratulations and thank you for sharing. \nOur team had thought about pp of Macro Recall, but we didn't have a good idea，so we suffered from shakeup.\nI tried you method, it is really impressived, EXP=-1.2 did not work for me, so I tried setting root EXP to -0.9, vowel EXP to -1.1 and consonant EXP to -0.6, and got public LB: 0.9920-&gt;0.9931, private LB: 0.9360-&gt;0.9528. ",
    "776760": "Hey Chris nice job! Thanks for sharing. Can you please explain at a high level why you created matricies and what the matrix multiplication is doing for the less mathematically inclined. ",
    "776116": "@cdeotte Congratulations! Such a cool method! I tried to leverage the recall information by designing custom loss but failed hahaha. I have a question: it seems that the expected value of recall increase has to be computed based on a ready distribution estimation, while the final distribution estimation is really computed after the recall hack...So using this trick is kind of like solving a problem with recursive solution dependence. How do you solve this? I take that you simply use first-time argmax predictions as estimation for distribution? ",
    "776109": "OMG! After the post-processing, my private LB and public LB massively increase by 0.022 and 0.048 respectively!! \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4008529%2F8a63beec93df3988d639a9c258ce86ac%2FScreen%20Shot%202020-03-16%20at%2010.55.41%20PM.png?generation=1584424713221162&amp;alt=media)\n",
    "776015": "WOW THIS IS REALLLLL\nMMMMMMAGICCCCCC!\n\nCongratulation! ;)",
    "945635": "Thank you very much, I learned a lot from this.  Somehow and unbelievably to my very novice level, I was able to apply your notebook examples to my own jupyter notebook, after watching your presentation on the Kaggle grandmaster powerhour youtube video.  My learning journey for me was a significant leap JUST to get to the Train Model stage w/o throwing errors! Unfortunately, however, I wasn't able to use the GPU strategy, as this iMac w TF just won't allow it as I've learned.  I'll give it another run on my linux laptop tomorrow to compare time and accuracy stats just for fun.  I sure did appreciate your example, and will be sure to spend more time studying and unpacking everything going forward. ",
    "901719": "Hi, Chris. Thanks for this discussion. I'd a huge jump when compared to now. Thanks for sharing. I learned a lot from this kernel. Even top 1,2,3 ranks are amused by your logic. Again, Thank you for sharing. \nBest score:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2Ffeff055aa7817f8e008faf7edd387950%2FScreenshot%202020-06-25%20at%2010.05.29%20PM.png?generation=1593103206943415&amp;alt=media)\n\nPrevious one:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1698259%2F5144f1dc8e6c29f34a95042e6918cb67%2FScreenshot%202020-06-25%20at%2010.04.38%20PM.png?generation=1593103103663912&amp;alt=media)\n",
    "776722": "@cdeotte congrats! Can you share the code for cam-cutmix btw (sorry, if you already shared!). I actually wanted to implement that - but, ultimately couldn't make it to work during the competition!",
    "776102": "Your CAM CutMix is awesome. I'm curious about the implementation. Did you use cutmix in your first model? And how much boost did CAM Cutmix gain?",
    "776098": "",
    "778178": "thanks for the clear code"
  }
}