{
  "id": 313981,
  "title": "Answering - \"Can CNNs Add Numbers?!\"  🔢 ➡ 🧠 ➡ 🔟 ➡ 🥳🎉",
  "url": "/competitions/ultra-mnist/discussion/313981",
  "author_name": "",
  "post_date": "2022-03-20T07:36:34.891693200Z",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p><img src=\"https://i.imgur.com/R0SdZzg.png\" alt=\"preds\"></p>\n<h2>Can CNNs add numbers without any background in mathematics?</h2>\n<h3>Take a step back and ask a fundamental question - is this a <strong>multi-class</strong> classification task?</h3>\n<blockquote>\n  <ul>\n  <li>In order, to build an elegant pipeline (atleast for ultra-mnist problem), we should be able to answer fundamental question like - \"Can CNNs add numbers without any background in mathematics?\" or \"Is this even a multi-class problem at all?\"</li>\n  <li>Intutively, CNNs must be able to add numbers as it is an <code>AND</code> operation between multiple <em>visual-features</em>. This notebook is proof of concept for the same.</li>\n  <li>If CNNs cannot solve, what method would slove the problem? This is a starter notebook to answer the question.</li>\n  </ul>\n</blockquote>\n<p>The main issue with current data (in <strong>innovation track</strong>) is 1. irregular sizes of digits and 2. high resolution image.</p>\n<p>So, here I provide custom data with both the above problems removed and simplified by huge extent. If we can achieve high accuracy here, we will work our way backwards later. My new dataset properties are:</p>\n<pre><code>- Uniform sized digits (28x28 pixels)\n- Low resolution image (64x64 pixels)\n- Only 3 possible digit combination\n- Every digit appears only once (i.e summax = 7+8+9 = 24)\n</code></pre>\n<p><strong>Note1:</strong> I chose imbalanced becuase - who knows - organizers might have put imbalanced dataset in test folder. Even so, it will be a very good idea to train with balanced dataset. But beware - real world data is never balanced (another reason for moving ahead with imbalanced instead of balanced distribution in training set)</p>\n<blockquote>\n  <p>I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.</p>\n</blockquote>\n<p><strong>Note2:</strong> This poblem can <strong>easily solved by multi-label classification</strong> (which is same as extracting bbox and running pretrained model). But we are not doing that. Sticking to <strong>multi-class</strong> problem.</p>\n<p>Dataset preview</p>\n<p><img src=\"https://i.imgur.com/wJG0rvz.jpg\" alt=\"dataset-preview\"></p>\n<blockquote>\n  <ul>\n  <li>test set is randomly generated just like training set</li>\n  <li>only difference between them is sample size</li>\n  </ul>\n</blockquote>\n<p><strong>CODE:</strong> <a href=\"https://www.kaggle.com/code/l0new0lf/umnistv2-best-case-scenario-imbalanced\" target=\"_blank\">https://www.kaggle.com/code/l0new0lf/umnistv2-best-case-scenario-imbalanced</a></p>",
  "messages": [
    {
      "id": "1729541",
      "postDate": "03/20/2022 07:36:34",
      "content": "<p><img src=\"https://i.imgur.com/R0SdZzg.png\" alt=\"preds\"></p>\n<h2>Can CNNs add numbers without any background in mathematics?</h2>\n<h3>Take a step back and ask a fundamental question - is this a <strong>multi-class</strong> classification task?</h3>\n<blockquote>\n  <ul>\n  <li>In order, to build an elegant pipeline (atleast for ultra-mnist problem), we should be able to answer fundamental question like - \"Can CNNs add numbers without any background in mathematics?\" or \"Is this even a multi-class problem at all?\"</li>\n  <li>Intutively, CNNs must be able to add numbers as it is an <code>AND</code> operation between multiple <em>visual-features</em>. This notebook is proof of concept for the same.</li>\n  <li>If CNNs cannot solve, what method would slove the problem? This is a starter notebook to answer the question.</li>\n  </ul>\n</blockquote>\n<p>The main issue with current data (in <strong>innovation track</strong>) is 1. irregular sizes of digits and 2. high resolution image.</p>\n<p>So, here I provide custom data with both the above problems removed and simplified by huge extent. If we can achieve high accuracy here, we will work our way backwards later. My new dataset properties are:</p>\n<pre><code>- Uniform sized digits (28x28 pixels)\n- Low resolution image (64x64 pixels)\n- Only 3 possible digit combination\n- Every digit appears only once (i.e summax = 7+8+9 = 24)\n</code></pre>\n<p><strong>Note1:</strong> I chose imbalanced becuase - who knows - organizers might have put imbalanced dataset in test folder. Even so, it will be a very good idea to train with balanced dataset. But beware - real world data is never balanced (another reason for moving ahead with imbalanced instead of balanced distribution in training set)</p>\n<blockquote>\n  <p>I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.</p>\n</blockquote>\n<p><strong>Note2:</strong> This poblem can <strong>easily solved by multi-label classification</strong> (which is same as extracting bbox and running pretrained model). But we are not doing that. Sticking to <strong>multi-class</strong> problem.</p>\n<p>Dataset preview</p>\n<p><img src=\"https://i.imgur.com/wJG0rvz.jpg\" alt=\"dataset-preview\"></p>\n<blockquote>\n  <ul>\n  <li>test set is randomly generated just like training set</li>\n  <li>only difference between them is sample size</li>\n  </ul>\n</blockquote>\n<p><strong>CODE:</strong> <a href=\"https://www.kaggle.com/code/l0new0lf/umnistv2-best-case-scenario-imbalanced\" target=\"_blank\">https://www.kaggle.com/code/l0new0lf/umnistv2-best-case-scenario-imbalanced</a></p>",
      "rawMarkdown": "![preds](https://i.imgur.com/R0SdZzg.png)\n\n## Can CNNs add numbers without any background in mathematics?\n\n### Take a step back and ask a fundamental question - is this a **multi-class** classification task?\n\n> - In order, to build an elegant pipeline (atleast for ultra-mnist problem), we should be able to answer fundamental question like - \"Can CNNs add numbers without any background in mathematics?\" or \"Is this even a multi-class problem at all?\"\n> - Intutively, CNNs must be able to add numbers as it is an `AND` operation between multiple *visual-features*. This notebook is proof of concept for the same.\n> - If CNNs cannot solve, what method would slove the problem? This is a starter notebook to answer the question.\n\nThe main issue with current data (in **innovation track**) is 1. irregular sizes of digits and 2. high resolution image.\n\nSo, here I provide custom data with both the above problems removed and simplified by huge extent. If we can achieve high accuracy here, we will work our way backwards later. My new dataset properties are:\n\n    - Uniform sized digits (28x28 pixels)\n    - Low resolution image (64x64 pixels)\n    - Only 3 possible digit combination\n    - Every digit appears only once (i.e summax = 7+8+9 = 24)\n\n\n**Note1:** I chose imbalanced becuase - who knows - organizers might have put imbalanced dataset in test folder. Even so, it will be a very good idea to train with balanced dataset. But beware - real world data is never balanced (another reason for moving ahead with imbalanced instead of balanced distribution in training set)\n\n> I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.\n\n**Note2:** This poblem can **easily solved by multi-label classification** (which is same as extracting bbox and running pretrained model). But we are not doing that. Sticking to **multi-class** problem.\n\nDataset preview\n\n![dataset-preview](https://i.imgur.com/wJG0rvz.jpg)\n\n> - test set is randomly generated just like training set\n> - only difference between them is sample size\n\n\n**CODE:** https://www.kaggle.com/code/l0new0lf/umnistv2-best-case-scenario-imbalanced",
      "votes": null
    },
    {
      "id": "1729565",
      "postDate": "03/20/2022 08:04:28",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/l0new0lf\" target=\"_blank\">@l0new0lf</a> </p>\n<p>You state </p>\n<blockquote>\n  <p>\"<em>Intutively, CNNs must be able to add numbers as it is an AND operation between multiple visual-features. This notebook is proof of concept for the same.</em>\"</p>\n</blockquote>\n<p>Out of curiosity, if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4?</p>\n<p>BTW: </p>\n<blockquote>\n  <p>\"<em>I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.</em>\"</p>\n</blockquote>\n<p>Perhaps see the notebook <a href=\"https://www.kaggle.com/code/carlmcbrideellis/ultramnist-baseline-all-class-20\" target=\"_blank\">\"UltraMNIST baseline: All class 20\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @l0new0lf \n\nYou state \n\n> \"*Intutively, CNNs must be able to add numbers as it is an AND operation between multiple visual-features. This notebook is proof of concept for the same.*\"\n\nOut of curiosity, if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4?\n\nBTW: \n\n> \"*I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.*\"\n\nPerhaps see the notebook [\"UltraMNIST baseline: All class 20\"](https://www.kaggle.com/code/carlmcbrideellis/ultramnist-baseline-all-class-20).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1729576",
      "postDate": "03/20/2022 08:34:52",
      "content": "<p>Interesting question. This is a notebook I drafted yesterday night without much thought - so have to think more before I give a concrete explanation.</p>\n<p>But let me briefly explain my intuition - in early days, humans used to represent number <code>2</code> with eyes 👀 or wings because they represented <em>\"features\"</em> called <em>\"pairs\"</em>. But, they did not have much trouble with adding numbers. So, our network must learn to internally create a feature-vs-number table. for example eye features 👀 will map to 2,  Earth features 🌎 will map to 1 and so on. </p>\n<p>Here, in-order to predict <code>3</code> model must learn that features of 🌎 <code>AND</code>  👀  are present in the image (just like multi-label problem) - this addition operation might be learnt by final classification linear units by operating on feature maps.</p>\n<p>Now, just replace 👀 and 🌎 with weird (yet coherent and visually distinct) symbols call digits - with analogy, you may understand my intuition.</p>\n<p>Regards, <br>\nAR</p>",
      "rawMarkdown": "Interesting question. This is a notebook I drafted yesterday night without much thought - so have to think more before I give a concrete explanation.\n\nBut let me briefly explain my intuition - in early days, humans used to represent number `2` with eyes 👀 or wings because they represented *\"features\"* called *\"pairs\"*. But, they did not have much trouble with adding numbers. So, our network must learn to internally create a feature-vs-number table. for example eye features 👀 will map to 2,  Earth features 🌎 will map to 1 and so on. \n\nHere, in-order to predict `3` model must learn that features of 🌎 `AND`  👀  are present in the image (just like multi-label problem) - this addition operation might be learnt by final classification linear units by operating on feature maps.\n\nNow, just replace 👀 and 🌎 with weird (yet coherent and visually distinct) symbols call digits - with analogy, you may understand my intuition.\n\nRegards, \nAR",
      "votes": null
    },
    {
      "id": "1729583",
      "postDate": "03/20/2022 08:44:19",
      "content": "<p>And to answer the other half of the question</p>\n<blockquote>\n  <p>if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4? </p>\n</blockquote>\n<p>Firstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, <code>1, 1, 1</code> or <code>1, 1, 1, 1</code> is not possible. (please have a look at data distribution graphs). But you can try for distinct combinations like <code>3, 4, 5</code> or <code>9, 2, 5</code> with at most 3 digits</p>\n<p>Saying that, let me answer the better part of the question - </p>\n<blockquote>\n  <p>… (rather than simply memorize results)…</p>\n</blockquote>\n<p>During data generation process, for every image I am picking digits randomly and co-ordinates to position them randomly. So, to answer if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.</p>\n<p>Thanks,<br>\nAR</p>",
      "rawMarkdown": "And to answer the other half of the question\n\n> if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4? \n\nFirstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, `1, 1, 1` or `1, 1, 1, 1` is not possible. (please have a look at data distribution graphs). But you can try for distinct combinations like `3, 4, 5` or `9, 2, 5` with at most 3 digits\n\nSaying that, let me answer the better part of the question - \n> ... (rather than simply memorize results)...\n\nDuring data generation process, for every image I am picking digits randomly and co-ordinates to position them randomly. So, to answer if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.\n\nThanks,\nAR",
      "votes": null
    },
    {
      "id": "1729596",
      "postDate": "03/20/2022 09:12:21",
      "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>  I'm confident that the test set data is uniformly distributed because they are using simple accuracy formula for evaluation ; D</p>",
      "rawMarkdown": "carlmcbrideellis  I'm confident that the test set data is uniformly distributed because they are using simple accuracy formula for evaluation ; D",
      "votes": null
    },
    {
      "id": "1729605",
      "postDate": "03/20/2022 09:21:35",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/l0new0lf\" target=\"_blank\">@l0new0lf</a> </p>\n<blockquote>\n  <p>\"<em>if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.</em>\"</p>\n</blockquote>\n<p>Given that you have generated all of the data yourself, and you know exactly which digits are in each image, would it not be a simple matter to create a test dataset composed of images of sets of digits that you know are not in the training set?</p>\n<p>As for </p>\n<blockquote>\n  <p>\"<em>Firstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, 1, 1, 1 or 1, 1, 1, 1 is not possible.</em>\"</p>\n</blockquote>\n<p>That is exactly why I asked. It would be an easy to perform test to see if your network can actually perform the operation of addition, or has simply encoded a dictionary of maps.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @l0new0lf \n\n> \"*if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.*\"\n\nGiven that you have generated all of the data yourself, and you know exactly which digits are in each image, would it not be a simple matter to create a test dataset composed of images of sets of digits that you know are not in the training set?\n\nAs for \n\n> \"*Firstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, 1, 1, 1 or 1, 1, 1, 1 is not possible.*\"\n\nThat is exactly why I asked. It would be an easy to perform test to see if your network can actually perform the operation of addition, or has simply encoded a dictionary of maps.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1729619",
      "postDate": "03/20/2022 09:44:11",
      "content": "<p>Questions like these can better be answered by visualising predictions. Here is a very good example from my model. Because 6 looks like 0, my model predicts 8+9+0=17 instead of 8+9+6=23.</p>\n<p>Link: <a href=\"https://twitter.com/inf800/status/1505476873484087297?t=PRUjmnBIKUElGaKemCnYUw&amp;s=19\" target=\"_blank\">https://twitter.com/inf800/status/1505476873484087297?t=PRUjmnBIKUElGaKemCnYUw&amp;s=19</a></p>\n<p>(For some reason cannot upload images)</p>",
      "rawMarkdown": "Questions like these can better be answered by visualising predictions. Here is a very good example from my model. Because 6 looks like 0, my model predicts 8+9+0=17 instead of 8+9+6=23.\n\nLink: https://twitter.com/inf800/status/1505476873484087297?t=PRUjmnBIKUElGaKemCnYUw&s=19\n\n(For some reason cannot upload images)",
      "votes": null
    },
    {
      "id": "1729622",
      "postDate": "03/20/2022 09:49:10",
      "content": "<p>and were there examples of 8+9+0=17 in the training data?….</p>",
      "rawMarkdown": "and were there examples of 8+9+0=17 in the training data?....",
      "votes": null
    },
    {
      "id": "1729627",
      "postDate": "03/20/2022 09:54:25",
      "content": "<p>This sample is from 60k samples validation set used for training. Notebook has more details about data.</p>",
      "rawMarkdown": "This sample is from 60k samples validation set used for training. Notebook has more details about data.",
      "votes": null
    },
    {
      "id": "1729647",
      "postDate": "03/20/2022 10:18:51",
      "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> Just curious. So, do you think CNNs cannot add numbers? If so, why? It will be of great help for me to know other person's perspective.</p>",
      "rawMarkdown": "carlmcbrideellis Just curious. So, do you think CNNs cannot add numbers? If so, why? It will be of great help for me to know other person's perspective.",
      "votes": null
    },
    {
      "id": "1729663",
      "postDate": "03/20/2022 10:28:18",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/l0new0lf\" target=\"_blank\">@l0new0lf</a> </p>\n<p>You may be interested in the two publications mentioned in the topic <a href=\"https://www.kaggle.com/c/ultra-mnist/discussion/312631\" target=\"_blank\">\"The Neural Addition Unit (NAU)\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @l0new0lf \n\nYou may be interested in the two publications mentioned in the topic [\"The Neural Addition Unit (NAU)\"](https://www.kaggle.com/c/ultra-mnist/discussion/312631).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1730448",
      "postDate": "03/21/2022 09:14:34",
      "content": "<p>Very simple but useful answer！Good to have publications after curiosity and exploration.</p>",
      "rawMarkdown": "Very simple but useful answer！Good to have publications after curiosity and exploration.",
      "votes": null
    },
    {
      "id": "1731117",
      "postDate": "03/22/2022 02:11:09",
      "content": "<blockquote>\n  <p>Very simple but useful answer！Good to have publications after curiosity and exploration.</p>\n</blockquote>\n<p>Thankyou!</p>",
      "rawMarkdown": "> Very simple but useful answer！Good to have publications after curiosity and exploration.\n\nThankyou!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1729565,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "03/20/2022 08:04:28",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/l0new0lf\" target=\"_blank\">@l0new0lf</a> </p>\n<p>You state </p>\n<blockquote>\n  <p>\"<em>Intutively, CNNs must be able to add numbers as it is an AND operation between multiple visual-features. This notebook is proof of concept for the same.</em>\"</p>\n</blockquote>\n<p>Out of curiosity, if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4?</p>\n<p>BTW: </p>\n<blockquote>\n  <p>\"<em>I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.</em>\"</p>\n</blockquote>\n<p>Perhaps see the notebook <a href=\"https://www.kaggle.com/code/carlmcbrideellis/ultramnist-baseline-all-class-20\" target=\"_blank\">\"UltraMNIST baseline: All class 20\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": [
        {
          "id": 1729576,
          "author_name": "l0new0lf",
          "author_url": "",
          "post_date": "03/20/2022 08:34:52",
          "content": "<p>Interesting question. This is a notebook I drafted yesterday night without much thought - so have to think more before I give a concrete explanation.</p>\n<p>But let me briefly explain my intuition - in early days, humans used to represent number <code>2</code> with eyes 👀 or wings because they represented <em>\"features\"</em> called <em>\"pairs\"</em>. But, they did not have much trouble with adding numbers. So, our network must learn to internally create a feature-vs-number table. for example eye features 👀 will map to 2,  Earth features 🌎 will map to 1 and so on. </p>\n<p>Here, in-order to predict <code>3</code> model must learn that features of 🌎 <code>AND</code>  👀  are present in the image (just like multi-label problem) - this addition operation might be learnt by final classification linear units by operating on feature maps.</p>\n<p>Now, just replace 👀 and 🌎 with weird (yet coherent and visually distinct) symbols call digits - with analogy, you may understand my intuition.</p>\n<p>Regards, <br>\nAR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729583,
          "author_name": "l0new0lf",
          "author_url": "",
          "post_date": "03/20/2022 08:44:19",
          "content": "<p>And to answer the other half of the question</p>\n<blockquote>\n  <p>if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4? </p>\n</blockquote>\n<p>Firstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, <code>1, 1, 1</code> or <code>1, 1, 1, 1</code> is not possible. (please have a look at data distribution graphs). But you can try for distinct combinations like <code>3, 4, 5</code> or <code>9, 2, 5</code> with at most 3 digits</p>\n<p>Saying that, let me answer the better part of the question - </p>\n<blockquote>\n  <p>… (rather than simply memorize results)…</p>\n</blockquote>\n<p>During data generation process, for every image I am picking digits randomly and co-ordinates to position them randomly. So, to answer if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.</p>\n<p>Thanks,<br>\nAR</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729605,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "03/20/2022 09:21:35",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/l0new0lf\" target=\"_blank\">@l0new0lf</a> </p>\n<blockquote>\n  <p>\"<em>if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.</em>\"</p>\n</blockquote>\n<p>Given that you have generated all of the data yourself, and you know exactly which digits are in each image, would it not be a simple matter to create a test dataset composed of images of sets of digits that you know are not in the training set?</p>\n<p>As for </p>\n<blockquote>\n  <p>\"<em>Firstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, 1, 1, 1 or 1, 1, 1, 1 is not possible.</em>\"</p>\n</blockquote>\n<p>That is exactly why I asked. It would be an easy to perform test to see if your network can actually perform the operation of addition, or has simply encoded a dictionary of maps.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729619,
          "author_name": "l0new0lf",
          "author_url": "",
          "post_date": "03/20/2022 09:44:11",
          "content": "<p>Questions like these can better be answered by visualising predictions. Here is a very good example from my model. Because 6 looks like 0, my model predicts 8+9+0=17 instead of 8+9+6=23.</p>\n<p>Link: <a href=\"https://twitter.com/inf800/status/1505476873484087297?t=PRUjmnBIKUElGaKemCnYUw&amp;s=19\" target=\"_blank\">https://twitter.com/inf800/status/1505476873484087297?t=PRUjmnBIKUElGaKemCnYUw&amp;s=19</a></p>\n<p>(For some reason cannot upload images)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729622,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "03/20/2022 09:49:10",
          "content": "<p>and were there examples of 8+9+0=17 in the training data?….</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729627,
          "author_name": "l0new0lf",
          "author_url": "",
          "post_date": "03/20/2022 09:54:25",
          "content": "<p>This sample is from 60k samples validation set used for training. Notebook has more details about data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729647,
          "author_name": "l0new0lf",
          "author_url": "",
          "post_date": "03/20/2022 10:18:51",
          "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> Just curious. So, do you think CNNs cannot add numbers? If so, why? It will be of great help for me to know other person's perspective.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729663,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "03/20/2022 10:28:18",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/l0new0lf\" target=\"_blank\">@l0new0lf</a> </p>\n<p>You may be interested in the two publications mentioned in the topic <a href=\"https://www.kaggle.com/c/ultra-mnist/discussion/312631\" target=\"_blank\">\"The Neural Addition Unit (NAU)\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730448,
          "author_name": "alanhabrony",
          "author_url": "",
          "post_date": "03/21/2022 09:14:34",
          "content": "<p>Very simple but useful answer！Good to have publications after curiosity and exploration.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1731117,
              "author_name": "l0new0lf",
              "author_url": "",
              "post_date": "03/22/2022 02:11:09",
              "content": "<blockquote>\n  <p>Very simple but useful answer！Good to have publications after curiosity and exploration.</p>\n</blockquote>\n<p>Thankyou!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1729596,
      "author_name": "l0new0lf",
      "author_url": "",
      "post_date": "03/20/2022 09:12:21",
      "content": "<p><a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>  I'm confident that the test set data is uniformly distributed because they are using simple accuracy formula for evaluation ; D</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1729541": "![preds](https://i.imgur.com/R0SdZzg.png)\n\n## Can CNNs add numbers without any background in mathematics?\n\n### Take a step back and ask a fundamental question - is this a **multi-class** classification task?\n\n> - In order, to build an elegant pipeline (atleast for ultra-mnist problem), we should be able to answer fundamental question like - \"Can CNNs add numbers without any background in mathematics?\" or \"Is this even a multi-class problem at all?\"\n> - Intutively, CNNs must be able to add numbers as it is an `AND` operation between multiple *visual-features*. This notebook is proof of concept for the same.\n> - If CNNs cannot solve, what method would slove the problem? This is a starter notebook to answer the question.\n\nThe main issue with current data (in **innovation track**) is 1. irregular sizes of digits and 2. high resolution image.\n\nSo, here I provide custom data with both the above problems removed and simplified by huge extent. If we can achieve high accuracy here, we will work our way backwards later. My new dataset properties are:\n\n    - Uniform sized digits (28x28 pixels)\n    - Low resolution image (64x64 pixels)\n    - Only 3 possible digit combination\n    - Every digit appears only once (i.e summax = 7+8+9 = 24)\n\n\n**Note1:** I chose imbalanced becuase - who knows - organizers might have put imbalanced dataset in test folder. Even so, it will be a very good idea to train with balanced dataset. But beware - real world data is never balanced (another reason for moving ahead with imbalanced instead of balanced distribution in training set)\n\n> I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.\n\n**Note2:** This poblem can **easily solved by multi-label classification** (which is same as extracting bbox and running pretrained model). But we are not doing that. Sticking to **multi-class** problem.\n\nDataset preview\n\n![dataset-preview](https://i.imgur.com/wJG0rvz.jpg)\n\n> - test set is randomly generated just like training set\n> - only difference between them is sample size\n\n\n**CODE:** https://www.kaggle.com/code/l0new0lf/umnistv2-best-case-scenario-imbalanced",
    "1729565": "Dear @l0new0lf \n\nYou state \n\n> \"*Intutively, CNNs must be able to add numbers as it is an AND operation between multiple visual-features. This notebook is proof of concept for the same.*\"\n\nOut of curiosity, if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4?\n\nBTW: \n\n> \"*I am little confident (w/o submitting any random baseline yet) that organizers did not put imbalanced data in test set. Anyway, lets proceed.*\"\n\nPerhaps see the notebook [\"UltraMNIST baseline: All class 20\"](https://www.kaggle.com/code/carlmcbrideellis/ultramnist-baseline-all-class-20).\n\nAll the best,\ncarl",
    "1729576": "Interesting question. This is a notebook I drafted yesterday night without much thought - so have to think more before I give a concrete explanation.\n\nBut let me briefly explain my intuition - in early days, humans used to represent number `2` with eyes 👀 or wings because they represented *\"features\"* called *\"pairs\"*. But, they did not have much trouble with adding numbers. So, our network must learn to internally create a feature-vs-number table. for example eye features 👀 will map to 2,  Earth features 🌎 will map to 1 and so on. \n\nHere, in-order to predict `3` model must learn that features of 🌎 `AND`  👀  are present in the image (just like multi-label problem) - this addition operation might be learnt by final classification linear units by operating on feature maps.\n\nNow, just replace 👀 and 🌎 with weird (yet coherent and visually distinct) symbols call digits - with analogy, you may understand my intuition.\n\nRegards, \nAR",
    "1729583": "And to answer the other half of the question\n\n> if your CNN has truly learnt to add numbers (rather than simply memorize results), when given images of only the digits 1 and 1 does it return that the sum=2? And when given 1, 1, 1, 1 does it return sum=4? \n\nFirstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, `1, 1, 1` or `1, 1, 1, 1` is not possible. (please have a look at data distribution graphs). But you can try for distinct combinations like `3, 4, 5` or `9, 2, 5` with at most 3 digits\n\nSaying that, let me answer the better part of the question - \n> ... (rather than simply memorize results)...\n\nDuring data generation process, for every image I am picking digits randomly and co-ordinates to position them randomly. So, to answer if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.\n\nThanks,\nAR",
    "1729596": "carlmcbrideellis  I'm confident that the test set data is uniformly distributed because they are using simple accuracy formula for evaluation ; D",
    "1729605": "Dear @l0new0lf \n\n> \"*if there is any data leakage - we need to answer what are the chances for a same digit combination to be placed as same places.*\"\n\nGiven that you have generated all of the data yourself, and you know exactly which digits are in each image, would it not be a simple matter to create a test dataset composed of images of sets of digits that you know are not in the training set?\n\nAs for \n\n> \"*Firstly, i added a constraint that in any image one digit will appear atmost once and there will be only 3 digits. So, 1, 1, 1 or 1, 1, 1, 1 is not possible.*\"\n\nThat is exactly why I asked. It would be an easy to perform test to see if your network can actually perform the operation of addition, or has simply encoded a dictionary of maps.\n\nAll the best,\ncarl",
    "1729619": "Questions like these can better be answered by visualising predictions. Here is a very good example from my model. Because 6 looks like 0, my model predicts 8+9+0=17 instead of 8+9+6=23.\n\nLink: https://twitter.com/inf800/status/1505476873484087297?t=PRUjmnBIKUElGaKemCnYUw&s=19\n\n(For some reason cannot upload images)",
    "1729622": "and were there examples of 8+9+0=17 in the training data?....",
    "1729627": "This sample is from 60k samples validation set used for training. Notebook has more details about data.",
    "1729647": "carlmcbrideellis Just curious. So, do you think CNNs cannot add numbers? If so, why? It will be of great help for me to know other person's perspective.",
    "1729663": "Dear @l0new0lf \n\nYou may be interested in the two publications mentioned in the topic [\"The Neural Addition Unit (NAU)\"](https://www.kaggle.com/c/ultra-mnist/discussion/312631).\n\nAll the best,\ncarl",
    "1730448": "Very simple but useful answer！Good to have publications after curiosity and exploration.",
    "1731117": "> Very simple but useful answer！Good to have publications after curiosity and exploration.\n\nThankyou!"
  },
  "source": "meta"
}