{
  "id": 312543,
  "title": "Abstaining From Further Participation - Sharing My Ideas",
  "url": "/competitions/ultra-mnist/discussion/312543",
  "author_name": "",
  "post_date": "2022-03-12T16:46:20.712194500Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi there, I will be abstaining from further participation in this competition following the rule changes/clarifications and I have made all my notebooks/datasets public.</p>\n<hr>\n<p>The specific rule clarification is what is considered disallowed public, open-source data.</p>\n<ul>\n<li>Previously the understanding was that using open-source datasets like MNIST was not allowed but using open source models (ImageNET, MNIST, COCO, etc) from reputable, historical sources (timm, tf, etc.) would be accepted.</li>\n<li>Now the understanding is that we are only allowed to use open source models trained on ImageNET. </li>\n</ul>\n<hr>\n<p>I would like to share some ideas that I think could be useful in developing novel solutions. Some of these ideas may be absurd, some may not work, but I just wanted to share my brainstorm a bit before I bowed out. These ideas will pertain strictly to approaches that only leverage ImageNET trained models and do not remove the background squares.</p>\n<hr>\n<p><strong>IDEA 1</strong> - Using RLE Encoding On Binarized Image <a href=\"https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338\" target=\"_blank\">[EXAMPLE NOTEBOOK]</a></p>\n<ul>\n<li>Say we resize the images to 1000x1000</li>\n<li>We then perform RLE encoding </li>\n<li>The result is that each image is converted into a sequence of 2882 'tokens'.</li>\n<li>Looking at 5 examples it would appear that our 'vocabulary' of 'tokens' is composed of '1501' unique tokens.</li>\n<li>As the RLE encoding generates digits we could also endeavour to incorporate the digits themselves (not just the idea of them as strings) into the algorithm somehow.</li>\n<li>From here we could use a whole bunch of different models<ul>\n<li>Sequence models on the sequence of tokens</li>\n<li>XGB/LGB models using each token location as a column and each token as the value.</li>\n<li>etc.</li></ul></li>\n</ul>\n<p><strong>EXTERNAL DATA DEPENDENCIES</strong>:</p>\n<ul>\n<li>None</li>\n</ul>\n<hr>\n<p><strong>IDEA 2</strong> - Two-Stage Modelling w/ Pseudo Labels</p>\n<ul>\n<li>Train a (FAST - LOW PARAM - XGB maybe? RLE?) image model to perform inference on large images (1000x1000 maybe?) in an attempt to simply go end-to-end (1000x1000 --&gt; 28 classes).</li>\n<li>Take this model and use explainability techniques to generate candidate ROIs (think 2 stage object detection). Essentially we don't care about whether the pred was right or wrong… simply what ROIs it was using to try and come up with a guess.</li>\n<li>Crop (pad generously) the generated candidate regions, taking up to N crops (let's say 10… even though we know the max is 5)</li>\n<li>Frame this now as a new problem where N input images (think sequence or something) yields the desired sum. </li>\n<li>Leverage multi-input Convolutional Neural Network, Recurrent Neural Network, Transformer, XGB, whatever.</li>\n</ul>\n<p><strong>EXTERNAL DATA DEPENDENCIES</strong>:</p>\n<ul>\n<li>You could use an ImageNET model for stage one… maybe something really fast like a MobileNet with a really small alpha (0.1)?</li>\n<li>Other than that None</li>\n</ul>\n<hr>\n<p><strong>IDEA 3</strong> - Figure Out How To Identify/Learn Based On Relationships Between Digits</p>\n<ul>\n<li>I don't know a lot about this one but maybe…<ul>\n<li>GNNs</li>\n<li>BarCNN</li>\n<li>etc.</li></ul></li>\n</ul>\n<hr>\n<p><strong>IDEA 4</strong> - Kind of Cheating with OD (I think this goes more against the spirit of the comp than the other ideas)</p>\n<ul>\n<li>Use Object Detection Pretrained Mode<ul>\n<li>Take logit output from object detection model for each detected object</li>\n<li>Take these logits as input for some other NN</li>\n<li>OR - Simply manually observe which pretrained labels correspond the most to which numbers and create a mapping to reassign binned pretrained predictions. </li></ul></li>\n</ul>\n<hr>\n<p><br></p>\n<p>I only had a couple of hours to come up with these ideas… but I wanted to share them anyway. I hope this helps some of you. If not, it was still fun to brainstorm some solutions. I look forward to seeing what the top solutions will be and good luck!!</p>",
  "messages": [
    {
      "id": "1720287",
      "postDate": "03/12/2022 16:46:20",
      "content": "<p>Hi there, I will be abstaining from further participation in this competition following the rule changes/clarifications and I have made all my notebooks/datasets public.</p>\n<hr>\n<p>The specific rule clarification is what is considered disallowed public, open-source data.</p>\n<ul>\n<li>Previously the understanding was that using open-source datasets like MNIST was not allowed but using open source models (ImageNET, MNIST, COCO, etc) from reputable, historical sources (timm, tf, etc.) would be accepted.</li>\n<li>Now the understanding is that we are only allowed to use open source models trained on ImageNET. </li>\n</ul>\n<hr>\n<p>I would like to share some ideas that I think could be useful in developing novel solutions. Some of these ideas may be absurd, some may not work, but I just wanted to share my brainstorm a bit before I bowed out. These ideas will pertain strictly to approaches that only leverage ImageNET trained models and do not remove the background squares.</p>\n<hr>\n<p><strong>IDEA 1</strong> - Using RLE Encoding On Binarized Image <a href=\"https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338\" target=\"_blank\">[EXAMPLE NOTEBOOK]</a></p>\n<ul>\n<li>Say we resize the images to 1000x1000</li>\n<li>We then perform RLE encoding </li>\n<li>The result is that each image is converted into a sequence of 2882 'tokens'.</li>\n<li>Looking at 5 examples it would appear that our 'vocabulary' of 'tokens' is composed of '1501' unique tokens.</li>\n<li>As the RLE encoding generates digits we could also endeavour to incorporate the digits themselves (not just the idea of them as strings) into the algorithm somehow.</li>\n<li>From here we could use a whole bunch of different models<ul>\n<li>Sequence models on the sequence of tokens</li>\n<li>XGB/LGB models using each token location as a column and each token as the value.</li>\n<li>etc.</li></ul></li>\n</ul>\n<p><strong>EXTERNAL DATA DEPENDENCIES</strong>:</p>\n<ul>\n<li>None</li>\n</ul>\n<hr>\n<p><strong>IDEA 2</strong> - Two-Stage Modelling w/ Pseudo Labels</p>\n<ul>\n<li>Train a (FAST - LOW PARAM - XGB maybe? RLE?) image model to perform inference on large images (1000x1000 maybe?) in an attempt to simply go end-to-end (1000x1000 --&gt; 28 classes).</li>\n<li>Take this model and use explainability techniques to generate candidate ROIs (think 2 stage object detection). Essentially we don't care about whether the pred was right or wrong… simply what ROIs it was using to try and come up with a guess.</li>\n<li>Crop (pad generously) the generated candidate regions, taking up to N crops (let's say 10… even though we know the max is 5)</li>\n<li>Frame this now as a new problem where N input images (think sequence or something) yields the desired sum. </li>\n<li>Leverage multi-input Convolutional Neural Network, Recurrent Neural Network, Transformer, XGB, whatever.</li>\n</ul>\n<p><strong>EXTERNAL DATA DEPENDENCIES</strong>:</p>\n<ul>\n<li>You could use an ImageNET model for stage one… maybe something really fast like a MobileNet with a really small alpha (0.1)?</li>\n<li>Other than that None</li>\n</ul>\n<hr>\n<p><strong>IDEA 3</strong> - Figure Out How To Identify/Learn Based On Relationships Between Digits</p>\n<ul>\n<li>I don't know a lot about this one but maybe…<ul>\n<li>GNNs</li>\n<li>BarCNN</li>\n<li>etc.</li></ul></li>\n</ul>\n<hr>\n<p><strong>IDEA 4</strong> - Kind of Cheating with OD (I think this goes more against the spirit of the comp than the other ideas)</p>\n<ul>\n<li>Use Object Detection Pretrained Mode<ul>\n<li>Take logit output from object detection model for each detected object</li>\n<li>Take these logits as input for some other NN</li>\n<li>OR - Simply manually observe which pretrained labels correspond the most to which numbers and create a mapping to reassign binned pretrained predictions. </li></ul></li>\n</ul>\n<hr>\n<p><br></p>\n<p>I only had a couple of hours to come up with these ideas… but I wanted to share them anyway. I hope this helps some of you. If not, it was still fun to brainstorm some solutions. I look forward to seeing what the top solutions will be and good luck!!</p>",
      "rawMarkdown": "Hi there, I will be abstaining from further participation in this competition following the rule changes/clarifications and I have made all my notebooks/datasets public.\n\n---\n\nThe specific rule clarification is what is considered disallowed public, open-source data.\n* Previously the understanding was that using open-source datasets like MNIST was not allowed but using open source models (ImageNET, MNIST, COCO, etc) from reputable, historical sources (timm, tf, etc.) would be accepted.\n* Now the understanding is that we are only allowed to use open source models trained on ImageNET. \n\n---\n\nI would like to share some ideas that I think could be useful in developing novel solutions. Some of these ideas may be absurd, some may not work, but I just wanted to share my brainstorm a bit before I bowed out. These ideas will pertain strictly to approaches that only leverage ImageNET trained models and do not remove the background squares.\n\n---\n\n**IDEA 1** - Using RLE Encoding On Binarized Image [[EXAMPLE NOTEBOOK]](https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338)\n* Say we resize the images to 1000x1000\n* We then perform RLE encoding \n* The result is that each image is converted into a sequence of 2882 'tokens'.\n* Looking at 5 examples it would appear that our 'vocabulary' of 'tokens' is composed of '1501' unique tokens.\n* As the RLE encoding generates digits we could also endeavour to incorporate the digits themselves (not just the idea of them as strings) into the algorithm somehow.\n* From here we could use a whole bunch of different models\n  * Sequence models on the sequence of tokens\n  * XGB/LGB models using each token location as a column and each token as the value.\n  * etc.\n\n**EXTERNAL DATA DEPENDENCIES**:\n* None\n\n---\n\n**IDEA 2** - Two-Stage Modelling w/ Pseudo Labels\n* Train a (FAST - LOW PARAM - XGB maybe? RLE?) image model to perform inference on large images (1000x1000 maybe?) in an attempt to simply go end-to-end (1000x1000 --> 28 classes).\n* Take this model and use explainability techniques to generate candidate ROIs (think 2 stage object detection). Essentially we don't care about whether the pred was right or wrong... simply what ROIs it was using to try and come up with a guess.\n* Crop (pad generously) the generated candidate regions, taking up to N crops (let's say 10... even though we know the max is 5)\n* Frame this now as a new problem where N input images (think sequence or something) yields the desired sum. \n* Leverage multi-input Convolutional Neural Network, Recurrent Neural Network, Transformer, XGB, whatever.\n\n**EXTERNAL DATA DEPENDENCIES**:\n* You could use an ImageNET model for stage one... maybe something really fast like a MobileNet with a really small alpha (0.1)?\n* Other than that None\n\n---\n\n**IDEA 3** - Figure Out How To Identify/Learn Based On Relationships Between Digits\n\n* I don't know a lot about this one but maybe...\n  * GNNs\n  * BarCNN\n  * etc.\n\n---\n\n**IDEA 4** - Kind of Cheating with OD (I think this goes more against the spirit of the comp than the other ideas)\n\n* Use Object Detection Pretrained Mode\n  * Take logit output from object detection model for each detected object\n  * Take these logits as input for some other NN\n  * OR - Simply manually observe which pretrained labels correspond the most to which numbers and create a mapping to reassign binned pretrained predictions. \n\n---\n\n<br>\n\nI only had a couple of hours to come up with these ideas... but I wanted to share them anyway. I hope this helps some of you. If not, it was still fun to brainstorm some solutions. I look forward to seeing what the top solutions will be and good luck!!",
      "votes": null
    },
    {
      "id": "1721581",
      "postDate": "03/13/2022 20:00:48",
      "content": "<p>Idea 1 is returning page 404<br>\n<a href=\"https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338\" target=\"_blank\">https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338</a></p>",
      "rawMarkdown": "Idea 1 is returning page 404\nhttps://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338",
      "votes": null
    },
    {
      "id": "1721586",
      "postDate": "03/13/2022 20:05:07",
      "content": "<p>Sorry about that, it was still private.</p>\n<p>I’ve updated it so it’s public now.</p>",
      "rawMarkdown": "Sorry about that, it was still private.\n\nI’ve updated it so it’s public now.",
      "votes": null
    },
    {
      "id": "1722492",
      "postDate": "03/14/2022 15:24:45",
      "content": "<p>you are just amazing <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> <br>\nYour comments are of quite help, especially for beginners like me! Thanks for this!</p>",
      "rawMarkdown": "you are just amazing @dschettler8845 \nYour comments are of quite help, especially for beginners like me! Thanks for this!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1721581,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "03/13/2022 20:00:48",
      "content": "<p>Idea 1 is returning page 404<br>\n<a href=\"https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338\" target=\"_blank\">https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1721586,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "03/13/2022 20:05:07",
          "content": "<p>Sorry about that, it was still private.</p>\n<p>I’ve updated it so it’s public now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1722492,
      "author_name": "imams2000",
      "author_url": "",
      "post_date": "03/14/2022 15:24:45",
      "content": "<p>you are just amazing <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> <br>\nYour comments are of quite help, especially for beginners like me! Thanks for this!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1720287": "Hi there, I will be abstaining from further participation in this competition following the rule changes/clarifications and I have made all my notebooks/datasets public.\n\n---\n\nThe specific rule clarification is what is considered disallowed public, open-source data.\n* Previously the understanding was that using open-source datasets like MNIST was not allowed but using open source models (ImageNET, MNIST, COCO, etc) from reputable, historical sources (timm, tf, etc.) would be accepted.\n* Now the understanding is that we are only allowed to use open source models trained on ImageNET. \n\n---\n\nI would like to share some ideas that I think could be useful in developing novel solutions. Some of these ideas may be absurd, some may not work, but I just wanted to share my brainstorm a bit before I bowed out. These ideas will pertain strictly to approaches that only leverage ImageNET trained models and do not remove the background squares.\n\n---\n\n**IDEA 1** - Using RLE Encoding On Binarized Image [[EXAMPLE NOTEBOOK]](https://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338)\n* Say we resize the images to 1000x1000\n* We then perform RLE encoding \n* The result is that each image is converted into a sequence of 2882 'tokens'.\n* Looking at 5 examples it would appear that our 'vocabulary' of 'tokens' is composed of '1501' unique tokens.\n* As the RLE encoding generates digits we could also endeavour to incorporate the digits themselves (not just the idea of them as strings) into the algorithm somehow.\n* From here we could use a whole bunch of different models\n  * Sequence models on the sequence of tokens\n  * XGB/LGB models using each token location as a column and each token as the value.\n  * etc.\n\n**EXTERNAL DATA DEPENDENCIES**:\n* None\n\n---\n\n**IDEA 2** - Two-Stage Modelling w/ Pseudo Labels\n* Train a (FAST - LOW PARAM - XGB maybe? RLE?) image model to perform inference on large images (1000x1000 maybe?) in an attempt to simply go end-to-end (1000x1000 --> 28 classes).\n* Take this model and use explainability techniques to generate candidate ROIs (think 2 stage object detection). Essentially we don't care about whether the pred was right or wrong... simply what ROIs it was using to try and come up with a guess.\n* Crop (pad generously) the generated candidate regions, taking up to N crops (let's say 10... even though we know the max is 5)\n* Frame this now as a new problem where N input images (think sequence or something) yields the desired sum. \n* Leverage multi-input Convolutional Neural Network, Recurrent Neural Network, Transformer, XGB, whatever.\n\n**EXTERNAL DATA DEPENDENCIES**:\n* You could use an ImageNET model for stage one... maybe something really fast like a MobileNet with a really small alpha (0.1)?\n* Other than that None\n\n---\n\n**IDEA 3** - Figure Out How To Identify/Learn Based On Relationships Between Digits\n\n* I don't know a lot about this one but maybe...\n  * GNNs\n  * BarCNN\n  * etc.\n\n---\n\n**IDEA 4** - Kind of Cheating with OD (I think this goes more against the spirit of the comp than the other ideas)\n\n* Use Object Detection Pretrained Mode\n  * Take logit output from object detection model for each detected object\n  * Take these logits as input for some other NN\n  * OR - Simply manually observe which pretrained labels correspond the most to which numbers and create a mapping to reassign binned pretrained predictions. \n\n---\n\n<br>\n\nI only had a couple of hours to come up with these ideas... but I wanted to share them anyway. I hope this helps some of you. If not, it was still fun to brainstorm some solutions. I look forward to seeing what the top solutions will be and good luck!!",
    "1721581": "Idea 1 is returning page 404\nhttps://www.kaggle.com/dschettler8845/rle-encode-idea?scriptVersionId=89928338",
    "1721586": "Sorry about that, it was still private.\n\nI’ve updated it so it’s public now.",
    "1722492": "you are just amazing @dschettler8845 \nYour comments are of quite help, especially for beginners like me! Thanks for this!"
  },
  "source": "meta"
}