{
  "id": 125185,
  "title": "Ideas for Saving Memory",
  "url": "/competitions/bengaliai-cv19/discussion/125185",
  "author_name": "",
  "post_date": "2020-01-09T05:03:53.975994200Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have some ideas for saving memory, however I am not really sure if this will actually work or implemented correctly.</p>\n\n<ol>\n<li><strong>Binary Thresholding</strong>\nSince most images are white over black then this should most likely work. However, I am not sure if this works for the whole set. I checked with around 100 images (from one parquet) and found it is could maybe pass although some image have quite grainy artifacts (not sure if for this image gonna is fine or not). Also is it better to use manual thresholding (setting our own arbitrary threshold) or algorithmic thresholding?</li>\n<li><strong>Cropping Images</strong>\nI found that most images are not centered and are space inefficient (most spaces are empty). So some form of autocropping should be useful, however I have no idea how I should approach this.</li>\n</ol>\n\n<p>Anyone willing to share their tricks?</p>",
  "messages": [
    {
      "id": "714127",
      "postDate": "01/09/2020 05:03:53",
      "content": "<p>I have some ideas for saving memory, however I am not really sure if this will actually work or implemented correctly.</p>\n\n<ol>\n<li><strong>Binary Thresholding</strong>\nSince most images are white over black then this should most likely work. However, I am not sure if this works for the whole set. I checked with around 100 images (from one parquet) and found it is could maybe pass although some image have quite grainy artifacts (not sure if for this image gonna is fine or not). Also is it better to use manual thresholding (setting our own arbitrary threshold) or algorithmic thresholding?</li>\n<li><strong>Cropping Images</strong>\nI found that most images are not centered and are space inefficient (most spaces are empty). So some form of autocropping should be useful, however I have no idea how I should approach this.</li>\n</ol>\n\n<p>Anyone willing to share their tricks?</p>",
      "rawMarkdown": "I have some ideas for saving memory, however I am not really sure if this will actually work or implemented correctly.\n\n1. **Binary Thresholding**\n    Since most images are white over black then this should most likely work. However, I am not sure if this works for the whole set. I checked with around 100 images (from one parquet) and found it is could maybe pass although some image have quite grainy artifacts (not sure if for this image gonna is fine or not). Also is it better to use manual thresholding (setting our own arbitrary threshold) or algorithmic thresholding?\n2.  **Cropping Images**\nI found that most images are not centered and are space inefficient (most spaces are empty). So some form of autocropping should be useful, however I have no idea how I should approach this.\n\nAnyone willing to share their tricks?",
      "votes": null
    },
    {
      "id": "714440",
      "postDate": "01/09/2020 12:51:17",
      "content": "<p>I made a notebook with that second idea that you mention, if ever you want <a href=\"https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images\">to take a look</a></p>\n\n<p>As for the first one, I'm guessing it could be useful but only for certain specific characters (like if one in particular is bigger than the others). But I also wouldn't be surprised if a model picks that up on its own already. </p>",
      "rawMarkdown": "I made a notebook with that second idea that you mention, if ever you want [to take a look](https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images)\n\nAs for the first one, I'm guessing it could be useful but only for certain specific characters (like if one in particular is bigger than the others). But I also wouldn't be surprised if a model picks that up on its own already.",
      "votes": null
    },
    {
      "id": "715670",
      "postDate": "01/10/2020 18:00:07",
      "content": "<p>first do manual delete of large objects and gc.collect \nsecond change data type to int8 and float32</p>",
      "rawMarkdown": "first do manual delete of large objects and gc.collect \nsecond change data type to int8 and float32",
      "votes": null
    },
    {
      "id": "715981",
      "postDate": "01/11/2020 04:54:19",
      "content": "<p>We can do GC.collect() and all that still if you get the problem you can try doing transfer learning one paraquet at a time!!</p>",
      "rawMarkdown": "We can do GC.collect() and all that still if you get the problem you can try doing transfer learning one paraquet at a time!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 714440,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "01/09/2020 12:51:17",
      "content": "<p>I made a notebook with that second idea that you mention, if ever you want <a href=\"https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images\">to take a look</a></p>\n\n<p>As for the first one, I'm guessing it could be useful but only for certain specific characters (like if one in particular is bigger than the others). But I also wouldn't be surprised if a model picks that up on its own already. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 715670,
      "author_name": "timaskr",
      "author_url": "",
      "post_date": "01/10/2020 18:00:07",
      "content": "<p>first do manual delete of large objects and gc.collect \nsecond change data type to int8 and float32</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 715981,
      "author_name": "chekoduadarsh",
      "author_url": "",
      "post_date": "01/11/2020 04:54:19",
      "content": "<p>We can do GC.collect() and all that still if you get the problem you can try doing transfer learning one paraquet at a time!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "714127": "I have some ideas for saving memory, however I am not really sure if this will actually work or implemented correctly.\n\n1. **Binary Thresholding**\n    Since most images are white over black then this should most likely work. However, I am not sure if this works for the whole set. I checked with around 100 images (from one parquet) and found it is could maybe pass although some image have quite grainy artifacts (not sure if for this image gonna is fine or not). Also is it better to use manual thresholding (setting our own arbitrary threshold) or algorithmic thresholding?\n2.  **Cropping Images**\nI found that most images are not centered and are space inefficient (most spaces are empty). So some form of autocropping should be useful, however I have no idea how I should approach this.\n\nAnyone willing to share their tricks?",
    "714440": "I made a notebook with that second idea that you mention, if ever you want [to take a look](https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images)\n\nAs for the first one, I'm guessing it could be useful but only for certain specific characters (like if one in particular is bigger than the others). But I also wouldn't be surprised if a model picks that up on its own already.",
    "715670": "first do manual delete of large objects and gc.collect \nsecond change data type to int8 and float32",
    "715981": "We can do GC.collect() and all that still if you get the problem you can try doing transfer learning one paraquet at a time!!"
  },
  "source": "meta"
}