{
  "id": 75365,
  "title": "vectorising images",
  "url": "/competitions/humpback-whale-identification/discussion/75365",
  "author_name": "",
  "post_date": "2018-12-20T21:57:16.458190500Z",
  "votes": -1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi! this is my first challenge.\nI would like to apply vectorisation to images,  but RAM get filled and Kernel dies.</p>\n\n<p>I am trying to:</p>\n\n<p>import imageio</p>\n\n<pre><code>def image_to_array(filename):\n    path = os.path.join(\"../input\", 'train', filename)\n    return imageio.imread( path )\n\nwhale_db = pd.read_csv(\"../input/train.csv\")\n\n# does not fit in CPU RAM\nwhale_db['ImageVector'] = whale_db['Image'].apply(image_to_array)\nwhale_db.describe()\n</code></pre>\n\n<p>Do you have advice that is possible to share on how to manipulate images data (not ML model)?\nAm I applying wrongly?</p>\n\n<p>By comparison, playing around with MNIST locally work.</p>\n\n<p>Should I work locally and only upload final results?</p>\n\n<p>thank you.</p>",
  "messages": [
    {
      "id": "443010",
      "postDate": "12/20/2018 21:57:16",
      "content": "<p>Hi! this is my first challenge.\nI would like to apply vectorisation to images,  but RAM get filled and Kernel dies.</p>\n\n<p>I am trying to:</p>\n\n<p>import imageio</p>\n\n<pre><code>def image_to_array(filename):\n    path = os.path.join(\"../input\", 'train', filename)\n    return imageio.imread( path )\n\nwhale_db = pd.read_csv(\"../input/train.csv\")\n\n# does not fit in CPU RAM\nwhale_db['ImageVector'] = whale_db['Image'].apply(image_to_array)\nwhale_db.describe()\n</code></pre>\n\n<p>Do you have advice that is possible to share on how to manipulate images data (not ML model)?\nAm I applying wrongly?</p>\n\n<p>By comparison, playing around with MNIST locally work.</p>\n\n<p>Should I work locally and only upload final results?</p>\n\n<p>thank you.</p>",
      "rawMarkdown": "Hi! this is my first challenge.\nI would like to apply vectorisation to images,  but RAM get filled and Kernel dies.\n\nI am trying to:\n\nimport imageio\n\n    def image_to_array(filename):\n        path = os.path.join(\"../input\", 'train', filename)\n        return imageio.imread( path )\n\n    whale_db = pd.read_csv(\"../input/train.csv\")\n\n    # does not fit in CPU RAM\n    whale_db['ImageVector'] = whale_db['Image'].apply(image_to_array)\n    whale_db.describe()\n\nDo you have advice that is possible to share on how to manipulate images data (not ML model)?\nAm I applying wrongly?\n\nBy comparison, playing around with MNIST locally work.\n\nShould I work locally and only upload final results?\n\nthank you.",
      "votes": null
    },
    {
      "id": "443082",
      "postDate": "12/21/2018 01:54:19",
      "content": "<p>The MNIST data set is too small so you can read it all at once in memory, whereas this data set is too large and you have to read it batch by batch.</p>",
      "rawMarkdown": "The MNIST data set is too small so you can read it all at once in memory, whereas this data set is too large and you have to read it batch by batch.",
      "votes": null
    },
    {
      "id": "444419",
      "postDate": "12/24/2018 02:29:34",
      "content": "<p>Not sure how much this will benefit you, but if you (or anyone else reading this) is using Keras, there is a <a href=\"https://keras.io/preprocessing/image/\"><code>flow_from_directory</code></a> feature which allows you to get images (and other input data) as you need it directly from the file. </p>",
      "rawMarkdown": "Not sure how much this will benefit you, but if you (or anyone else reading this) is using Keras, there is a [```flow_from_directory```](https://keras.io/preprocessing/image/) feature which allows you to get images (and other input data) as you need it directly from the file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 443082,
      "author_name": "luyang2018",
      "author_url": "",
      "post_date": "12/21/2018 01:54:19",
      "content": "<p>The MNIST data set is too small so you can read it all at once in memory, whereas this data set is too large and you have to read it batch by batch.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 444419,
      "author_name": "rajatrasal",
      "author_url": "",
      "post_date": "12/24/2018 02:29:34",
      "content": "<p>Not sure how much this will benefit you, but if you (or anyone else reading this) is using Keras, there is a <a href=\"https://keras.io/preprocessing/image/\"><code>flow_from_directory</code></a> feature which allows you to get images (and other input data) as you need it directly from the file. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "443010": "Hi! this is my first challenge.\nI would like to apply vectorisation to images,  but RAM get filled and Kernel dies.\n\nI am trying to:\n\nimport imageio\n\n    def image_to_array(filename):\n        path = os.path.join(\"../input\", 'train', filename)\n        return imageio.imread( path )\n\n    whale_db = pd.read_csv(\"../input/train.csv\")\n\n    # does not fit in CPU RAM\n    whale_db['ImageVector'] = whale_db['Image'].apply(image_to_array)\n    whale_db.describe()\n\nDo you have advice that is possible to share on how to manipulate images data (not ML model)?\nAm I applying wrongly?\n\nBy comparison, playing around with MNIST locally work.\n\nShould I work locally and only upload final results?\n\nthank you.",
    "443082": "The MNIST data set is too small so you can read it all at once in memory, whereas this data set is too large and you have to read it batch by batch.",
    "444419": "Not sure how much this will benefit you, but if you (or anyone else reading this) is using Keras, there is a [```flow_from_directory```](https://keras.io/preprocessing/image/) feature which allows you to get images (and other input data) as you need it directly from the file."
  },
  "source": "meta"
}