{
  "id": 40827,
  "title": "Caution for EXT4 file counts. And any advice?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40827",
  "author_name": "",
  "post_date": "2017-10-09T05:51:00.915162800Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>While extracting image files from bson into separate jpg files,  I got an error that disk is full even though df shows it has enough space. It turns out that EXT4 file system has limitation on # of files in file system.</p>\n\n<p><a href=\"https://serverfault.com/questions/104986/what-is-the-maximum-number-of-files-a-file-system-can-contain\">https://serverfault.com/questions/104986/what-is-the-maximum-number-of-files-a-file-system-can-contain</a></p>\n\n<p>You can check the current inode usage with df -i</p>\n\n<p>And any advice on this issue?</p>\n\n<p>Solutions that I can think of are:</p>\n\n<ol>\n<li>Re-format partition with more inode count (hard to do since this is\nmy root partition :( )</li>\n<li>Read directly from BSON file (not an option for me since I want to\nresize the files into smaller resolution)</li>\n<li>Mount a file as a file system and save images to this. </li>\n<li>Convert BSON into different format - any suggestion?</li>\n<li>Any tool to increase inode without formatting?</li>\n</ol>",
  "messages": [
    {
      "id": "229211",
      "postDate": "10/09/2017 05:51:00",
      "content": "<p>While extracting image files from bson into separate jpg files,  I got an error that disk is full even though df shows it has enough space. It turns out that EXT4 file system has limitation on # of files in file system.</p>\n\n<p><a href=\"https://serverfault.com/questions/104986/what-is-the-maximum-number-of-files-a-file-system-can-contain\">https://serverfault.com/questions/104986/what-is-the-maximum-number-of-files-a-file-system-can-contain</a></p>\n\n<p>You can check the current inode usage with df -i</p>\n\n<p>And any advice on this issue?</p>\n\n<p>Solutions that I can think of are:</p>\n\n<ol>\n<li>Re-format partition with more inode count (hard to do since this is\nmy root partition :( )</li>\n<li>Read directly from BSON file (not an option for me since I want to\nresize the files into smaller resolution)</li>\n<li>Mount a file as a file system and save images to this. </li>\n<li>Convert BSON into different format - any suggestion?</li>\n<li>Any tool to increase inode without formatting?</li>\n</ol>",
      "rawMarkdown": "While extracting image files from bson into separate jpg files,  I got an error that disk is full even though df shows it has enough space. It turns out that EXT4 file system has limitation on # of files in file system.\n\nhttps://serverfault.com/questions/104986/what-is-the-maximum-number-of-files-a-file-system-can-contain\n\nYou can check the current inode usage with df -i\n\nAnd any advice on this issue?\n\nSolutions that I can think of are:\n\n 1. Re-format partition with more inode count (hard to do since this is\n    my root partition :( )\n 2. Read directly from BSON file (not an option for me since I want to\n    resize the files into smaller resolution)\n 3. Mount a file as a file system and save images to this. \n 4. Convert BSON into different format - any suggestion?\n 5. Any tool to increase inode without formatting?",
      "votes": null
    },
    {
      "id": "229224",
      "postDate": "10/09/2017 06:27:17",
      "content": "<p>Why on earth would you like to have images stored as separate files?</p>\n\n<p>Regarding different file format, you may convert them to TFRecords, see kernels for a sample converter I have made.</p>",
      "rawMarkdown": "Why on earth would you like to have images stored as separate files?\n\nRegarding different file format, you may convert them to TFRecords, see kernels for a sample converter I have made.",
      "votes": null
    },
    {
      "id": "229236",
      "postDate": "10/09/2017 07:07:31",
      "content": "<p>I thought it would be easier to handle despite file system overhead. </p>\n\n<p>Thanks for the kernel - I didn't see your kernel.</p>\n\n<p>About TFRecords, is it possible to random access element?</p>\n\n<p>I can find some tricks to overcome it, but just creating lookup table + one big file stream seems to be simple enough also.</p>",
      "rawMarkdown": "I thought it would be easier to handle despite file system overhead. \n\nThanks for the kernel - I didn't see your kernel.\n\nAbout TFRecords, is it possible to random access element?\n\nI can find some tricks to overcome it, but just creating lookup table + one big file stream seems to be simple enough also.",
      "votes": null
    },
    {
      "id": "229270",
      "postDate": "10/09/2017 08:56:23",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "229275",
      "postDate": "10/09/2017 09:18:38",
      "content": "<p>I have never thought about it. I have divided the base BSON file into about 200 TFRecords and just shuffle them. That is also my way of dividing sample into train and test (just by assigning files to each group, not single images).</p>",
      "rawMarkdown": "I have never thought about it. I have divided the base BSON file into about 200 TFRecords and just shuffle them. That is also my way of dividing sample into train and test (just by assigning files to each group, not single images).",
      "votes": null
    },
    {
      "id": "229441",
      "postDate": "10/09/2017 16:11:57",
      "content": "<p>It looks like distribution of product category is highly skewed along product id. \nMost frequent product in all products is just 19th frequent among the first 100000 products. So I guess it's better to split them in stratified way.</p>",
      "rawMarkdown": "It looks like distribution of product category is highly skewed along product id. \nMost frequent product in all products is just 19th frequent among the first 100000 products. So I guess it's better to split them in stratified way.",
      "votes": null
    },
    {
      "id": "229542",
      "postDate": "10/09/2017 21:05:01",
      "content": "<p>Wow, thanks for possible explanation. I had the same prob writing images from linux to ntfs (showing that disk is full even though it has plenty of free space). The next day my hdd died)) So I decided to use generator from bson directly to pytorch. Haven't tested it yet in action though.</p>",
      "rawMarkdown": "Wow, thanks for possible explanation. I had the same prob writing images from linux to ntfs (showing that disk is full even though it has plenty of free space). The next day my hdd died)) So I decided to use generator from bson directly to pytorch. Haven't tested it yet in action though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 229224,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "10/09/2017 06:27:17",
      "content": "<p>Why on earth would you like to have images stored as separate files?</p>\n\n<p>Regarding different file format, you may convert them to TFRecords, see kernels for a sample converter I have made.</p>",
      "votes": null,
      "replies": [
        {
          "id": 229236,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "10/09/2017 07:07:31",
          "content": "<p>I thought it would be easier to handle despite file system overhead. </p>\n\n<p>Thanks for the kernel - I didn't see your kernel.</p>\n\n<p>About TFRecords, is it possible to random access element?</p>\n\n<p>I can find some tricks to overcome it, but just creating lookup table + one big file stream seems to be simple enough also.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 229275,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "10/09/2017 09:18:38",
          "content": "<p>I have never thought about it. I have divided the base BSON file into about 200 TFRecords and just shuffle them. That is also my way of dividing sample into train and test (just by assigning files to each group, not single images).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 229441,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "10/09/2017 16:11:57",
          "content": "<p>It looks like distribution of product category is highly skewed along product id. \nMost frequent product in all products is just 19th frequent among the first 100000 products. So I guess it's better to split them in stratified way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 229270,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/09/2017 08:56:23",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 229542,
      "author_name": "heyt0ny",
      "author_url": "",
      "post_date": "10/09/2017 21:05:01",
      "content": "<p>Wow, thanks for possible explanation. I had the same prob writing images from linux to ntfs (showing that disk is full even though it has plenty of free space). The next day my hdd died)) So I decided to use generator from bson directly to pytorch. Haven't tested it yet in action though.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "229211": "While extracting image files from bson into separate jpg files,  I got an error that disk is full even though df shows it has enough space. It turns out that EXT4 file system has limitation on # of files in file system.\n\nhttps://serverfault.com/questions/104986/what-is-the-maximum-number-of-files-a-file-system-can-contain\n\nYou can check the current inode usage with df -i\n\nAnd any advice on this issue?\n\nSolutions that I can think of are:\n\n 1. Re-format partition with more inode count (hard to do since this is\n    my root partition :( )\n 2. Read directly from BSON file (not an option for me since I want to\n    resize the files into smaller resolution)\n 3. Mount a file as a file system and save images to this. \n 4. Convert BSON into different format - any suggestion?\n 5. Any tool to increase inode without formatting?",
    "229224": "Why on earth would you like to have images stored as separate files?\n\nRegarding different file format, you may convert them to TFRecords, see kernels for a sample converter I have made.",
    "229236": "I thought it would be easier to handle despite file system overhead. \n\nThanks for the kernel - I didn't see your kernel.\n\nAbout TFRecords, is it possible to random access element?\n\nI can find some tricks to overcome it, but just creating lookup table + one big file stream seems to be simple enough also.",
    "229270": "",
    "229275": "I have never thought about it. I have divided the base BSON file into about 200 TFRecords and just shuffle them. That is also my way of dividing sample into train and test (just by assigning files to each group, not single images).",
    "229441": "It looks like distribution of product category is highly skewed along product id. \nMost frequent product in all products is just 19th frequent among the first 100000 products. So I guess it's better to split them in stratified way.",
    "229542": "Wow, thanks for possible explanation. I had the same prob writing images from linux to ntfs (showing that disk is full even though it has plenty of free space). The next day my hdd died)) So I decided to use generator from bson directly to pytorch. Haven't tested it yet in action though."
  },
  "source": "meta"
}