{
  "id": 40715,
  "title": "Keras starter",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40715",
  "author_name": "Petros Giannakopoulos",
  "post_date": "2017-10-06T21:45:39.263000",
  "votes": 18,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello everyone! This looks like a challenging classification task due to the very large dataset and large number of classes. </p>\n\n<p>This is my code so far for this competition, I hope that it helps more people to join and that it provides a decent baseline to build a robust solution on:</p>\n\n<p><a href=\"https://github.com/petrosgk/Kaggle-Cdiscount-Image-Classification-Challenge\">https://github.com/petrosgk/Kaggle-Cdiscount-Image-Classification-Challenge</a></p>\n\n<p>I am still processing the .bson files so no results on the full dataset yet. These are 5 epochs on a subset of 10k samples with pre-trained VGG16 (with top dense layers removed) and image size 180x180:</p>\n\n<pre><code>Epoch 1/20\n231s - loss: 3.7500 - acc: 0.4056 - val_loss: 3.2170 - val_acc: 0.4783\nEpoch 2/20\n227s - loss: 2.8793 - acc: 0.5010 - val_loss: 3.0092 - val_acc: 0.5149\nEpoch 3/20\n228s - loss: 2.4361 - acc: 0.5570 - val_loss: 2.9464 - val_acc: 0.5323\nEpoch 4/20\n229s - loss: 2.1199 - acc: 0.5996 - val_loss: 2.9545 - val_acc: 0.5447\nEpoch 5/20\n229s - loss: 1.8476 - acc: 0.6346 - val_loss: 2.9870 - val_acc: 0.5440\n</code></pre>\n\n<p>There is overfitting with 9k training samples (90/10 train-val split) but that may change when training on the full ~7m samples.</p>",
  "messages": [
    {
      "id": 228504,
      "postDate": "2017-10-06T21:45:39.263Z",
      "content": "<p>Hello everyone! This looks like a challenging classification task due to the very large dataset and large number of classes. </p>\n\n<p>This is my code so far for this competition, I hope that it helps more people to join and that it provides a decent baseline to build a robust solution on:</p>\n\n<p><a href=\"https://github.com/petrosgk/Kaggle-Cdiscount-Image-Classification-Challenge\">https://github.com/petrosgk/Kaggle-Cdiscount-Image-Classification-Challenge</a></p>\n\n<p>I am still processing the .bson files so no results on the full dataset yet. These are 5 epochs on a subset of 10k samples with pre-trained VGG16 (with top dense layers removed) and image size 180x180:</p>\n\n<pre><code>Epoch 1/20\n231s - loss: 3.7500 - acc: 0.4056 - val_loss: 3.2170 - val_acc: 0.4783\nEpoch 2/20\n227s - loss: 2.8793 - acc: 0.5010 - val_loss: 3.0092 - val_acc: 0.5149\nEpoch 3/20\n228s - loss: 2.4361 - acc: 0.5570 - val_loss: 2.9464 - val_acc: 0.5323\nEpoch 4/20\n229s - loss: 2.1199 - acc: 0.5996 - val_loss: 2.9545 - val_acc: 0.5447\nEpoch 5/20\n229s - loss: 1.8476 - acc: 0.6346 - val_loss: 2.9870 - val_acc: 0.5440\n</code></pre>\n\n<p>There is overfitting with 9k training samples (90/10 train-val split) but that may change when training on the full ~7m samples.</p>",
      "rawMarkdown": "Hello everyone! This looks like a challenging classification task due to the very large dataset and large number of classes. \n\nThis is my code so far for this competition, I hope that it helps more people to join and that it provides a decent baseline to build a robust solution on:\n\nhttps://github.com/petrosgk/Kaggle-Cdiscount-Image-Classification-Challenge\n\nI am still processing the .bson files so no results on the full dataset yet. These are 5 epochs on a subset of 10k samples with pre-trained VGG16 (with top dense layers removed) and image size 180x180:\n\n    Epoch 1/20\n    231s - loss: 3.7500 - acc: 0.4056 - val_loss: 3.2170 - val_acc: 0.4783\n    Epoch 2/20\n    227s - loss: 2.8793 - acc: 0.5010 - val_loss: 3.0092 - val_acc: 0.5149\n    Epoch 3/20\n    228s - loss: 2.4361 - acc: 0.5570 - val_loss: 2.9464 - val_acc: 0.5323\n    Epoch 4/20\n    229s - loss: 2.1199 - acc: 0.5996 - val_loss: 2.9545 - val_acc: 0.5447\n    Epoch 5/20\n    229s - loss: 1.8476 - acc: 0.6346 - val_loss: 2.9870 - val_acc: 0.5440\n\nThere is overfitting with 9k training samples (90/10 train-val split) but that may change when training on the full ~7m samples.",
      "votes": 18
    },
    {
      "id": 231530,
      "postDate": "2017-10-15T09:07:16.680Z",
      "content": "<p>I suffered from training very slow when using mxnet. I found that if random read raw images, it is very slow. Now I use the rec file (it is a file generated from all images, hence is fast) of mxnet, it improves from 140 imgs/second to 600 imgs/second.</p>",
      "rawMarkdown": "I suffered from training very slow when using mxnet. I found that if random read raw images, it is very slow. Now I use the rec file (it is a file generated from all images, hence is fast) of mxnet, it improves from 140 imgs/second to 600 imgs/second.",
      "votes": 1
    },
    {
      "id": 233722,
      "postDate": "2017-10-20T23:07:54.557Z",
      "content": "<p>I took at look at your Keras starter - a few good learning points, I did a trip to Keras.io a few times, I like the 'callbacks'. \nA few observations and a question: The png files generated for the train folder are in the hundreds of gigabytes (caveat emptor) so it goes without saying -training and manipulation- can take a very very long time to those stepping into this. Don't waste time using a standard HDD, nothing less than an m.2 SSD will do. Just doing a simple ls -la can take 5 minutes. My SSD which is much quicker unfortunately is too small to hold that much data! None-the-less what variables did you tinker with to run just a 10k sample using train.py? Usually I can spot this quickly but things look a little nested. As a base line I want to replicate your result as shown. Thus far I cannot. It would help square things up.</p>",
      "rawMarkdown": "I took at look at your Keras starter - a few good learning points, I did a trip to Keras.io a few times, I like the 'callbacks'. \nA few observations and a question: The png files generated for the train folder are in the hundreds of gigabytes (caveat emptor) so it goes without saying -training and manipulation- can take a very very long time to those stepping into this. Don't waste time using a standard HDD, nothing less than an m.2 SSD will do. Just doing a simple ls -la can take 5 minutes. My SSD which is much quicker unfortunately is too small to hold that much data! None-the-less what variables did you tinker with to run just a 10k sample using train.py? Usually I can spot this quickly but things look a little nested. As a base line I want to replicate your result as shown. Thus far I cannot. It would help square things up.",
      "replies": [
        {
          "id": 233796,
          "postDate": "2017-10-21T09:17:18.270Z",
          "content": "<p>In the code at the beginning of train.py that reads the product_id's, category_id's and num of images per product_id:</p>\n\n<pre><code>df = pd.read_csv(csv_file)\nproduct_ids, category_ids, n_pics = df['product_id'], df['category_id'], df['n_pics']\n</code></pre>\n\n<p>You can simply only use the first 10k elements of the product_ids, category_ids, n_pics arrays. So you can add something like this:</p>\n\n<pre><code> product_ids = product_ids[:10000]\n category_ids = category_ids[:10000]\n n_pics = n_pics[:10000]\n</code></pre>",
          "rawMarkdown": "In the code at the beginning of train.py that reads the product_id's, category_id's and num of images per product_id:\n\n    df = pd.read_csv(csv_file)\n    product_ids, category_ids, n_pics = df['product_id'], df['category_id'], df['n_pics']\n\nYou can simply only use the first 10k elements of the product_ids, category_ids, n_pics arrays. So you can add something like this:\n\n     product_ids = product_ids[:10000]\n     category_ids = category_ids[:10000]\n     n_pics = n_pics[:10000]\n\n\n"
        },
        {
          "id": 233939,
          "postDate": "2017-10-21T21:22:32.193Z",
          "content": "<p>Peter - much appreciated I will give that a try and see how things play out.</p>",
          "rawMarkdown": "Peter - much appreciated I will give that a try and see how things play out."
        },
        {
          "id": 233983,
          "postDate": "2017-10-22T00:33:02.953Z",
          "content": "<p>Peter - I am getting an error on line 55 stratify=category_ids resulting in this error output:</p>\n\n<p>ValueError: The least populated class in y has only 1 member, which is too few. The minimum number of groups for any class cannot be less than 2.</p>\n\n<p>I had a play around with no effect, but I am running out of time this week (I need to leave town). My guess is its the train/validation split not going to plan with only 10000 values? Something else needs tweaking.</p>",
          "rawMarkdown": "Peter - I am getting an error on line 55 stratify=category_ids resulting in this error output:\n\nValueError: The least populated class in y has only 1 member, which is too few. The minimum number of groups for any class cannot be less than 2.\n\nI had a play around with no effect, but I am running out of time this week (I need to leave town). My guess is its the train/validation split not going to plan with only 10000 values? Something else needs tweaking.\n\n"
        },
        {
          "id": 234068,
          "postDate": "2017-10-22T07:57:29.200Z",
          "content": "<p>Your guess is right, sklearn complains about the stratified split with 10000 values because some classes end up with just 1 product. You can do a simple split instead of stratified by removing the 'stratify=category_ids' parameter and then it'll work.</p>",
          "rawMarkdown": "Your guess is right, sklearn complains about the stratified split with 10000 values because some classes end up with just 1 product. You can do a simple split instead of stratified by removing the 'stratify=category_ids' parameter and then it'll work."
        },
        {
          "id": 236759,
          "postDate": "2017-10-28T03:42:26.420Z",
          "content": "<p>I am getting a new error on line 63: 'for pic in range(n_pics_train[prod]):'</p>\n\n<p>Its a key error and is a result of the length of the data changes made above. I have had a look at the variables but so far I could not locate the issue. So supposedly its got no key for some value I assume someplace in one of the created lists?</p>",
          "rawMarkdown": "I am getting a new error on line 63: 'for pic in range(n_pics_train[prod]):'\n\nIts a key error and is a result of the length of the data changes made above. I have had a look at the variables but so far I could not locate the issue. So supposedly its got no key for some value I assume someplace in one of the created lists?"
        },
        {
          "id": 236787,
          "postDate": "2017-10-28T07:30:23.940Z",
          "content": "<p>ok this is what I needed to do, 'print' couldn't save me this time but I figured it out below. BUT the valid generator function has 'blown a gasket' so I am trying to figure that out as it runs through the test and stops just on the last iteration!</p>\n\n<p>print('Creating training data:')\nfor prod in tqdm(range(len(product_ids_train))):\n    for pic in range(n_pics_train[9000]):\n        product_ids_pics_train_arr.append(str(product_ids_train[9000]) + '_' + str(pic))\n        category_ids_train_arr.append(category_ids_train[9000])\nproduct_ids_pics_test_arr = []\ncategory_ids_test_arr = []\nprint('Creating validation data:')\nfor prod in tqdm(range(len(product_ids_test))):\n    for pic in range(1000):\n        product_ids_pics_test_arr.append(str(1000) + '_' + str(pic))\n        category_ids_test_arr.append(1000)\ntrain_samples = len(product_ids_pics_train_arr)\ntest_samples = len(product_ids_pics_test_arr)</p>",
          "rawMarkdown": "ok this is what I needed to do, 'print' couldn't save me this time but I figured it out below. BUT the valid generator function has 'blown a gasket' so I am trying to figure that out as it runs through the test and stops just on the last iteration!\n\nprint('Creating training data:')\nfor prod in tqdm(range(len(product_ids_train))):\n    for pic in range(n_pics_train[9000]):\n        product_ids_pics_train_arr.append(str(product_ids_train[9000]) + '_' + str(pic))\n        category_ids_train_arr.append(category_ids_train[9000])\nproduct_ids_pics_test_arr = []\ncategory_ids_test_arr = []\nprint('Creating validation data:')\nfor prod in tqdm(range(len(product_ids_test))):\n    for pic in range(1000):\n        product_ids_pics_test_arr.append(str(1000) + '_' + str(pic))\n        category_ids_test_arr.append(1000)\ntrain_samples = len(product_ids_pics_train_arr)\ntest_samples = len(product_ids_pics_test_arr)"
        }
      ]
    },
    {
      "id": 228768,
      "postDate": "2017-10-07T19:39:06.710Z",
      "content": "<p>Hello, thanks for sharing</p>\n\n<p>How long should a VGG16 or inception v3 epoch take?</p>",
      "rawMarkdown": "Hello, thanks for sharing\n\nHow long should a VGG16 or inception v3 epoch take?",
      "replies": [
        {
          "id": 228778,
          "postDate": "2017-10-07T20:30:14.043Z",
          "content": "<p>I'm currently training on the full dataset. 1 epoch will take ~80 hours on my GTX 1070...</p>",
          "rawMarkdown": "I'm currently training on the full dataset. 1 epoch will take ~80 hours on my GTX 1070...",
          "votes": 1
        },
        {
          "id": 228963,
          "postDate": "2017-10-08T12:23:03.650Z",
          "content": "<p>It is too slow.  I think that  you  missed something important.  Do you have a look the post :<a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a> ?  It used 8 workers to speed up the reading of datas and got a better speed.  On my 1080ti gpu, it will take about 22 hours per epoch. </p>",
          "rawMarkdown": "It is too slow.  I think that  you  missed something important.  Do you have a look the post :https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson ?  It used 8 workers to speed up the reading of datas and got a better speed.  On my 1080ti gpu, it will take about 22 hours per epoch. "
        },
        {
          "id": 228967,
          "postDate": "2017-10-08T12:30:11.730Z",
          "content": "<p>Perhaps a reason it's so slow for me is because i'm reading the data from HDD as I don't have enough space on my SSD's.</p>",
          "rawMarkdown": "Perhaps a reason it's so slow for me is because i'm reading the data from HDD as I don't have enough space on my SSD's."
        },
        {
          "id": 228982,
          "postDate": "2017-10-08T13:14:54.343Z",
          "content": "<p>for your information, this is also discussed here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a>:</p>\n\n<p>\"Hi, Heng, I use similar way to load data as you, converting bson to files.</p>\n\n<p>The loading process is very fast at first(0.3s, CPU 30%, RAM 10% of 30G), but some iterations later, it becomes very slow(6s, CPU 5%, RAM 10% of 30G).</p>\n\n<p>I set the CPU mode to performance, and try different num_workers(0,4,8,12 …), but the situation stays the same. ...\"</p>",
          "rawMarkdown": "for your information, this is also discussed here:\n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498:\n\n\"Hi, Heng, I use similar way to load data as you, converting bson to files.\n\nThe loading process is very fast at first(0.3s, CPU 30%, RAM 10% of 30G), but some iterations later, it becomes very slow(6s, CPU 5%, RAM 10% of 30G).\n\nI set the CPU mode to performance, and try different num_workers(0,4,8,12 …), but the situation stays the same. ...\"",
          "votes": 3
        }
      ]
    },
    {
      "id": 228529,
      "postDate": "2017-10-07T00:24:41.793Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 231530,
      "author_name": "heroxrq",
      "author_url": "",
      "post_date": "2017-10-15T09:07:16.680000",
      "content": "<p>I suffered from training very slow when using mxnet. I found that if random read raw images, it is very slow. Now I use the rec file (it is a file generated from all images, hence is fast) of mxnet, it improves from 140 imgs/second to 600 imgs/second.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 233722,
      "author_name": "stephenl",
      "author_url": "",
      "post_date": "2017-10-20T23:07:54.557000",
      "content": "<p>I took at look at your Keras starter - a few good learning points, I did a trip to Keras.io a few times, I like the 'callbacks'. \nA few observations and a question: The png files generated for the train folder are in the hundreds of gigabytes (caveat emptor) so it goes without saying -training and manipulation- can take a very very long time to those stepping into this. Don't waste time using a standard HDD, nothing less than an m.2 SSD will do. Just doing a simple ls -la can take 5 minutes. My SSD which is much quicker unfortunately is too small to hold that much data! None-the-less what variables did you tinker with to run just a 10k sample using train.py? Usually I can spot this quickly but things look a little nested. As a base line I want to replicate your result as shown. Thus far I cannot. It would help square things up.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 233796,
          "author_name": "Petros Giannakopoulos",
          "author_url": "",
          "post_date": "2017-10-21T09:17:18.270000",
          "content": "<p>In the code at the beginning of train.py that reads the product_id's, category_id's and num of images per product_id:</p>\n\n<pre><code>df = pd.read_csv(csv_file)\nproduct_ids, category_ids, n_pics = df['product_id'], df['category_id'], df['n_pics']\n</code></pre>\n\n<p>You can simply only use the first 10k elements of the product_ids, category_ids, n_pics arrays. So you can add something like this:</p>\n\n<pre><code> product_ids = product_ids[:10000]\n category_ids = category_ids[:10000]\n n_pics = n_pics[:10000]\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 233939,
          "author_name": "stephenl",
          "author_url": "",
          "post_date": "2017-10-21T21:22:32.193000",
          "content": "<p>Peter - much appreciated I will give that a try and see how things play out.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 233983,
          "author_name": "stephenl",
          "author_url": "",
          "post_date": "2017-10-22T00:33:02.953000",
          "content": "<p>Peter - I am getting an error on line 55 stratify=category_ids resulting in this error output:</p>\n\n<p>ValueError: The least populated class in y has only 1 member, which is too few. The minimum number of groups for any class cannot be less than 2.</p>\n\n<p>I had a play around with no effect, but I am running out of time this week (I need to leave town). My guess is its the train/validation split not going to plan with only 10000 values? Something else needs tweaking.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 234068,
          "author_name": "Petros Giannakopoulos",
          "author_url": "",
          "post_date": "2017-10-22T07:57:29.200000",
          "content": "<p>Your guess is right, sklearn complains about the stratified split with 10000 values because some classes end up with just 1 product. You can do a simple split instead of stratified by removing the 'stratify=category_ids' parameter and then it'll work.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 236759,
          "author_name": "stephenl",
          "author_url": "",
          "post_date": "2017-10-28T03:42:26.420000",
          "content": "<p>I am getting a new error on line 63: 'for pic in range(n_pics_train[prod]):'</p>\n\n<p>Its a key error and is a result of the length of the data changes made above. I have had a look at the variables but so far I could not locate the issue. So supposedly its got no key for some value I assume someplace in one of the created lists?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 236787,
          "author_name": "stephenl",
          "author_url": "",
          "post_date": "2017-10-28T07:30:23.940000",
          "content": "<p>ok this is what I needed to do, 'print' couldn't save me this time but I figured it out below. BUT the valid generator function has 'blown a gasket' so I am trying to figure that out as it runs through the test and stops just on the last iteration!</p>\n\n<p>print('Creating training data:')\nfor prod in tqdm(range(len(product_ids_train))):\n    for pic in range(n_pics_train[9000]):\n        product_ids_pics_train_arr.append(str(product_ids_train[9000]) + '_' + str(pic))\n        category_ids_train_arr.append(category_ids_train[9000])\nproduct_ids_pics_test_arr = []\ncategory_ids_test_arr = []\nprint('Creating validation data:')\nfor prod in tqdm(range(len(product_ids_test))):\n    for pic in range(1000):\n        product_ids_pics_test_arr.append(str(1000) + '_' + str(pic))\n        category_ids_test_arr.append(1000)\ntrain_samples = len(product_ids_pics_train_arr)\ntest_samples = len(product_ids_pics_test_arr)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 228768,
      "author_name": "Juan Pizarro",
      "author_url": "",
      "post_date": "2017-10-07T19:39:06.710000",
      "content": "<p>Hello, thanks for sharing</p>\n\n<p>How long should a VGG16 or inception v3 epoch take?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 228778,
          "author_name": "Petros Giannakopoulos",
          "author_url": "",
          "post_date": "2017-10-07T20:30:14.043000",
          "content": "<p>I'm currently training on the full dataset. 1 epoch will take ~80 hours on my GTX 1070...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 228963,
          "author_name": "huiqin",
          "author_url": "",
          "post_date": "2017-10-08T12:23:03.650000",
          "content": "<p>It is too slow.  I think that  you  missed something important.  Do you have a look the post :<a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a> ?  It used 8 workers to speed up the reading of datas and got a better speed.  On my 1080ti gpu, it will take about 22 hours per epoch. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 228967,
          "author_name": "Petros Giannakopoulos",
          "author_url": "",
          "post_date": "2017-10-08T12:30:11.730000",
          "content": "<p>Perhaps a reason it's so slow for me is because i'm reading the data from HDD as I don't have enough space on my SSD's.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 228982,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2017-10-08T13:14:54.343000",
          "content": "<p>for your information, this is also discussed here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a>:</p>\n\n<p>\"Hi, Heng, I use similar way to load data as you, converting bson to files.</p>\n\n<p>The loading process is very fast at first(0.3s, CPU 30%, RAM 10% of 30G), but some iterations later, it becomes very slow(6s, CPU 5%, RAM 10% of 30G).</p>\n\n<p>I set the CPU mode to performance, and try different num_workers(0,4,8,12 …), but the situation stays the same. ...\"</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 228529,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-10-07T00:24:41.793000",
      "content": "",
      "votes": -3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "228504": "Hello everyone! This looks like a challenging classification task due to the very large dataset and large number of classes. \n\nThis is my code so far for this competition, I hope that it helps more people to join and that it provides a decent baseline to build a robust solution on:\n\nhttps://github.com/petrosgk/Kaggle-Cdiscount-Image-Classification-Challenge\n\nI am still processing the .bson files so no results on the full dataset yet. These are 5 epochs on a subset of 10k samples with pre-trained VGG16 (with top dense layers removed) and image size 180x180:\n\n    Epoch 1/20\n    231s - loss: 3.7500 - acc: 0.4056 - val_loss: 3.2170 - val_acc: 0.4783\n    Epoch 2/20\n    227s - loss: 2.8793 - acc: 0.5010 - val_loss: 3.0092 - val_acc: 0.5149\n    Epoch 3/20\n    228s - loss: 2.4361 - acc: 0.5570 - val_loss: 2.9464 - val_acc: 0.5323\n    Epoch 4/20\n    229s - loss: 2.1199 - acc: 0.5996 - val_loss: 2.9545 - val_acc: 0.5447\n    Epoch 5/20\n    229s - loss: 1.8476 - acc: 0.6346 - val_loss: 2.9870 - val_acc: 0.5440\n\nThere is overfitting with 9k training samples (90/10 train-val split) but that may change when training on the full ~7m samples.",
    "231530": "I suffered from training very slow when using mxnet. I found that if random read raw images, it is very slow. Now I use the rec file (it is a file generated from all images, hence is fast) of mxnet, it improves from 140 imgs/second to 600 imgs/second.",
    "233722": "I took at look at your Keras starter - a few good learning points, I did a trip to Keras.io a few times, I like the 'callbacks'. \nA few observations and a question: The png files generated for the train folder are in the hundreds of gigabytes (caveat emptor) so it goes without saying -training and manipulation- can take a very very long time to those stepping into this. Don't waste time using a standard HDD, nothing less than an m.2 SSD will do. Just doing a simple ls -la can take 5 minutes. My SSD which is much quicker unfortunately is too small to hold that much data! None-the-less what variables did you tinker with to run just a 10k sample using train.py? Usually I can spot this quickly but things look a little nested. As a base line I want to replicate your result as shown. Thus far I cannot. It would help square things up.",
    "228768": "Hello, thanks for sharing\n\nHow long should a VGG16 or inception v3 epoch take?",
    "228529": ""
  }
}