{
  "id": 19971,
  "title": "Simple solution (Keras)",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/19971",
  "author_name": "",
  "post_date": "2016-04-06T10:53:18.427Z",
  "votes": 46,
  "comment_count": 64,
  "views": 18451,
  "content": "<p>Here is simple solution using CNN to start from:</p>\n\n<p>Ver. 1: <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py</a></p>\n\n<p>Ver. 2 (<strong>UPD 07.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Ver. 3 (<strong>UPD 09.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>Ver. 4 (<strong>UPD 03.05</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py</a></p>\n\n<p>Ver. 5 (<strong>UPD 29.07</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py</a> - pretrained VGG16 Net.</p>",
  "messages": [
    {
      "id": "113939",
      "postDate": "04/06/2016 10:53:18",
      "content": "<p>Here is simple solution using CNN to start from:</p>\n\n<p>Ver. 1: <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py</a></p>\n\n<p>Ver. 2 (<strong>UPD 07.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Ver. 3 (<strong>UPD 09.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>Ver. 4 (<strong>UPD 03.05</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py</a></p>\n\n<p>Ver. 5 (<strong>UPD 29.07</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py</a> - pretrained VGG16 Net.</p>",
      "rawMarkdown": "Here is simple solution using CNN to start from:\r\n\r\nVer. 1: https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\r\n\r\nVer. 2 (**UPD 07.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nVer. 3 (**UPD 09.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\nVer. 4 (**UPD 03.05**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\r\n\r\nVer. 5 (**UPD 29.07**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py - pretrained VGG16 Net.",
      "votes": null
    },
    {
      "id": "113963",
      "postDate": "04/06/2016 13:41:49",
      "content": "<p>The training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.</p>",
      "rawMarkdown": "The training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.",
      "votes": null
    },
    {
      "id": "114014",
      "postDate": "04/06/2016 21:22:13",
      "content": "<p>[quote=Mike;113963]\nThe training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.\n[/quote]</p>\n\n<p>You are right. Need to find out the way how to deal with it. )</p>",
      "rawMarkdown": "[quote=Mike;113963]\r\nThe training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.\r\n[/quote]\r\n\r\nYou are right. Need to find out the way how to deal with it. )",
      "votes": null
    },
    {
      "id": "114041",
      "postDate": "04/07/2016 07:21:51",
      "content": "<p>How did you come up with (128, 96) for resized images?</p>",
      "rawMarkdown": "How did you come up with (128, 96) for resized images?",
      "votes": null
    },
    {
      "id": "114046",
      "postDate": "04/07/2016 08:06:49",
      "content": "<p>[quote=Khanh;114041]\nHow did you come up with (128, 96) for resized images?\n[/quote]</p>\n\n<p>I needed to decrease training time consumption. So I choose reasonable picture size where I still can classify images by eyes. Actually in current model reducing pictures even more to (64, 48) gives better LB result. It still requires more experiments with parameters and layers tuning.</p>",
      "rawMarkdown": "[quote=Khanh;114041]\r\nHow did you come up with (128, 96) for resized images?\r\n[/quote]\r\n\r\nI needed to decrease training time consumption. So I choose reasonable picture size where I still can classify images by eyes. Actually in current model reducing pictures even more to (64, 48) gives better LB result. It still requires more experiments with parameters and layers tuning.",
      "votes": null
    },
    {
      "id": "114059",
      "postDate": "04/07/2016 10:40:30",
      "content": "<p>May I ask how long does it take to train the model for (128, 96)  and on what kind of setup? </p>",
      "rawMarkdown": "May I ask how long does it take to train the model for (128, 96)  and on what kind of setup?",
      "votes": null
    },
    {
      "id": "114062",
      "postDate": "04/07/2016 10:57:15",
      "content": "<p>Processor: Intel(R) Core(TM) i7-2600K CPU @ 3.40GHz (8 CPUs), ~3.4GHz</p>\n\n<p>Memory: 8192MB RAM</p>\n\n<p>GPU: NVIDIA GeForce GTX 560 Ti 1GB</p>\n\n<p>Requires around 10-15 minutes overall in GPU mode. Half of the time is image reading.</p>",
      "rawMarkdown": "Processor: Intel(R) Core(TM) i7-2600K CPU @ 3.40GHz (8 CPUs), ~3.4GHz\r\n\r\nMemory: 8192MB RAM\r\n\r\nGPU: NVIDIA GeForce GTX 560 Ti 1GB\r\n\r\nRequires around 10-15 minutes overall in GPU mode. Half of the time is image reading.",
      "votes": null
    },
    {
      "id": "114129",
      "postDate": "04/07/2016 17:39:19",
      "content": "<p>ZFTurbo, are you using exactly what you have on Github to get 1.3 on LB? Im not doing as well locally. I just wanted to be sure that I can produce the expected result before I start messing with things  </p>",
      "rawMarkdown": "ZFTurbo, are you using exactly what you have on Github to get 1.3 on LB? Im not doing as well locally. I just wanted to be sure that I can produce the expected result before I start messing with things",
      "votes": null
    },
    {
      "id": "114132",
      "postDate": "04/07/2016 17:46:40",
      "content": "<p>No. Code provided on GitHub, will be around 2.20. </p>\n\n<p>Change: nb_epoch = 1, img_rows, img_cols = 48, 64 to achieve better results. </p>\n\n<p>And as I said in the first post: local and leaderboard score is totally different.</p>",
      "rawMarkdown": "No. Code provided on GitHub, will be around 2.20. \r\n\r\nChange: nb_epoch = 1, img_rows, img_cols = 48, 64 to achieve better results. \r\n\r\nAnd as I said in the first post: local and leaderboard score is totally different.",
      "votes": null
    },
    {
      "id": "114139",
      "postDate": "04/07/2016 18:32:01",
      "content": "<p>Cool thanks, 2.2 is in the neighborhood of what I'm getting, what should I expect if I make the above changes? Thanks this is helpful for benchmarking. </p>",
      "rawMarkdown": "Cool thanks, 2.2 is in the neighborhood of what I'm getting, what should I expect if I make the above changes? Thanks this is helpful for benchmarking.",
      "votes": null
    },
    {
      "id": "114151",
      "postDate": "04/07/2016 20:02:26",
      "content": "<p><strong>DrewWham</strong>: it'll be around ~2.0. The next steps to optimize solution is to go for cross-validation technique, also even more reduction of initial images made solution better for some purpose. Looks like small resolution make CNN focus on overall picture, than on driver clothes or something. ) My current Keras code with cross-validation:</p>\n\n<p><a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Allows to reach ~1.4 on leaderboard. It also pretty fast, requires around 15 minutes on GPU.</p>\n\n<p><strong>Question</strong>: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?</p>",
      "rawMarkdown": "**DrewWham**: it'll be around ~2.0. The next steps to optimize solution is to go for cross-validation technique, also even more reduction of initial images made solution better for some purpose. Looks like small resolution make CNN focus on overall picture, than on driver clothes or something. ) My current Keras code with cross-validation:\r\n\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nAllows to reach ~1.4 on leaderboard. It also pretty fast, requires around 15 minutes on GPU.\r\n\r\n**Question**: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?",
      "votes": null
    },
    {
      "id": "114154",
      "postDate": "04/07/2016 20:39:58",
      "content": "<p>[quote=ZFTurbo;114151]</p>\n\n<p><strong>Question</strong>: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?</p>\n\n<p>[/quote]\nThank you very much for your code. Usually geometric mean works better for logloss like metrics. And you could also try stacking  <a href=\"http://mlwave.com/kaggle-ensembling-guide/\">http://mlwave.com/kaggle-ensembling-guide/</a></p>",
      "rawMarkdown": "[quote=ZFTurbo;114151]\r\n\r\n**Question**: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?\r\n\r\n[/quote]\r\nThank you very much for your code. Usually geometric mean works better for logloss like metrics. And you could also try stacking  http://mlwave.com/kaggle-ensembling-guide/",
      "votes": null
    },
    {
      "id": "114183",
      "postDate": "04/08/2016 00:55:44",
      "content": "<p>This is weird, I implemented a very similar model before I found this thread and while I was using some slightly different parameters and larger images, my LB score is waaay off from my logloss (anywhere from 7-15).  Saw this thread and thought maybe just a bad model, and tried running your  run_keras_cv.py and my lb score was ~9.  </p>",
      "rawMarkdown": "This is weird, I implemented a very similar model before I found this thread and while I was using some slightly different parameters and larger images, my LB score is waaay off from my logloss (anywhere from 7-15).  Saw this thread and thought maybe just a bad model, and tried running your  run_keras_cv.py and my lb score was ~9.",
      "votes": null
    },
    {
      "id": "114188",
      "postDate": "04/08/2016 02:17:37",
      "content": "<p>There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.</p>\n\n<p>Using the newly added &quot;driver_imgs_list.csv&quot; should make it easier to validate models.</p>",
      "rawMarkdown": "There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.\r\n\r\nUsing the newly added \"driver_imgs_list.csv\" should make it easier to validate models.",
      "votes": null
    },
    {
      "id": "114215",
      "postDate": "04/08/2016 08:42:47",
      "content": "<p>[quote=Ryan Pream;114188]</p>\n\n<p>There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.</p>\n\n<p>Using the newly added &quot;driver_imgs_list.csv&quot; should make it easier to validate models.</p>\n\n<p>[/quote]</p>\n\n<p>Yeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:</p>\n\n<pre><code>def select_subset_of_driver(percentage_split=.2):\n    &quot;&quot;&quot;\n    &quot;&quot;&quot;\n    driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\n    all_ids = list(driver_df.subject.unique())\n    # for testing\n    np.random.seed(69)\n    np.random.shuffle(all_ids)\n    valid_amount = int(len(all_ids) * percentage_split)\n    train_x_drivers = all_ids[valid_amount:]\n    test_x_drivers = all_ids[:valid_amount]\n    print('-' * 50)\n    print('using subset of drivers as validation')\n    print('using following ids for training: ', train_x_drivers)\n    print('using following ids for testing: ', test_x_drivers)\n    print('-' * 50)\n    test_imgs = driver_df[\n        driver_df['subject'].isin(test_x_drivers)]['img'].values\n    train_imgs = driver_df[\n        driver_df['subject'].isin(train_x_drivers)]['img'].values\n\n    return train_imgs, test_imgs\n\nX, y, train_ids = load_train_data()\ntrain_imgs, validate_imgs = select_subset_of_driver()\n# select indices of validate and train data\nvalidate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\ntrain_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\n# validate subset\nX_validate = X[validate_idx]\ny_validate = y[validate_idx]\nvalidate_ids = train_ids[validate_idx]\n# train subset\nX = X[train_idx]\ny = y[train_idx]\ntrain_ids = train_ids[train_idx]\n</code></pre>\n\n<p>My hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on</p>",
      "rawMarkdown": "[quote=Ryan Pream;114188]\r\n\r\nThere are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.\r\n\r\nUsing the newly added \"driver_imgs_list.csv\" should make it easier to validate models.\r\n\r\n[/quote]\r\n\r\n\r\nYeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:\r\n\r\n    def select_subset_of_driver(percentage_split=.2):\r\n        \"\"\"\r\n        \"\"\"\r\n        driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\r\n        all_ids = list(driver_df.subject.unique())\r\n        # for testing\r\n        np.random.seed(69)\r\n        np.random.shuffle(all_ids)\r\n        valid_amount = int(len(all_ids) * percentage_split)\r\n        train_x_drivers = all_ids[valid_amount:]\r\n        test_x_drivers = all_ids[:valid_amount]\r\n        print('-' * 50)\r\n        print('using subset of drivers as validation')\r\n        print('using following ids for training: ', train_x_drivers)\r\n        print('using following ids for testing: ', test_x_drivers)\r\n        print('-' * 50)\r\n        test_imgs = driver_df[\r\n            driver_df['subject'].isin(test_x_drivers)]['img'].values\r\n        train_imgs = driver_df[\r\n            driver_df['subject'].isin(train_x_drivers)]['img'].values\r\n    \r\n        return train_imgs, test_imgs\r\n\r\n    X, y, train_ids = load_train_data()\r\n    train_imgs, validate_imgs = select_subset_of_driver()\r\n    # select indices of validate and train data\r\n    validate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\r\n    train_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\r\n    # validate subset\r\n    X_validate = X[validate_idx]\r\n    y_validate = y[validate_idx]\r\n    validate_ids = train_ids[validate_idx]\r\n    # train subset\r\n    X = X[train_idx]\r\n    y = y[train_idx]\r\n    train_ids = train_ids[train_idx]\r\n\r\n\r\nMy hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on",
      "votes": null
    },
    {
      "id": "114220",
      "postDate": "04/08/2016 10:21:03",
      "content": "<p>Hi hassiktir\nYou mentioned that trying to run run_keras_cv.py produced strange results. Can you say how you bypassed this as I get results around 4 when running it for 10 epochs\nthnx</p>",
      "rawMarkdown": "Hi hassiktir\r\nYou mentioned that trying to run run_keras_cv.py produced strange results. Can you say how you bypassed this as I get results around 4 when running it for 10 epochs\r\nthnx",
      "votes": null
    },
    {
      "id": "114232",
      "postDate": "04/08/2016 11:51:46",
      "content": "<p>[quote=hassiktir;114215]</p>\n\n<p>[quote=Ryan Pream;114188]</p>\n\n<p>There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.</p>\n\n<p>Using the newly added &quot;driver_imgs_list.csv&quot; should make it easier to validate models.</p>\n\n<p>[/quote]</p>\n\n<p>Yeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:</p>\n\n<pre><code>def select_subset_of_driver(percentage_split=.2):\n    &quot;&quot;&quot;\n    &quot;&quot;&quot;\n    driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\n    all_ids = list(driver_df.subject.unique())\n    # for testing\n    np.random.seed(69)\n    np.random.shuffle(all_ids)\n    valid_amount = int(len(all_ids) * percentage_split)\n    train_x_drivers = all_ids[valid_amount:]\n    test_x_drivers = all_ids[:valid_amount]\n    print('-' * 50)\n    print('using subset of drivers as validation')\n    print('using following ids for training: ', train_x_drivers)\n    print('using following ids for testing: ', test_x_drivers)\n    print('-' * 50)\n    test_imgs = driver_df[\n        driver_df['subject'].isin(test_x_drivers)]['img'].values\n    train_imgs = driver_df[\n        driver_df['subject'].isin(train_x_drivers)]['img'].values\n\n    return train_imgs, test_imgs\n\nX, y, train_ids = load_train_data()\ntrain_imgs, validate_imgs = select_subset_of_driver()\n# select indices of validate and train data\nvalidate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\ntrain_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\n# validate subset\nX_validate = X[validate_idx]\ny_validate = y[validate_idx]\nvalidate_ids = train_ids[validate_idx]\n# train subset\nX = X[train_idx]\ny = y[train_idx]\ntrain_ids = train_ids[train_idx]\n</code></pre>\n\n<p>My hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on</p>\n\n<p>[/quote]</p>\n\n<p>You probably need the LeavePLabelOut function (<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.LeavePLabelOut.html\">http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.LeavePLabelOut.html</a>)</p>",
      "rawMarkdown": "[quote=hassiktir;114215]\r\n\r\n[quote=Ryan Pream;114188]\r\n\r\nThere are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.\r\n\r\nUsing the newly added \"driver_imgs_list.csv\" should make it easier to validate models.\r\n\r\n[/quote]\r\n\r\n\r\nYeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:\r\n\r\n    def select_subset_of_driver(percentage_split=.2):\r\n        \"\"\"\r\n        \"\"\"\r\n        driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\r\n        all_ids = list(driver_df.subject.unique())\r\n        # for testing\r\n        np.random.seed(69)\r\n        np.random.shuffle(all_ids)\r\n        valid_amount = int(len(all_ids) * percentage_split)\r\n        train_x_drivers = all_ids[valid_amount:]\r\n        test_x_drivers = all_ids[:valid_amount]\r\n        print('-' * 50)\r\n        print('using subset of drivers as validation')\r\n        print('using following ids for training: ', train_x_drivers)\r\n        print('using following ids for testing: ', test_x_drivers)\r\n        print('-' * 50)\r\n        test_imgs = driver_df[\r\n            driver_df['subject'].isin(test_x_drivers)]['img'].values\r\n        train_imgs = driver_df[\r\n            driver_df['subject'].isin(train_x_drivers)]['img'].values\r\n    \r\n        return train_imgs, test_imgs\r\n\r\n    X, y, train_ids = load_train_data()\r\n    train_imgs, validate_imgs = select_subset_of_driver()\r\n    # select indices of validate and train data\r\n    validate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\r\n    train_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\r\n    # validate subset\r\n    X_validate = X[validate_idx]\r\n    y_validate = y[validate_idx]\r\n    validate_ids = train_ids[validate_idx]\r\n    # train subset\r\n    X = X[train_idx]\r\n    y = y[train_idx]\r\n    train_ids = train_ids[train_idx]\r\n\r\n\r\nMy hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on\r\n\r\n[/quote]\r\n\r\nYou probably need the LeavePLabelOut function (http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.LeavePLabelOut.html)",
      "votes": null
    },
    {
      "id": "114234",
      "postDate": "04/08/2016 12:03:53",
      "content": "<p><em>There are 28 drivers</em></p>\n\n<p>Actually there are 26 drivers. )</p>",
      "rawMarkdown": "*There are 28 drivers*\r\n\r\nActually there are 26 drivers. )",
      "votes": null
    },
    {
      "id": "114339",
      "postDate": "04/09/2016 14:21:57",
      "content": "<p>I created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.</p>\n\n<p><a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>Now it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:</p>\n\n<ol>\n<li>Is there any docs/papers/faqs for dummies how to construct the\nCNN models? Most of the docs concentrate on MNIST, which is not the\ncase.</li>\n<li>I tried to follow VGG-16 model scheme for construction of\nCNNs, but adding second conv/pool block actually make LOSS function\nworse. After tuning of filter number and convolution kernel it's\nsometimes became better. How to find out the optimal number of\nconvolution/pooling/dense layers? Does it somehow depends on input\nimage width/height? Is there some intuitive predictions which CNN\nmodel will be better for given task? </li>\n<li>On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\nquess. And it doesn't improve with epoch number. Otherwise on simple\nmodels while train loss decrease, valid loss either jump, or\ncontinue increasing with each epoch (overfit?). How to make it\ndecrease with each step? I find out that Dropout layers sometimes\nhelp with overfitting. May be some other tricks exists? </li>\n<li>Is there any way in Keras to feed different picture areas to different CNNs\nwith later merging them in one bigger CNN at some stage? How do you\nthink will it make model better? </li>\n<li>As I can see &quot;fit&quot; on CNN works\ntotally unpredictable comapring to XGBoost. Should successfull\nmodels use many epochs? Is there some CNNs, which used for some\nreallife problems to predict something, with only 1 epoch? What\n&quot;optimizer&quot; is the best (I tried adadelta and SGD)? </li>\n<li>How to find out which initial picture size is optimal as input for CNN? Should\nwe use gray or full RGB for this problem?</li>\n</ol>",
      "rawMarkdown": "I created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.\r\n\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\nNow it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:\r\n\r\n 1. Is there any docs/papers/faqs for dummies how to construct the\r\n    CNN models? Most of the docs concentrate on MNIST, which is not the\r\n    case.\r\n 2. I tried to follow VGG-16 model scheme for construction of\r\n    CNNs, but adding second conv/pool block actually make LOSS function\r\n    worse. After tuning of filter number and convolution kernel it's\r\n    sometimes became better. How to find out the optimal number of\r\n    convolution/pooling/dense layers? Does it somehow depends on input\r\n    image width/height? Is there some intuitive predictions which CNN\r\n    model will be better for given task? \r\n 3. On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\r\n    quess. And it doesn't improve with epoch number. Otherwise on simple\r\n    models while train loss decrease, valid loss either jump, or\r\n    continue increasing with each epoch (overfit?). How to make it\r\n    decrease with each step? I find out that Dropout layers sometimes\r\n    help with overfitting. May be some other tricks exists? \r\n 4. Is there any way in Keras to feed different picture areas to different CNNs\r\n    with later merging them in one bigger CNN at some stage? How do you\r\n    think will it make model better? \r\n 5. As I can see \"fit\" on CNN works\r\n    totally unpredictable comapring to XGBoost. Should successfull\r\n    models use many epochs? Is there some CNNs, which used for some\r\n    reallife problems to predict something, with only 1 epoch? What\r\n    \"optimizer\" is the best (I tried adadelta and SGD)? \r\n 6. How to find out which initial picture size is optimal as input for CNN? Should\r\n    we use gray or full RGB for this problem?",
      "votes": null
    },
    {
      "id": "114345",
      "postDate": "04/09/2016 15:08:31",
      "content": "<p>[quote=ZFTurbo;114339]</p>\n\n<p>I created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.</p>\n\n<p><a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>Now it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:</p>\n\n<ol>\n<li>Is there any docs/papers/faqs for dummies how to construct the\nCNN models? Most of the docs concentrate on MNIST, which is not the\ncase.</li>\n<li>I tried to follow VGG-16 model scheme for construction of\nCNNs, but adding second conv/pool block actually make LOSS function\nworse. After tuning of filter number and convolution kernel it's\nsometimes became better. How to find out the optimal number of\nconvolution/pooling/dense layers? Does it somehow depends on input\nimage width/height? Is there some intuitive predictions which CNN\nmodel will be better for given task? </li>\n<li>On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\nquess. And it doesn't improve with epoch number. Otherwise on simple\nmodels while train loss decrease, valid loss either jump, or\ncontinue increasing with each epoch (overfit?). How to make it\ndecrease with each step? I find out that Dropout layers sometimes\nhelp with overfitting. May be some other tricks exists? </li>\n<li>Is there any way in Keras to feed different picture areas to different CNNs\nwith later merging them in one bigger CNN at some stage? How do you\nthink will it make model better? </li>\n<li>As I can see &quot;fit&quot; on CNN works\ntotally unpredictable comapring to XGBoost. Should successfull\nmodels use many epochs? Is there some CNNs, which used for some\nreallife problems to predict something, with only 1 epoch? What\n&quot;optimizer&quot; is the best (I tried adadelta and SGD)? </li>\n<li>How to find out which initial picture size is optimal as input for CNN? Should\nwe use gray or full RGB for this problem?</li>\n</ol>\n\n<p>[/quote]</p>\n\n<p>Here are my personal opinions regarding your questions:</p>\n\n<p>1, Another benchmark contest in Computer Vision is ILSVRC (<a href=\"http://www.image-net.org/\">http://www.image-net.org/</a>). The most famous paper is this one (<a href=\"http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf\">http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf</a>).</p>\n\n<p>2, It is tricky to fine-tune a NN. We may need to explorer more structures.</p>\n\n<p>3, One possibility is using lower learning rate. You may replace the optimizer with SGD (<a href=\"http://keras.io/optimizers/\">http://keras.io/optimizers/</a>). You may also find and save the optimal model by using ModelCheckpoint and EarlyStopping (<a href=\"http://keras.io/callbacks/\">http://keras.io/callbacks/</a>).</p>\n\n<p>4, You may use ImageDataGenerator to manipulate the images a little bit (<a href=\"http://keras.io/preprocessing/image/\">http://keras.io/preprocessing/image/</a>). I don't think dividing the whole image to several smaller images could help in this competition.</p>\n\n<p>5, The same as question 3. I prefer to use SGD. batch_size also makes a difference.</p>\n\n<p>6, If the computing power is not a constraint, I prefer to use color images. The image should not be too small. At least, humans should be able to identify the categories of the images.</p>",
      "rawMarkdown": "[quote=ZFTurbo;114339]\r\n\r\nI created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.\r\n\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\nNow it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:\r\n\r\n 1. Is there any docs/papers/faqs for dummies how to construct the\r\n    CNN models? Most of the docs concentrate on MNIST, which is not the\r\n    case.\r\n 2. I tried to follow VGG-16 model scheme for construction of\r\n    CNNs, but adding second conv/pool block actually make LOSS function\r\n    worse. After tuning of filter number and convolution kernel it's\r\n    sometimes became better. How to find out the optimal number of\r\n    convolution/pooling/dense layers? Does it somehow depends on input\r\n    image width/height? Is there some intuitive predictions which CNN\r\n    model will be better for given task? \r\n 3. On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\r\n    quess. And it doesn't improve with epoch number. Otherwise on simple\r\n    models while train loss decrease, valid loss either jump, or\r\n    continue increasing with each epoch (overfit?). How to make it\r\n    decrease with each step? I find out that Dropout layers sometimes\r\n    help with overfitting. May be some other tricks exists? \r\n 4. Is there any way in Keras to feed different picture areas to different CNNs\r\n    with later merging them in one bigger CNN at some stage? How do you\r\n    think will it make model better? \r\n 5. As I can see \"fit\" on CNN works\r\n    totally unpredictable comapring to XGBoost. Should successfull\r\n    models use many epochs? Is there some CNNs, which used for some\r\n    reallife problems to predict something, with only 1 epoch? What\r\n    \"optimizer\" is the best (I tried adadelta and SGD)? \r\n 6. How to find out which initial picture size is optimal as input for CNN? Should\r\n    we use gray or full RGB for this problem?\r\n\r\n[/quote]\r\n\r\nHere are my personal opinions regarding your questions:\r\n\r\n1, Another benchmark contest in Computer Vision is ILSVRC (http://www.image-net.org/). The most famous paper is this one (http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf).\r\n\r\n2, It is tricky to fine-tune a NN. We may need to explorer more structures.\r\n\r\n3, One possibility is using lower learning rate. You may replace the optimizer with SGD (http://keras.io/optimizers/). You may also find and save the optimal model by using ModelCheckpoint and EarlyStopping (http://keras.io/callbacks/).\r\n\r\n4, You may use ImageDataGenerator to manipulate the images a little bit (http://keras.io/preprocessing/image/). I don't think dividing the whole image to several smaller images could help in this competition.\r\n\r\n5, The same as question 3. I prefer to use SGD. batch_size also makes a difference.\r\n\r\n6, If the computing power is not a constraint, I prefer to use color images. The image should not be too small. At least, humans should be able to identify the categories of the images.",
      "votes": null
    },
    {
      "id": "114348",
      "postDate": "04/09/2016 16:28:22",
      "content": "<p>[quote=ZFTurbo;113939]</p>\n\n<p>Here is simple solution using CNN to start from:</p>\n\n<p>Ver. 1: <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py</a></p>\n\n<p>Ver. 2 (<strong>UPD 07.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Ver. 3 (<strong>UPD 09.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a>\n[/quote]</p>\n\n<p>Since I've benefited from your code a bit, I think it would be fair to give you some tips:</p>\n\n<p>1.) You don't need opencv for image processing, there's scipy.misc.imread, imresize <br>\nSaves on memory.</p>\n\n<p>2.) You lose a lot of information by going greyscale</p>\n\n<p>3.) Mean normalization</p>\n\n<p>4.) Hyper-parameter (i.e. learning rate) tuning is very important, the default ones are really bad</p>\n\n<p>5.) 1 epoch is not enough for model to generalize well</p>",
      "rawMarkdown": "[quote=ZFTurbo;113939]\r\n\r\nHere is simple solution using CNN to start from:\r\n\r\nVer. 1: https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\r\n\r\nVer. 2 (**UPD 07.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nVer. 3 (**UPD 09.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n[/quote]\r\n\r\nSince I've benefited from your code a bit, I think it would be fair to give you some tips:\r\n\r\n1.) You don't need opencv for image processing, there's scipy.misc.imread, imresize  \r\nSaves on memory.\r\n\r\n2.) You lose a lot of information by going greyscale\r\n\r\n3.) Mean normalization\r\n\r\n4.) Hyper-parameter (i.e. learning rate) tuning is very important, the default ones are really bad\r\n\r\n5.) 1 epoch is not enough for model to generalize well",
      "votes": null
    },
    {
      "id": "114420",
      "postDate": "04/10/2016 17:02:13",
      "content": "<p>I dont understand how you separate the validation set from the trainset.</p>\n\n<p>Would you clarify?\nI wanted to pass a percentage parameter as well, to get the validation set..</p>\n\n<p>This is rather weird for me!</p>\n\n<p>def load(...):</p>\n\n<pre><code>def copy_selected_drivers(train_data, train_target, driver_id, driver_list):\n        data = []\n        target = []\n        index = []\n        for i in range(len(driver_id)):\n            if driver_id[i] in driver_list:\n                data.append(train_data[i])\n                target.append(train_target[i])\n                index.append(i)\n        data = np.array(data, dtype=np.float32)\n        target = np.array(target, dtype=np.float32)\n        index = np.array(index, dtype=np.uint32)\n        return data, target, index\n\n    train_data, train_target, driver_id, unique_drivers = read_and_normalize_train_data(img_rows, img_cols, color_type_global)\n    test_data, test_id = read_and_normalize_test_data(img_rows, img_cols, color_type_global)\n\n    unique_list_train = ['p002', 'p012', 'p014', 'p015', 'p016', 'p021', 'p022', 'p024',\n                         'p026', 'p035', 'p039', 'p041', 'p042', 'p045', 'p047', 'p049',\n                         'p050', 'p051', 'p052', 'p056', 'p061', 'p064', 'p066', 'p072',\n                         'p075']\n    X_train, Y_train, train_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_train)\n\n    unique_list_valid = ['p081']\n    X_valid, Y_valid, test_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_valid)\n\n    print('Split train: ', len(X_train), len(Y_train))\n    print('Split valid: ', len(X_valid), len(Y_valid))\n    print('Train drivers: ', unique_list_train)\n    print('Test drivers: ', unique_list_valid)\n\nreturn X_train, Y_train, train_index, X_valid, Y_valid, test_index, test_data, test_id\n</code></pre>",
      "rawMarkdown": "I dont understand how you separate the validation set from the trainset.\r\n\r\nWould you clarify?\r\nI wanted to pass a percentage parameter as well, to get the validation set..\r\n\r\nThis is rather weird for me!\r\n\r\ndef load(...):\r\n\r\n    def copy_selected_drivers(train_data, train_target, driver_id, driver_list):\r\n            data = []\r\n            target = []\r\n            index = []\r\n            for i in range(len(driver_id)):\r\n                if driver_id[i] in driver_list:\r\n                    data.append(train_data[i])\r\n                    target.append(train_target[i])\r\n                    index.append(i)\r\n            data = np.array(data, dtype=np.float32)\r\n            target = np.array(target, dtype=np.float32)\r\n            index = np.array(index, dtype=np.uint32)\r\n            return data, target, index\r\n    \r\n        train_data, train_target, driver_id, unique_drivers = read_and_normalize_train_data(img_rows, img_cols, color_type_global)\r\n        test_data, test_id = read_and_normalize_test_data(img_rows, img_cols, color_type_global)\r\n    \r\n        unique_list_train = ['p002', 'p012', 'p014', 'p015', 'p016', 'p021', 'p022', 'p024',\r\n                             'p026', 'p035', 'p039', 'p041', 'p042', 'p045', 'p047', 'p049',\r\n                             'p050', 'p051', 'p052', 'p056', 'p061', 'p064', 'p066', 'p072',\r\n                             'p075']\r\n        X_train, Y_train, train_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_train)\r\n    \r\n        unique_list_valid = ['p081']\r\n        X_valid, Y_valid, test_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_valid)\r\n    \r\n        print('Split train: ', len(X_train), len(Y_train))\r\n        print('Split valid: ', len(X_valid), len(Y_valid))\r\n        print('Train drivers: ', unique_list_train)\r\n        print('Test drivers: ', unique_list_valid)\r\n\r\n    return X_train, Y_train, train_index, X_valid, Y_valid, test_index, test_data, test_id",
      "votes": null
    },
    {
      "id": "114422",
      "postDate": "04/10/2016 17:07:35",
      "content": "<p><strong>Andre lopes</strong>, I split not by images, but by drivers. There are hundreds of images of the same driver. There are only 26 drivers in train set. </p>\n\n<p>In the version of code you provide. 25 drivers used for train and 1 driver for validation. Check &quot;<em>driver_imgs_list.csv</em>&quot;</p>",
      "rawMarkdown": "**Andre lopes**, I split not by images, but by drivers. There are hundreds of images of the same driver. There are only 26 drivers in train set. \r\n\r\nIn the version of code you provide. 25 drivers used for train and 1 driver for validation. Check \"*driver_imgs_list.csv*\"",
      "votes": null
    },
    {
      "id": "114423",
      "postDate": "04/10/2016 17:13:32",
      "content": "<p>I see. Any special reason for doing that?</p>",
      "rawMarkdown": "I see. Any special reason for doing that?",
      "votes": null
    },
    {
      "id": "114427",
      "postDate": "04/10/2016 18:17:22",
      "content": "<p>[quote=Andre lopes;114423]\nI see. Any special reason for doing that?\n[/quote]</p>\n\n<p>The reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.</p>",
      "rawMarkdown": "[quote=Andre lopes;114423]\r\nI see. Any special reason for doing that?\r\n[/quote]\r\n\r\nThe reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.",
      "votes": null
    },
    {
      "id": "114433",
      "postDate": "04/10/2016 19:34:46",
      "content": "<p>[quote=ZFTurbo;114427]</p>\n\n<p>[quote=Andre lopes;114423]\nI see. Any special reason for doing that?\n[/quote]</p>\n\n<p>The reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.</p>\n\n<p>[/quote]</p>\n\n<p>There is a <em>huge</em> difference and improvement in accuracy of CV metrics when splitting validation by driver. It's still far from perfect, but with ~3 drivers split out for CV (using 8-fold CV by driver) I'm getting within 20% of my LB score (1.5 CV compared to 1.3 LB). Whilst with a random split of e.g. 25% of images - I get CV loss of 0.1 and accuracy of 98%+ for the same meta-params, which is nonsense. </p>\n\n<p>As far as I can see a random split for CV provides next to no useful information - you may as well train with everything and just fit to the leaderboard score. Although still a bad idea, I don't think that would be a disaster here - i.e. I think it will be hard to badly overfit the public LB - except for the limited number of times you can test per day.</p>",
      "rawMarkdown": "[quote=ZFTurbo;114427]\r\n\r\n[quote=Andre lopes;114423]\r\nI see. Any special reason for doing that?\r\n[/quote]\r\n\r\nThe reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.\r\n\r\n[/quote]\r\n\r\nThere is a *huge* difference and improvement in accuracy of CV metrics when splitting validation by driver. It's still far from perfect, but with ~3 drivers split out for CV (using 8-fold CV by driver) I'm getting within 20% of my LB score (1.5 CV compared to 1.3 LB). Whilst with a random split of e.g. 25% of images - I get CV loss of 0.1 and accuracy of 98%+ for the same meta-params, which is nonsense. \r\n\r\nAs far as I can see a random split for CV provides next to no useful information - you may as well train with everything and just fit to the leaderboard score. Although still a bad idea, I don't think that would be a disaster here - i.e. I think it will be hard to badly overfit the public LB - except for the limited number of times you can test per day.",
      "votes": null
    },
    {
      "id": "114518",
      "postDate": "04/11/2016 15:38:18",
      "content": "<p>Hi, I am new to Keras, I like the idea to manually select the validation set. However, I receive similar results when I don't use the for loop and just set a constant validation set using the last 20% of the data:</p>\n\n<pre><code>    model.fit(X_train_in, Y_train, batch_size=32, nb_epoch=20, show_accuracy=True, verbose=1, validation_split=0.2)\n</code></pre>\n\n<p>As I am not very familiar with Keras, are the weights actually updated within the for loop, or is it not rather the case that they are reset every time and it is just using the weights from the last iteration?</p>\n\n<p>Thanks for the clarification.</p>",
      "rawMarkdown": "Hi, I am new to Keras, I like the idea to manually select the validation set. However, I receive similar results when I don't use the for loop and just set a constant validation set using the last 20% of the data:\r\n\r\n        model.fit(X_train_in, Y_train, batch_size=32, nb_epoch=20, show_accuracy=True, verbose=1, validation_split=0.2)\r\n\r\nAs I am not very familiar with Keras, are the weights actually updated within the for loop, or is it not rather the case that they are reset every time and it is just using the weights from the last iteration?\r\n\r\nThanks for the clarification.",
      "votes": null
    },
    {
      "id": "115765",
      "postDate": "04/19/2016 23:34:33",
      "content": "<p>I think there is a problem in using <code>np.reshape()</code> versus <code>np.transpose()</code> when loading the images. The current code is</p>\n\n<pre><code>train_data = train_data.reshape(train_data.shape[0], color_type, img_rows, img_cols)\n</code></pre>\n\n<p>I think it should be</p>\n\n<pre><code>train_data = train_data.transpose((0, 3, 1, 2))\n</code></pre>\n\n<p>Reshape scrambles the indices with the results shown in the attached image.</p>\n\n<p>The code I used to create those images is in the <code>load_train</code> method:</p>\n\n<pre><code>img_raw = train_data[100, ...]  # shape = (color_type, img_rows, img_cols)\nimg = np.zeros((img_rows, img_cols, color_type), dtype=np.uint8)\nimg[...,0] = img_raw[0]\nimg[...,1] = img_raw[1]\nimg[...,2] = img_raw[2]\nimg = cv2.resize(img, (640, 480))\ncv2.imwrite('reshape.jpg', img)\ncv2.imshow('reshape', img)\ncv2.waitKey(0)\ncv2.destroyAllWindows()\n</code></pre>",
      "rawMarkdown": "I think there is a problem in using `np.reshape()` versus `np.transpose()` when loading the images. The current code is\r\n\r\n    train_data = train_data.reshape(train_data.shape[0], color_type, img_rows, img_cols)\r\n\r\nI think it should be\r\n\r\n    train_data = train_data.transpose((0, 3, 1, 2))\r\n\r\nReshape scrambles the indices with the results shown in the attached image.\r\n\r\nThe code I used to create those images is in the `load_train` method:\r\n\r\n    img_raw = train_data[100, ...]  # shape = (color_type, img_rows, img_cols)\r\n    img = np.zeros((img_rows, img_cols, color_type), dtype=np.uint8)\r\n    img[...,0] = img_raw[0]\r\n    img[...,1] = img_raw[1]\r\n    img[...,2] = img_raw[2]\r\n    img = cv2.resize(img, (640, 480))\r\n    cv2.imwrite('reshape.jpg', img)\r\n    cv2.imshow('reshape', img)\r\n    cv2.waitKey(0)\r\n    cv2.destroyAllWindows()",
      "votes": null
    },
    {
      "id": "115810",
      "postDate": "04/20/2016 08:47:55",
      "content": "<p>gauss256, this fix only needed for colored images. It won't work for gray_scale images (I checked image looks correct). So my current fix is the following:</p>\n\n<pre><code>if color_type == 1:\n    train_data = train_data.reshape(train_data.shape[0], 1, img_rows, img_cols)\nelse:\n    train_data = train_data.transpose((0, 3, 1, 2))\n</code></pre>\n\n<p>And the same for test:</p>\n\n<pre><code>if color_type == 1:\n    test_data = test_data.reshape(test_data.shape[0], 1, img_rows, img_cols)\nelse:\n    test_data = test_data.transpose((0, 3, 1, 2))\n</code></pre>\n\n<p>I made the changes on GITHUB as well.</p>",
      "rawMarkdown": "gauss256, this fix only needed for colored images. It won't work for gray_scale images (I checked image looks correct). So my current fix is the following:\r\n\r\n    if color_type == 1:\r\n        train_data = train_data.reshape(train_data.shape[0], 1, img_rows, img_cols)\r\n    else:\r\n        train_data = train_data.transpose((0, 3, 1, 2))\r\n\r\nAnd the same for test:\r\n\r\n    if color_type == 1:\r\n        test_data = test_data.reshape(test_data.shape[0], 1, img_rows, img_cols)\r\n    else:\r\n        test_data = test_data.transpose((0, 3, 1, 2))\r\n\r\nI made the changes on GITHUB as well.",
      "votes": null
    },
    {
      "id": "115814",
      "postDate": "04/20/2016 09:29:32",
      "content": "<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>",
      "rawMarkdown": "BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.",
      "votes": null
    },
    {
      "id": "115838",
      "postDate": "04/20/2016 12:36:32",
      "content": "<p>[quote=ZFTurbo;115814]</p>\n\n<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>\n\n<p>[/quote]</p>\n\n<p>I am seeing this sort of thing too, especially with certain drivers in the CV set. However, I think it is expected behaviour, even though it is not wanted. The model is predicting incorrect class <em>very confidently</em> on the CV set for some proportion of the images. </p>\n\n<p>So I don't think this is a bug, instead the challenge is to find parameters (or maybe different overall approaches) which are less vulnerable to the effect.</p>",
      "rawMarkdown": "[quote=ZFTurbo;115814]\r\n\r\nBTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.\r\n\r\n[/quote]\r\n\r\nI am seeing this sort of thing too, especially with certain drivers in the CV set. However, I think it is expected behaviour, even though it is not wanted. The model is predicting incorrect class *very confidently* on the CV set for some proportion of the images. \r\n\r\nSo I don't think this is a bug, instead the challenge is to find parameters (or maybe different overall approaches) which are less vulnerable to the effect.",
      "votes": null
    },
    {
      "id": "115862",
      "postDate": "04/20/2016 14:34:37",
      "content": "<p>[quote=ZFTurbo;115814]</p>\n\n<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>\n\n<p>[/quote]\nAre you sure you're not overfitting? Sure looks like it, but I didn't check the code.</p>",
      "rawMarkdown": "[quote=ZFTurbo;115814]\r\n\r\nBTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.\r\n\r\n[/quote]\r\nAre you sure you're not overfitting? Sure looks like it, but I didn't check the code.",
      "votes": null
    },
    {
      "id": "115893",
      "postDate": "04/20/2016 16:24:52",
      "content": "<p>[quote=ZFTurbo;115814]</p>\n\n<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>\n\n<p>[/quote]\nI see the same thing and similar results for the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20129/cloud-gpu-starter-project\">Fomoro</a> model. These appear to be symptoms of overfitting, but adding L2 regularization and other tweaks have not made much of a difference for me.</p>\n\n<p>My other theory is that an image size of (24,32) is just too small. As a human looking at that size of image I don't think I could classify them very well at all. But I've tried (48,64) and it wasn't much better. Larger images start bogging down the computer pretty quickly.</p>",
      "rawMarkdown": "[quote=ZFTurbo;115814]\r\n\r\nBTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.\r\n\r\n[/quote]\r\nI see the same thing and similar results for the [Fomoro][1] model. These appear to be symptoms of overfitting, but adding L2 regularization and other tweaks have not made much of a difference for me.\r\n\r\nMy other theory is that an image size of (24,32) is just too small. As a human looking at that size of image I don't think I could classify them very well at all. But I've tried (48,64) and it wasn't much better. Larger images start bogging down the computer pretty quickly.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20129/cloud-gpu-starter-project",
      "votes": null
    },
    {
      "id": "116980",
      "postDate": "04/26/2016 19:39:15",
      "content": "<p>[quote=ZFTurbo;113939]</p>\n\n<p>Here is simple solution using CNN to start from:</p>\n\n<p>Ver. 1: <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py</a></p>\n\n<p>Ver. 2 (<strong>UPD 07.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Ver. 3 (<strong>UPD 09.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>[/quote]</p>\n\n<p>Hi, I am using your script as a starter. I was trying to play around with the model. I wanted to ask you 2 things:\n1) I tried adding another Convolutional layer, but I was surprised that log loss on test set did not improve, in fact, it actually got worse. Do you know why? I am attaching my script for you to have a look(turbo_script_v5.py)</p>\n\n<p>2)I tried using a pre-trained model VGG_16. I integrated it with your code, but there seems to be some error regarding the shape of input and output. I just call this model instead of create_model_v1 in your scripts run_keras_cv_drivers.py. Any idea what am I doing wrong here? </p>\n\n<p>Here's the code snippet:</p>\n\n<pre><code>def VCG_16(img_rows, img_cols, color_type=1,weights_path=None):\n        model = Sequential()\n        model.add(ZeroPadding2D((1,1),input_shape=(color_type,img_rows,img_cols)))\n        model.add(Convolution2D(64, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(64, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(Flatten())\n\n        model.add(Dense(4096, activation='relu'))\n        model.add(Dropout(0.5))\n        model.add(Dense(4096, activation='relu'))\n        model.add(Dropout(0.5))\n        model.add(Dense(10, activation='softmax'))\n\n        if weights_path:\n            model.load_weights(weights_path)\n\n        sgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)\n        model.compile(optimizer=sgd, loss='categorical_crossentropy')\n\n        return model\n</code></pre>",
      "rawMarkdown": "[quote=ZFTurbo;113939]\r\n\r\nHere is simple solution using CNN to start from:\r\n\r\nVer. 1: https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\r\n\r\nVer. 2 (**UPD 07.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nVer. 3 (**UPD 09.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\n\r\n[/quote]\r\n\r\nHi, I am using your script as a starter. I was trying to play around with the model. I wanted to ask you 2 things:\r\n1) I tried adding another Convolutional layer, but I was surprised that log loss on test set did not improve, in fact, it actually got worse. Do you know why? I am attaching my script for you to have a look(turbo_script_v5.py)\r\n\r\n2)I tried using a pre-trained model VGG_16. I integrated it with your code, but there seems to be some error regarding the shape of input and output. I just call this model instead of create_model_v1 in your scripts run_keras_cv_drivers.py. Any idea what am I doing wrong here? \r\n\r\nHere's the code snippet:\r\n\r\n    def VCG_16(img_rows, img_cols, color_type=1,weights_path=None):\r\n    \t\tmodel = Sequential()\r\n    \t\tmodel.add(ZeroPadding2D((1,1),input_shape=(color_type,img_rows,img_cols)))\r\n    \t\tmodel.add(Convolution2D(64, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(64, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(128, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(128, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(256, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(256, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(256, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(Flatten())\r\n    \r\n    \t\tmodel.add(Dense(4096, activation='relu'))\r\n    \t\tmodel.add(Dropout(0.5))\r\n    \t\tmodel.add(Dense(4096, activation='relu'))\r\n    \t\tmodel.add(Dropout(0.5))\r\n    \t\tmodel.add(Dense(10, activation='softmax'))\r\n    \r\n    \t\tif weights_path:\r\n    \t\t    model.load_weights(weights_path)\r\n    \r\n    \t\tsgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)\r\n    \t\tmodel.compile(optimizer=sgd, loss='categorical_crossentropy')\r\n    \t\r\n    \t\treturn model",
      "votes": null
    },
    {
      "id": "116981",
      "postDate": "04/26/2016 19:47:08",
      "content": "<p>Try this from one of my experiments: </p>\n\n<pre><code>  def create_model_v2(img_rows, img_cols, color_type=1):\n    model = Sequential()\n\n    # 1 block 48x64\n    model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    # model.add(Dropout(0.25))\n\n    # 2 block 24x32\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 3 block 12x16\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 4 block 6x8\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    model.add(Flatten())\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10))\n    model.add(Activation('softmax'))\n\n    sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\n    model.compile(loss='categorical_crossentropy', optimizer=sgd)\n    return model\n</code></pre>\n\n<p>or this:</p>\n\n<pre><code>def create_model_v4(img_rows, img_cols, color_type=1):\n    model = Sequential()\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Flatten())\n    model.add(Dense(128, activation='sigmoid', init='he_normal'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10, activation='softmax', init='he_normal'))\n    model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\n    return model\n</code></pre>",
      "rawMarkdown": "Try this from one of my experiments: \r\n \r\n\r\n      def create_model_v2(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n    \r\n        # 1 block 48x64\r\n        model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n        # model.add(Dropout(0.25))\r\n    \r\n        # 2 block 24x32\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 3 block 12x16\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 4 block 6x8\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        model.add(Flatten())\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10))\r\n        model.add(Activation('softmax'))\r\n    \r\n        sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\r\n        model.compile(loss='categorical_crossentropy', optimizer=sgd)\r\n        return model\r\n\r\nor this:\r\n\r\n    def create_model_v4(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Flatten())\r\n        model.add(Dense(128, activation='sigmoid', init='he_normal'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10, activation='softmax', init='he_normal'))\r\n        model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\r\n        return model",
      "votes": null
    },
    {
      "id": "116985",
      "postDate": "04/26/2016 19:59:20",
      "content": "<p>[quote=ZFTurbo;116981]</p>\n\n<p>Try this from one of my experiments: </p>\n\n<pre><code>  def create_model_v2(img_rows, img_cols, color_type=1):\n    model = Sequential()\n\n    # 1 block 48x64\n    model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    # model.add(Dropout(0.25))\n\n    # 2 block 24x32\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 3 block 12x16\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 4 block 6x8\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    model.add(Flatten())\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10))\n    model.add(Activation('softmax'))\n\n    sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\n    model.compile(loss='categorical_crossentropy', optimizer=sgd)\n    return model\n</code></pre>\n\n<p>or this:</p>\n\n<pre><code>def create_model_v4(img_rows, img_cols, color_type=1):\n    model = Sequential()\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Flatten())\n    model.add(Dense(128, activation='sigmoid', init='he_normal'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10, activation='softmax', init='he_normal'))\n    model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\n    return model\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>Thanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? </p>",
      "rawMarkdown": "[quote=ZFTurbo;116981]\r\n\r\nTry this from one of my experiments: \r\n \r\n\r\n      def create_model_v2(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n    \r\n        # 1 block 48x64\r\n        model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n        # model.add(Dropout(0.25))\r\n    \r\n        # 2 block 24x32\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 3 block 12x16\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 4 block 6x8\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        model.add(Flatten())\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10))\r\n        model.add(Activation('softmax'))\r\n    \r\n        sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\r\n        model.compile(loss='categorical_crossentropy', optimizer=sgd)\r\n        return model\r\n\r\nor this:\r\n\r\n    def create_model_v4(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Flatten())\r\n        model.add(Dense(128, activation='sigmoid', init='he_normal'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10, activation='softmax', init='he_normal'))\r\n        model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\r\n        return model\r\n\r\n[/quote]\r\n\r\nThanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights?",
      "votes": null
    },
    {
      "id": "116988",
      "postDate": "04/26/2016 20:11:11",
      "content": "<p>[quote=V.AbhijayArora;116985]</p>\n\n<p>Thanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? </p>\n\n<p>[/quote]</p>\n\n<p>Not yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.</p>",
      "rawMarkdown": "[quote=V.AbhijayArora;116985]\r\n\r\nThanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? \r\n\r\n[/quote]\r\n\r\nNot yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.",
      "votes": null
    },
    {
      "id": "117923",
      "postDate": "05/01/2016 19:20:50",
      "content": "<p>[quote=ZFTurbo;116988]</p>\n\n<p>[quote=V.AbhijayArora;116985]</p>\n\n<p>Thanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? </p>\n\n<p>[/quote]</p>\n\n<p>Not yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.</p>\n\n<p>[/quote]\nWhat image size works the best? I have tried (48 x 48 x1) or (64 x 48 x 1) or (32 x 24 x 1). Have you tried increasing the number of epochs?I'm not able to get a better score than 1.22 on LB. Any tips?</p>",
      "rawMarkdown": "[quote=ZFTurbo;116988]\r\n\r\n[quote=V.AbhijayArora;116985]\r\n\r\nThanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? \r\n\r\n[/quote]\r\n\r\nNot yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.\r\n\r\n[/quote]\r\nWhat image size works the best? I have tried (48 x 48 x1) or (64 x 48 x 1) or (32 x 24 x 1). Have you tried increasing the number of epochs?I'm not able to get a better score than 1.22 on LB. Any tips?",
      "votes": null
    },
    {
      "id": "117983",
      "postDate": "05/02/2016 08:13:34",
      "content": "<p><strong>Abhijay Arora</strong>, it's just the matter of experiments. In current code 32x24 grayscale is better than 64x48 or RGB color mode. I'll plan to post my current code a little bit later, which allows to obtain around 0.85 on LB.</p>",
      "rawMarkdown": "**Abhijay Arora**, it's just the matter of experiments. In current code 32x24 grayscale is better than 64x48 or RGB color mode. I'll plan to post my current code a little bit later, which allows to obtain around 0.85 on LB.",
      "votes": null
    },
    {
      "id": "118264",
      "postDate": "05/03/2016 04:19:13",
      "content": "<p>[quote=ZFTurbo;116981]</p>\n\n<p>Try this from one of my experiments: </p>\n\n<pre><code>  def create_model_v2(img_rows, img_cols, color_type=1):\n    model = Sequential()\n\n    # 1 block 48x64\n    model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    # model.add(Dropout(0.25))\n\n    # 2 block 24x32\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 3 block 12x16\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 4 block 6x8\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    model.add(Flatten())\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10))\n    model.add(Activation('softmax'))\n\n    sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\n    model.compile(loss='categorical_crossentropy', optimizer=sgd)\n    return model\n</code></pre>\n\n<p>or this:</p>\n\n<pre><code>def create_model_v4(img_rows, img_cols, color_type=1):\n    model = Sequential()\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Flatten())\n    model.add(Dense(128, activation='sigmoid', init='he_normal'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10, activation='softmax', init='he_normal'))\n    model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\n    return model\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>What score did these achieve on LB? The first one took 12 hours to train on my laptop.</p>",
      "rawMarkdown": "[quote=ZFTurbo;116981]\r\n\r\nTry this from one of my experiments: \r\n \r\n\r\n      def create_model_v2(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n    \r\n        # 1 block 48x64\r\n        model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n        # model.add(Dropout(0.25))\r\n    \r\n        # 2 block 24x32\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 3 block 12x16\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 4 block 6x8\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        model.add(Flatten())\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10))\r\n        model.add(Activation('softmax'))\r\n    \r\n        sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\r\n        model.compile(loss='categorical_crossentropy', optimizer=sgd)\r\n        return model\r\n\r\nor this:\r\n\r\n    def create_model_v4(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Flatten())\r\n        model.add(Dense(128, activation='sigmoid', init='he_normal'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10, activation='softmax', init='he_normal'))\r\n        model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\r\n        return model\r\n\r\n[/quote]\r\n\r\nWhat score did these achieve on LB? The first one took 12 hours to train on my laptop.",
      "votes": null
    },
    {
      "id": "118398",
      "postDate": "05/03/2016 14:20:15",
      "content": "<p>I posted latest version of my Keras code (1st topic updated):\n<a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py</a></p>\n\n<p>I used some ideas from this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne</a></p>\n\n<ol>\n<li>Code randomly rotate images +-10 degrees</li>\n<li>Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.</li>\n<li>Code uses 64x64 pixel grayscale images</li>\n<li>CNN is actually simple enough to be run on ordinary computer in reasonable time</li>\n<li>I added some useful callback functions:\nEarlyStopping - to stop early after loss stop decreasing\nModelCheckpoint - save best weights and restore them for minimum loss after &quot;fit&quot; ends. Some kind of XGBoost's best_ntree_limit</li>\n</ol>\n\n<p>Notes:</p>\n\n<ol>\n<li>Crossfold score is actually much lower than leaderboard one </li>\n<li>In most cases best loss achieved right after first epoch </li>\n<li><p>I feel like this line: </p>\n\n<p>train_data = train_data.reshape(train_data.shape[0], 1,\n    img_rows, img_cols) </p></li>\n</ol>\n\n<p>should be replaced with some other function. It would be good if someone propose best replacement.</p>\n\n<ol start=\"4\">\n<li>&quot;batch_size&quot; quite strongly affects learning process </li>\n<li>It seems KFold = 26 is best for this problem.</li>\n</ol>\n\n<p>Running &quot;as is&quot; from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.</p>",
      "rawMarkdown": "I posted latest version of my Keras code (1st topic updated):\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\r\n\r\nI used some ideas from this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\r\n\r\n1. Code randomly rotate images +-10 degrees\r\n2. Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.\r\n3. Code uses 64x64 pixel grayscale images\r\n4. CNN is actually simple enough to be run on ordinary computer in reasonable time\r\n5. I added some useful callback functions:\r\nEarlyStopping - to stop early after loss stop decreasing\r\nModelCheckpoint - save best weights and restore them for minimum loss after \"fit\" ends. Some kind of XGBoost's best_ntree_limit\r\n\r\nNotes:\r\n\r\n 1. Crossfold score is actually much lower than leaderboard one \r\n 2. In most cases best loss achieved right after first epoch \r\n 3. I feel like this line: \r\n\r\n    train_data = train_data.reshape(train_data.shape[0], 1,\r\n        img_rows, img_cols) \r\n\r\nshould be replaced with some other function. It would be good if someone propose best replacement.\r\n\r\n  4. \"batch_size\" quite strongly affects learning process \r\n  5. It seems KFold = 26 is best for this problem.\r\n\r\nRunning \"as is\" from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.",
      "votes": null
    },
    {
      "id": "118794",
      "postDate": "05/05/2016 11:13:17",
      "content": "<p>[quote=ZFTurbo;118398]</p>\n\n<p>I posted latest version of my Keras code (1st topic updated):\n<a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py</a></p>\n\n<p>I used some ideas from this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne</a></p>\n\n<ol>\n<li>Code randomly rotate images +-10 degrees</li>\n<li>Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.</li>\n<li>Code uses 64x64 pixel grayscale images</li>\n<li>CNN is actually simple enough to be run on ordinary computer in reasonable time</li>\n<li>I added some useful callback functions:\nEarlyStopping - to stop early after loss stop decreasing\nModelCheckpoint - save best weights and restore them for minimum loss after &quot;fit&quot; ends. Some kind of XGBoost's best_ntree_limit</li>\n</ol>\n\n<p>Notes:</p>\n\n<ol>\n<li>Crossfold score is actually much lower than leaderboard one </li>\n<li>In most cases best loss achieved right after first epoch </li>\n<li><p>I feel like this line: </p>\n\n<p>train_data = train_data.reshape(train_data.shape[0], 1,\n    img_rows, img_cols) </p></li>\n</ol>\n\n<p>should be replaced with some other function. It would be good if someone propose best replacement.</p>\n\n<ol start=\"4\">\n<li>&quot;batch_size&quot; quite strongly affects learning process </li>\n<li>It seems KFold = 26 is best for this problem.</li>\n</ol>\n\n<p>Running &quot;as is&quot; from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.</p>\n\n<p>[/quote]\n1) Did you try increasing the number of convolutional layers? I added a 32 x 32 convolutional layer and a 64 X 64 convolutional layer only to get worse results(LB: 0.88)\n2) Why don't you try adding more dense layers? </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "[quote=ZFTurbo;118398]\r\n\r\nI posted latest version of my Keras code (1st topic updated):\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\r\n\r\nI used some ideas from this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\r\n\r\n1. Code randomly rotate images +-10 degrees\r\n2. Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.\r\n3. Code uses 64x64 pixel grayscale images\r\n4. CNN is actually simple enough to be run on ordinary computer in reasonable time\r\n5. I added some useful callback functions:\r\nEarlyStopping - to stop early after loss stop decreasing\r\nModelCheckpoint - save best weights and restore them for minimum loss after \"fit\" ends. Some kind of XGBoost's best_ntree_limit\r\n\r\nNotes:\r\n\r\n 1. Crossfold score is actually much lower than leaderboard one \r\n 2. In most cases best loss achieved right after first epoch \r\n 3. I feel like this line: \r\n\r\n    train_data = train_data.reshape(train_data.shape[0], 1,\r\n        img_rows, img_cols) \r\n\r\nshould be replaced with some other function. It would be good if someone propose best replacement.\r\n\r\n  4. \"batch_size\" quite strongly affects learning process \r\n  5. It seems KFold = 26 is best for this problem.\r\n\r\nRunning \"as is\" from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.\r\n\r\n[/quote]\r\n1) Did you try increasing the number of convolutional layers? I added a 32 x 32 convolutional layer and a 64 X 64 convolutional layer only to get worse results(LB: 0.88)\r\n2) Why don't you try adding more dense layers? \r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "129029",
      "postDate": "07/26/2016 00:47:07",
      "content": "<p>Thanks for sharing! It really helped me to get started.</p>",
      "rawMarkdown": "Thanks for sharing! It really helped me to get started.",
      "votes": null
    },
    {
      "id": "129338",
      "postDate": "07/28/2016 22:01:21",
      "content": "<p>I've added my code to run pretrained VGG16 Neural Net:\n<a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py</a></p>\n\n<p>This code allows to get <strong>0.20646</strong> on LB.</p>\n\n<p>Requirements:</p>\n\n<ol>\n<li>16 GB of RAM (with some swap at preprocessing peaks). Code splits &quot;test&quot; in 5 chunks, so it allowed to reduce memory usage.</li>\n<li>Powerful NVIDIA GPU. It takes around a day on 980Ti 6GB.</li>\n<li><strong>Important</strong>: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.</li>\n<li>Weights: <a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3</a></li>\n</ol>\n\n<p>Notes:</p>\n\n<ol>\n<li>Code doesn't use data augmentation</li>\n<li>Code uses totally random test split (doesn't use split by drivers). So don't look at local validation loss, it will be around ~0.01 at the end.</li>\n</ol>",
      "rawMarkdown": "I've added my code to run pretrained VGG16 Neural Net:\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py\r\n\r\nThis code allows to get **0.20646** on LB.\r\n\r\nRequirements:\r\n\r\n 1. 16 GB of RAM (with some swap at preprocessing peaks). Code splits \"test\" in 5 chunks, so it allowed to reduce memory usage.\r\n 2. Powerful NVIDIA GPU. It takes around a day on 980Ti 6GB.\r\n 3. **Important**: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.\r\n 4. Weights: https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n\r\nNotes:\r\n\r\n 1. Code doesn't use data augmentation\r\n 2. Code uses totally random test split (doesn't use split by drivers). So don't look at local validation loss, it will be around ~0.01 at the end.",
      "votes": null
    },
    {
      "id": "129350",
      "postDate": "07/29/2016 01:17:02",
      "content": "<p>@ZFTurbo :  just wanted to confirm that during training of the model the data is split based on drivers ( was looking at your code for function  run_cross_validation_create_models) </p>",
      "rawMarkdown": "ZFTurbo :  just wanted to confirm that during training of the model the data is split based on drivers ( was looking at your code for function  run_cross_validation_create_models)",
      "votes": null
    },
    {
      "id": "129362",
      "postDate": "07/29/2016 06:59:04",
      "content": "<p><strong>Hetal Chandaria</strong> it initially splitted by drivers, but later shuffled (look for the following code):</p>\n\n<pre><code># Shuffle experiment START !!!\nperm = permutation(len(train_target))\ntrain_data = train_data[perm]\ntrain_target = train_target[perm]\n# Shuffle experiment END !!!\n</code></pre>",
      "rawMarkdown": "**Hetal Chandaria** it initially splitted by drivers, but later shuffled (look for the following code):\r\n\r\n    # Shuffle experiment START !!!\r\n    perm = permutation(len(train_target))\r\n    train_data = train_data[perm]\r\n    train_target = train_target[perm]\r\n    # Shuffle experiment END !!!",
      "votes": null
    },
    {
      "id": "129369",
      "postDate": "07/29/2016 08:28:49",
      "content": "<p>Thanks for sharing!</p>\n\n<p>So the key is version of keras? ToT</p>",
      "rawMarkdown": "Thanks for sharing!\r\n\r\nSo the key is version of keras? ToT",
      "votes": null
    },
    {
      "id": "129370",
      "postDate": "07/29/2016 08:32:34",
      "content": "<p><strong>Ferris</strong>, yes. At least this code works much better on outdated version.</p>",
      "rawMarkdown": "**Ferris**, yes. At least this code works much better on outdated version.",
      "votes": null
    },
    {
      "id": "129483",
      "postDate": "07/30/2016 09:28:26",
      "content": "<p>Thanks for the efforts, very educative!</p>\n\n<p>One question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day). </p>",
      "rawMarkdown": "Thanks for the efforts, very educative!\r\n\r\nOne question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day).",
      "votes": null
    },
    {
      "id": "129484",
      "postDate": "07/30/2016 09:33:08",
      "content": "<p>[quote=liviu;129483]</p>\n\n<p>Thanks for the efforts, very educative!</p>\n\n<p>One question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day). </p>\n\n<p>[/quote]</p>\n\n<p>I suspect not. I think you'll find that anything scoring 0.5 or better would take literally weeks to train on a CPU-based setup, and you would need to have chosen the meta-params correctly. I got down to 0.89 on a CPU rig after a few iterations, and that took 2.5 days training time. </p>",
      "rawMarkdown": "[quote=liviu;129483]\r\n\r\nThanks for the efforts, very educative!\r\n\r\nOne question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day). \r\n\r\n[/quote]\r\n\r\nI suspect not. I think you'll find that anything scoring 0.5 or better would take literally weeks to train on a CPU-based setup, and you would need to have chosen the meta-params correctly. I got down to 0.89 on a CPU rig after a few iterations, and that took 2.5 days training time.",
      "votes": null
    },
    {
      "id": "129485",
      "postDate": "07/30/2016 09:48:27",
      "content": "<p>Not what I was hoping to hear, but it gives me the green light to move on :P. Thanks!</p>",
      "rawMarkdown": "Not what I was hoping to hear, but it gives me the green light to move on :P. Thanks!",
      "votes": null
    },
    {
      "id": "129519",
      "postDate": "07/30/2016 16:18:09",
      "content": "<p>Hi, @ZFTurbo,\nThank you for your nice code sharing. Is there any reason for not split by driver this time?</p>",
      "rawMarkdown": "Hi, @ZFTurbo,\r\nThank you for your nice code sharing. Is there any reason for not split by driver this time?",
      "votes": null
    },
    {
      "id": "129520",
      "postDate": "07/30/2016 16:37:38",
      "content": "<p><strong>Li Li</strong>, for some reason score while splitting by drivers was worse.</p>",
      "rawMarkdown": "**Li Li**, for some reason score while splitting by drivers was worse.",
      "votes": null
    },
    {
      "id": "129531",
      "postDate": "07/30/2016 18:25:23",
      "content": "<p>[quote=ZFTurbo;129338]\n 3. <strong>Important</strong>: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.\n[/quote]</p>\n\n<p>Can you post score on latest Keras version you tried? I tried on 1.0.6 and got 0.23 with 8 folds with some modifications (GAP instead of Dense Layers for memory optimization). I will try to debug your code with new Keras version, but wander if you have ideas why its different and what score you got.</p>",
      "rawMarkdown": "[quote=ZFTurbo;129338]\r\n 3. **Important**: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.\r\n[/quote]\r\n\r\nCan you post score on latest Keras version you tried? I tried on 1.0.6 and got 0.23 with 8 folds with some modifications (GAP instead of Dense Layers for memory optimization). I will try to debug your code with new Keras version, but wander if you have ideas why its different and what score you got.",
      "votes": null
    },
    {
      "id": "129534",
      "postDate": "07/30/2016 19:02:40",
      "content": "<p><strong>kyv</strong>, It was run by my teammates. And they had problems with my code. I see 0.29 and 0.27 in list (it was calculated on 1.0.* version).</p>\n\n<p>I don't have ideas, except that learning process is somehow different in 0.2.0 and 1.0.*. I probably rerun this code on latest version after contest ends to check the difference myself.</p>\n\n<p>I also notice difference in loss in Nerve contest. When I debug Kernel at local machine, loss was much different from version on Kaggle servers.</p>",
      "rawMarkdown": "**kyv**, It was run by my teammates. And they had problems with my code. I see 0.29 and 0.27 in list (it was calculated on 1.0.* version).\r\n\r\nI don't have ideas, except that learning process is somehow different in 0.2.0 and 1.0.*. I probably rerun this code on latest version after contest ends to check the difference myself.\r\n\r\nI also notice difference in loss in Nerve contest. When I debug Kernel at local machine, loss was much different from version on Kaggle servers.",
      "votes": null
    },
    {
      "id": "129642",
      "postDate": "08/01/2016 08:21:34",
      "content": "<p>Could be because of different cudnn or theano versions. You also have to clean theano cache once in a while - a for sure if you change something like cuda drivers, cudnn or theano version.</p>",
      "rawMarkdown": "Could be because of different cudnn or theano versions. You also have to clean theano cache once in a while - a for sure if you change something like cuda drivers, cudnn or theano version.",
      "votes": null
    },
    {
      "id": "129644",
      "postDate": "08/01/2016 08:46:00",
      "content": "<p>@kyv I'm sorry for not knowing. What is &quot;GAP instead of Dense Layers for memory optimization&quot;?</p>",
      "rawMarkdown": "kyv I'm sorry for not knowing. What is \"GAP instead of Dense Layers for memory optimization\"?",
      "votes": null
    },
    {
      "id": "129737",
      "postDate": "08/02/2016 01:13:48",
      "content": "<p>[quote=Keiku;129644]</p>\n\n<p>@kyv I'm sorry for not knowing. What is &quot;GAP instead of Dense Layers for memory optimization&quot;?</p>\n\n<p>[/quote]\nI think that's global average pooling layer. You can find this concept in paper:Network In Network</p>",
      "rawMarkdown": "[quote=Keiku;129644]\r\n\r\n@kyv I'm sorry for not knowing. What is \"GAP instead of Dense Layers for memory optimization\"?\r\n\r\n[/quote]\r\nI think that's global average pooling layer. You can find this concept in paper:Network In Network",
      "votes": null
    },
    {
      "id": "129740",
      "postDate": "08/02/2016 01:27:15",
      "content": "<p>@Tian Zhou Thank you for telling me about.</p>",
      "rawMarkdown": "Tian Zhou Thank you for telling me about.",
      "votes": null
    },
    {
      "id": "129759",
      "postDate": "08/02/2016 05:35:28",
      "content": "<p>[quote=Keiku;129644]</p>\n\n<p>@kyv I'm sorry for not knowing. What is &quot;GAP instead of Dense Layers for memory optimization&quot;?</p>\n\n<p>[/quote]\nYes its AveragePooling layer with stride equal to image size and channels equal to number of classes followed by softmax .\nBy doing this you reduce model size by 80% (as most parameters are in dense layers) and usually you don't sacrifice model accuracy. New models often uses this approach. </p>",
      "rawMarkdown": "[quote=Keiku;129644]\r\n\r\n@kyv I'm sorry for not knowing. What is \"GAP instead of Dense Layers for memory optimization\"?\r\n\r\n[/quote]\r\nYes its AveragePooling layer with stride equal to image size and channels equal to number of classes followed by softmax .\r\nBy doing this you reduce model size by 80% (as most parameters are in dense layers) and usually you don't sacrifice model accuracy. New models often uses this approach.",
      "votes": null
    },
    {
      "id": "129763",
      "postDate": "08/02/2016 06:07:52",
      "content": "<p>@kyv Thank you for your advice. I will try with new Keras version for future reference, too.</p>",
      "rawMarkdown": "kyv Thank you for your advice. I will try with new Keras version for future reference, too.",
      "votes": null
    },
    {
      "id": "129766",
      "postDate": "08/02/2016 07:12:52",
      "content": "<p><strong>kyv</strong>, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )</p>",
      "rawMarkdown": "**kyv**, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )",
      "votes": null
    },
    {
      "id": "129770",
      "postDate": "08/02/2016 08:05:30",
      "content": "<p>[quote=ZFTurbo;129766]</p>\n\n<p><strong>kyv</strong>, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )\n[/quote]</p>\n\n<p>I replaced last dense layers with</p>\n\n<pre><code> for _ in range(conv_layers):\n    model.add(Convolution2D(nb_filter, nb_cols, nb_cols, &quot;relu&quot;)) \nmodel.add(AveragePooling2D((7, 7)))\nmodel.add(Flatten())\nmodel.add(Dropout(dropout))\nmodel.add(Dense(10, W_regularizer=REG, b_regularizer=REG))\nmodel.add(Activation(&quot;softmax&quot;))\n</code></pre>\n\n<p>conv_layers=3 and nb_filter=50 to 100.\nIt was comparable with original vgg16 on fine-tuning convergence speed and eventually I stopped using dense layers as it was limiting me in batch size. </p>",
      "rawMarkdown": "[quote=ZFTurbo;129766]\r\n\r\n**kyv**, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )\r\n[/quote]\r\n\r\nI replaced last dense layers with\r\n\r\n     for _ in range(conv_layers):\r\n        model.add(Convolution2D(nb_filter, nb_cols, nb_cols, \"relu\")) \r\n    model.add(AveragePooling2D((7, 7)))\r\n    model.add(Flatten())\r\n    model.add(Dropout(dropout))\r\n    model.add(Dense(10, W_regularizer=REG, b_regularizer=REG))\r\n    model.add(Activation(\"softmax\"))\r\n\r\nconv_layers=3 and nb_filter=50 to 100.\r\nIt was comparable with original vgg16 on fine-tuning convergence speed and eventually I stopped using dense layers as it was limiting me in batch size.",
      "votes": null
    },
    {
      "id": "129773",
      "postDate": "08/02/2016 08:12:37",
      "content": "<p><strong>kyv</strong>, thank you. Does it decrease complexity of computations? Or allows to save memory? I see there are pretty much new layers with many filters. Can you clarify it?</p>",
      "rawMarkdown": "**kyv**, thank you. Does it decrease complexity of computations? Or allows to save memory? I see there are pretty much new layers with many filters. Can you clarify it?",
      "votes": null
    },
    {
      "id": "129874",
      "postDate": "08/02/2016 17:15:26",
      "content": "<p>ZFTurbo,\nI got initial config from here\n<a href=\"https://github.com/tdeboissiere/VGG16CAM-keras/blob/master/VGGCAM-keras.py\">https://github.com/tdeboissiere/VGG16CAM-keras/blob/master/VGGCAM-keras.py</a> - so it can be just Conv, Avgpool + flatten+softmax. \nand made top classifier more complex just because I trained it separately with parameter optimizations to see how far can I get without fine-tuning (78% accuracy) and then combined with main vgg16.</p>\n\n<p>My model size was 62mb compared to 550mb. It was slightly faster 500 - 600sec per epoch compared to 700 - 800 with full vgg16 and increased batch_size from 8 to 16 (haven't tried more)</p>",
      "rawMarkdown": "ZFTurbo,\r\nI got initial config from here\r\nhttps://github.com/tdeboissiere/VGG16CAM-keras/blob/master/VGGCAM-keras.py - so it can be just Conv, Avgpool + flatten+softmax. \r\nand made top classifier more complex just because I trained it separately with parameter optimizations to see how far can I get without fine-tuning (78% accuracy) and then combined with main vgg16.\r\n\r\nMy model size was 62mb compared to 550mb. It was slightly faster 500 - 600sec per epoch compared to 700 - 800 with full vgg16 and increased batch_size from 8 to 16 (haven't tried more)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 113963,
      "author_name": "zerrxy",
      "author_url": "",
      "post_date": "04/06/2016 13:41:49",
      "content": "<p>The training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114014,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/06/2016 21:22:13",
      "content": "<p>[quote=Mike;113963]\nThe training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.\n[/quote]</p>\n\n<p>You are right. Need to find out the way how to deal with it. )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114041,
      "author_name": "knguyen",
      "author_url": "",
      "post_date": "04/07/2016 07:21:51",
      "content": "<p>How did you come up with (128, 96) for resized images?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114046,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/07/2016 08:06:49",
      "content": "<p>[quote=Khanh;114041]\nHow did you come up with (128, 96) for resized images?\n[/quote]</p>\n\n<p>I needed to decrease training time consumption. So I choose reasonable picture size where I still can classify images by eyes. Actually in current model reducing pictures even more to (64, 48) gives better LB result. It still requires more experiments with parameters and layers tuning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114059,
      "author_name": "nirmalyaghosh",
      "author_url": "",
      "post_date": "04/07/2016 10:40:30",
      "content": "<p>May I ask how long does it take to train the model for (128, 96)  and on what kind of setup? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114062,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/07/2016 10:57:15",
      "content": "<p>Processor: Intel(R) Core(TM) i7-2600K CPU @ 3.40GHz (8 CPUs), ~3.4GHz</p>\n\n<p>Memory: 8192MB RAM</p>\n\n<p>GPU: NVIDIA GeForce GTX 560 Ti 1GB</p>\n\n<p>Requires around 10-15 minutes overall in GPU mode. Half of the time is image reading.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114129,
      "author_name": "whamfish",
      "author_url": "",
      "post_date": "04/07/2016 17:39:19",
      "content": "<p>ZFTurbo, are you using exactly what you have on Github to get 1.3 on LB? Im not doing as well locally. I just wanted to be sure that I can produce the expected result before I start messing with things  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114132,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/07/2016 17:46:40",
      "content": "<p>No. Code provided on GitHub, will be around 2.20. </p>\n\n<p>Change: nb_epoch = 1, img_rows, img_cols = 48, 64 to achieve better results. </p>\n\n<p>And as I said in the first post: local and leaderboard score is totally different.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114139,
      "author_name": "whamfish",
      "author_url": "",
      "post_date": "04/07/2016 18:32:01",
      "content": "<p>Cool thanks, 2.2 is in the neighborhood of what I'm getting, what should I expect if I make the above changes? Thanks this is helpful for benchmarking. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114151,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/07/2016 20:02:26",
      "content": "<p><strong>DrewWham</strong>: it'll be around ~2.0. The next steps to optimize solution is to go for cross-validation technique, also even more reduction of initial images made solution better for some purpose. Looks like small resolution make CNN focus on overall picture, than on driver clothes or something. ) My current Keras code with cross-validation:</p>\n\n<p><a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Allows to reach ~1.4 on leaderboard. It also pretty fast, requires around 15 minutes on GPU.</p>\n\n<p><strong>Question</strong>: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114154,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "04/07/2016 20:39:58",
      "content": "<p>[quote=ZFTurbo;114151]</p>\n\n<p><strong>Question</strong>: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?</p>\n\n<p>[/quote]\nThank you very much for your code. Usually geometric mean works better for logloss like metrics. And you could also try stacking  <a href=\"http://mlwave.com/kaggle-ensembling-guide/\">http://mlwave.com/kaggle-ensembling-guide/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114183,
      "author_name": "hassiktir",
      "author_url": "",
      "post_date": "04/08/2016 00:55:44",
      "content": "<p>This is weird, I implemented a very similar model before I found this thread and while I was using some slightly different parameters and larger images, my LB score is waaay off from my logloss (anywhere from 7-15).  Saw this thread and thought maybe just a bad model, and tried running your  run_keras_cv.py and my lb score was ~9.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114188,
      "author_name": "ryanpream",
      "author_url": "",
      "post_date": "04/08/2016 02:17:37",
      "content": "<p>There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.</p>\n\n<p>Using the newly added &quot;driver_imgs_list.csv&quot; should make it easier to validate models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114215,
      "author_name": "hassiktir",
      "author_url": "",
      "post_date": "04/08/2016 08:42:47",
      "content": "<p>[quote=Ryan Pream;114188]</p>\n\n<p>There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.</p>\n\n<p>Using the newly added &quot;driver_imgs_list.csv&quot; should make it easier to validate models.</p>\n\n<p>[/quote]</p>\n\n<p>Yeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:</p>\n\n<pre><code>def select_subset_of_driver(percentage_split=.2):\n    &quot;&quot;&quot;\n    &quot;&quot;&quot;\n    driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\n    all_ids = list(driver_df.subject.unique())\n    # for testing\n    np.random.seed(69)\n    np.random.shuffle(all_ids)\n    valid_amount = int(len(all_ids) * percentage_split)\n    train_x_drivers = all_ids[valid_amount:]\n    test_x_drivers = all_ids[:valid_amount]\n    print('-' * 50)\n    print('using subset of drivers as validation')\n    print('using following ids for training: ', train_x_drivers)\n    print('using following ids for testing: ', test_x_drivers)\n    print('-' * 50)\n    test_imgs = driver_df[\n        driver_df['subject'].isin(test_x_drivers)]['img'].values\n    train_imgs = driver_df[\n        driver_df['subject'].isin(train_x_drivers)]['img'].values\n\n    return train_imgs, test_imgs\n\nX, y, train_ids = load_train_data()\ntrain_imgs, validate_imgs = select_subset_of_driver()\n# select indices of validate and train data\nvalidate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\ntrain_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\n# validate subset\nX_validate = X[validate_idx]\ny_validate = y[validate_idx]\nvalidate_ids = train_ids[validate_idx]\n# train subset\nX = X[train_idx]\ny = y[train_idx]\ntrain_ids = train_ids[train_idx]\n</code></pre>\n\n<p>My hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114220,
      "author_name": "potamitis",
      "author_url": "",
      "post_date": "04/08/2016 10:21:03",
      "content": "<p>Hi hassiktir\nYou mentioned that trying to run run_keras_cv.py produced strange results. Can you say how you bypassed this as I get results around 4 when running it for 10 epochs\nthnx</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114232,
      "author_name": "xingyang",
      "author_url": "",
      "post_date": "04/08/2016 11:51:46",
      "content": "<p>[quote=hassiktir;114215]</p>\n\n<p>[quote=Ryan Pream;114188]</p>\n\n<p>There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.</p>\n\n<p>Using the newly added &quot;driver_imgs_list.csv&quot; should make it easier to validate models.</p>\n\n<p>[/quote]</p>\n\n<p>Yeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:</p>\n\n<pre><code>def select_subset_of_driver(percentage_split=.2):\n    &quot;&quot;&quot;\n    &quot;&quot;&quot;\n    driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\n    all_ids = list(driver_df.subject.unique())\n    # for testing\n    np.random.seed(69)\n    np.random.shuffle(all_ids)\n    valid_amount = int(len(all_ids) * percentage_split)\n    train_x_drivers = all_ids[valid_amount:]\n    test_x_drivers = all_ids[:valid_amount]\n    print('-' * 50)\n    print('using subset of drivers as validation')\n    print('using following ids for training: ', train_x_drivers)\n    print('using following ids for testing: ', test_x_drivers)\n    print('-' * 50)\n    test_imgs = driver_df[\n        driver_df['subject'].isin(test_x_drivers)]['img'].values\n    train_imgs = driver_df[\n        driver_df['subject'].isin(train_x_drivers)]['img'].values\n\n    return train_imgs, test_imgs\n\nX, y, train_ids = load_train_data()\ntrain_imgs, validate_imgs = select_subset_of_driver()\n# select indices of validate and train data\nvalidate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\ntrain_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\n# validate subset\nX_validate = X[validate_idx]\ny_validate = y[validate_idx]\nvalidate_ids = train_ids[validate_idx]\n# train subset\nX = X[train_idx]\ny = y[train_idx]\ntrain_ids = train_ids[train_idx]\n</code></pre>\n\n<p>My hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on</p>\n\n<p>[/quote]</p>\n\n<p>You probably need the LeavePLabelOut function (<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.LeavePLabelOut.html\">http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.LeavePLabelOut.html</a>)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114234,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/08/2016 12:03:53",
      "content": "<p><em>There are 28 drivers</em></p>\n\n<p>Actually there are 26 drivers. )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114339,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/09/2016 14:21:57",
      "content": "<p>I created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.</p>\n\n<p><a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>Now it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:</p>\n\n<ol>\n<li>Is there any docs/papers/faqs for dummies how to construct the\nCNN models? Most of the docs concentrate on MNIST, which is not the\ncase.</li>\n<li>I tried to follow VGG-16 model scheme for construction of\nCNNs, but adding second conv/pool block actually make LOSS function\nworse. After tuning of filter number and convolution kernel it's\nsometimes became better. How to find out the optimal number of\nconvolution/pooling/dense layers? Does it somehow depends on input\nimage width/height? Is there some intuitive predictions which CNN\nmodel will be better for given task? </li>\n<li>On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\nquess. And it doesn't improve with epoch number. Otherwise on simple\nmodels while train loss decrease, valid loss either jump, or\ncontinue increasing with each epoch (overfit?). How to make it\ndecrease with each step? I find out that Dropout layers sometimes\nhelp with overfitting. May be some other tricks exists? </li>\n<li>Is there any way in Keras to feed different picture areas to different CNNs\nwith later merging them in one bigger CNN at some stage? How do you\nthink will it make model better? </li>\n<li>As I can see &quot;fit&quot; on CNN works\ntotally unpredictable comapring to XGBoost. Should successfull\nmodels use many epochs? Is there some CNNs, which used for some\nreallife problems to predict something, with only 1 epoch? What\n&quot;optimizer&quot; is the best (I tried adadelta and SGD)? </li>\n<li>How to find out which initial picture size is optimal as input for CNN? Should\nwe use gray or full RGB for this problem?</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114345,
      "author_name": "xingyang",
      "author_url": "",
      "post_date": "04/09/2016 15:08:31",
      "content": "<p>[quote=ZFTurbo;114339]</p>\n\n<p>I created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.</p>\n\n<p><a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>Now it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:</p>\n\n<ol>\n<li>Is there any docs/papers/faqs for dummies how to construct the\nCNN models? Most of the docs concentrate on MNIST, which is not the\ncase.</li>\n<li>I tried to follow VGG-16 model scheme for construction of\nCNNs, but adding second conv/pool block actually make LOSS function\nworse. After tuning of filter number and convolution kernel it's\nsometimes became better. How to find out the optimal number of\nconvolution/pooling/dense layers? Does it somehow depends on input\nimage width/height? Is there some intuitive predictions which CNN\nmodel will be better for given task? </li>\n<li>On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\nquess. And it doesn't improve with epoch number. Otherwise on simple\nmodels while train loss decrease, valid loss either jump, or\ncontinue increasing with each epoch (overfit?). How to make it\ndecrease with each step? I find out that Dropout layers sometimes\nhelp with overfitting. May be some other tricks exists? </li>\n<li>Is there any way in Keras to feed different picture areas to different CNNs\nwith later merging them in one bigger CNN at some stage? How do you\nthink will it make model better? </li>\n<li>As I can see &quot;fit&quot; on CNN works\ntotally unpredictable comapring to XGBoost. Should successfull\nmodels use many epochs? Is there some CNNs, which used for some\nreallife problems to predict something, with only 1 epoch? What\n&quot;optimizer&quot; is the best (I tried adadelta and SGD)? </li>\n<li>How to find out which initial picture size is optimal as input for CNN? Should\nwe use gray or full RGB for this problem?</li>\n</ol>\n\n<p>[/quote]</p>\n\n<p>Here are my personal opinions regarding your questions:</p>\n\n<p>1, Another benchmark contest in Computer Vision is ILSVRC (<a href=\"http://www.image-net.org/\">http://www.image-net.org/</a>). The most famous paper is this one (<a href=\"http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf\">http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf</a>).</p>\n\n<p>2, It is tricky to fine-tune a NN. We may need to explorer more structures.</p>\n\n<p>3, One possibility is using lower learning rate. You may replace the optimizer with SGD (<a href=\"http://keras.io/optimizers/\">http://keras.io/optimizers/</a>). You may also find and save the optimal model by using ModelCheckpoint and EarlyStopping (<a href=\"http://keras.io/callbacks/\">http://keras.io/callbacks/</a>).</p>\n\n<p>4, You may use ImageDataGenerator to manipulate the images a little bit (<a href=\"http://keras.io/preprocessing/image/\">http://keras.io/preprocessing/image/</a>). I don't think dividing the whole image to several smaller images could help in this competition.</p>\n\n<p>5, The same as question 3. I prefer to use SGD. batch_size also makes a difference.</p>\n\n<p>6, If the computing power is not a constraint, I prefer to use color images. The image should not be too small. At least, humans should be able to identify the categories of the images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114348,
      "author_name": "inoryy",
      "author_url": "",
      "post_date": "04/09/2016 16:28:22",
      "content": "<p>[quote=ZFTurbo;113939]</p>\n\n<p>Here is simple solution using CNN to start from:</p>\n\n<p>Ver. 1: <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py</a></p>\n\n<p>Ver. 2 (<strong>UPD 07.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Ver. 3 (<strong>UPD 09.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a>\n[/quote]</p>\n\n<p>Since I've benefited from your code a bit, I think it would be fair to give you some tips:</p>\n\n<p>1.) You don't need opencv for image processing, there's scipy.misc.imread, imresize <br>\nSaves on memory.</p>\n\n<p>2.) You lose a lot of information by going greyscale</p>\n\n<p>3.) Mean normalization</p>\n\n<p>4.) Hyper-parameter (i.e. learning rate) tuning is very important, the default ones are really bad</p>\n\n<p>5.) 1 epoch is not enough for model to generalize well</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114420,
      "author_name": "andrelopes1705",
      "author_url": "",
      "post_date": "04/10/2016 17:02:13",
      "content": "<p>I dont understand how you separate the validation set from the trainset.</p>\n\n<p>Would you clarify?\nI wanted to pass a percentage parameter as well, to get the validation set..</p>\n\n<p>This is rather weird for me!</p>\n\n<p>def load(...):</p>\n\n<pre><code>def copy_selected_drivers(train_data, train_target, driver_id, driver_list):\n        data = []\n        target = []\n        index = []\n        for i in range(len(driver_id)):\n            if driver_id[i] in driver_list:\n                data.append(train_data[i])\n                target.append(train_target[i])\n                index.append(i)\n        data = np.array(data, dtype=np.float32)\n        target = np.array(target, dtype=np.float32)\n        index = np.array(index, dtype=np.uint32)\n        return data, target, index\n\n    train_data, train_target, driver_id, unique_drivers = read_and_normalize_train_data(img_rows, img_cols, color_type_global)\n    test_data, test_id = read_and_normalize_test_data(img_rows, img_cols, color_type_global)\n\n    unique_list_train = ['p002', 'p012', 'p014', 'p015', 'p016', 'p021', 'p022', 'p024',\n                         'p026', 'p035', 'p039', 'p041', 'p042', 'p045', 'p047', 'p049',\n                         'p050', 'p051', 'p052', 'p056', 'p061', 'p064', 'p066', 'p072',\n                         'p075']\n    X_train, Y_train, train_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_train)\n\n    unique_list_valid = ['p081']\n    X_valid, Y_valid, test_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_valid)\n\n    print('Split train: ', len(X_train), len(Y_train))\n    print('Split valid: ', len(X_valid), len(Y_valid))\n    print('Train drivers: ', unique_list_train)\n    print('Test drivers: ', unique_list_valid)\n\nreturn X_train, Y_train, train_index, X_valid, Y_valid, test_index, test_data, test_id\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114422,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/10/2016 17:07:35",
      "content": "<p><strong>Andre lopes</strong>, I split not by images, but by drivers. There are hundreds of images of the same driver. There are only 26 drivers in train set. </p>\n\n<p>In the version of code you provide. 25 drivers used for train and 1 driver for validation. Check &quot;<em>driver_imgs_list.csv</em>&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114423,
      "author_name": "andrelopes1705",
      "author_url": "",
      "post_date": "04/10/2016 17:13:32",
      "content": "<p>I see. Any special reason for doing that?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114427,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/10/2016 18:17:22",
      "content": "<p>[quote=Andre lopes;114423]\nI see. Any special reason for doing that?\n[/quote]</p>\n\n<p>The reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114433,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/10/2016 19:34:46",
      "content": "<p>[quote=ZFTurbo;114427]</p>\n\n<p>[quote=Andre lopes;114423]\nI see. Any special reason for doing that?\n[/quote]</p>\n\n<p>The reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.</p>\n\n<p>[/quote]</p>\n\n<p>There is a <em>huge</em> difference and improvement in accuracy of CV metrics when splitting validation by driver. It's still far from perfect, but with ~3 drivers split out for CV (using 8-fold CV by driver) I'm getting within 20% of my LB score (1.5 CV compared to 1.3 LB). Whilst with a random split of e.g. 25% of images - I get CV loss of 0.1 and accuracy of 98%+ for the same meta-params, which is nonsense. </p>\n\n<p>As far as I can see a random split for CV provides next to no useful information - you may as well train with everything and just fit to the leaderboard score. Although still a bad idea, I don't think that would be a disaster here - i.e. I think it will be hard to badly overfit the public LB - except for the limited number of times you can test per day.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114518,
      "author_name": "maderafunk",
      "author_url": "",
      "post_date": "04/11/2016 15:38:18",
      "content": "<p>Hi, I am new to Keras, I like the idea to manually select the validation set. However, I receive similar results when I don't use the for loop and just set a constant validation set using the last 20% of the data:</p>\n\n<pre><code>    model.fit(X_train_in, Y_train, batch_size=32, nb_epoch=20, show_accuracy=True, verbose=1, validation_split=0.2)\n</code></pre>\n\n<p>As I am not very familiar with Keras, are the weights actually updated within the for loop, or is it not rather the case that they are reset every time and it is just using the weights from the last iteration?</p>\n\n<p>Thanks for the clarification.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115765,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "04/19/2016 23:34:33",
      "content": "<p>I think there is a problem in using <code>np.reshape()</code> versus <code>np.transpose()</code> when loading the images. The current code is</p>\n\n<pre><code>train_data = train_data.reshape(train_data.shape[0], color_type, img_rows, img_cols)\n</code></pre>\n\n<p>I think it should be</p>\n\n<pre><code>train_data = train_data.transpose((0, 3, 1, 2))\n</code></pre>\n\n<p>Reshape scrambles the indices with the results shown in the attached image.</p>\n\n<p>The code I used to create those images is in the <code>load_train</code> method:</p>\n\n<pre><code>img_raw = train_data[100, ...]  # shape = (color_type, img_rows, img_cols)\nimg = np.zeros((img_rows, img_cols, color_type), dtype=np.uint8)\nimg[...,0] = img_raw[0]\nimg[...,1] = img_raw[1]\nimg[...,2] = img_raw[2]\nimg = cv2.resize(img, (640, 480))\ncv2.imwrite('reshape.jpg', img)\ncv2.imshow('reshape', img)\ncv2.waitKey(0)\ncv2.destroyAllWindows()\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115810,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/20/2016 08:47:55",
      "content": "<p>gauss256, this fix only needed for colored images. It won't work for gray_scale images (I checked image looks correct). So my current fix is the following:</p>\n\n<pre><code>if color_type == 1:\n    train_data = train_data.reshape(train_data.shape[0], 1, img_rows, img_cols)\nelse:\n    train_data = train_data.transpose((0, 3, 1, 2))\n</code></pre>\n\n<p>And the same for test:</p>\n\n<pre><code>if color_type == 1:\n    test_data = test_data.reshape(test_data.shape[0], 1, img_rows, img_cols)\nelse:\n    test_data = test_data.transpose((0, 3, 1, 2))\n</code></pre>\n\n<p>I made the changes on GITHUB as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115814,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/20/2016 09:29:32",
      "content": "<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115838,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/20/2016 12:36:32",
      "content": "<p>[quote=ZFTurbo;115814]</p>\n\n<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>\n\n<p>[/quote]</p>\n\n<p>I am seeing this sort of thing too, especially with certain drivers in the CV set. However, I think it is expected behaviour, even though it is not wanted. The model is predicting incorrect class <em>very confidently</em> on the CV set for some proportion of the images. </p>\n\n<p>So I don't think this is a bug, instead the challenge is to find parameters (or maybe different overall approaches) which are less vulnerable to the effect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115862,
      "author_name": "inoryy",
      "author_url": "",
      "post_date": "04/20/2016 14:34:37",
      "content": "<p>[quote=ZFTurbo;115814]</p>\n\n<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>\n\n<p>[/quote]\nAre you sure you're not overfitting? Sure looks like it, but I didn't check the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115893,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "04/20/2016 16:24:52",
      "content": "<p>[quote=ZFTurbo;115814]</p>\n\n<p>BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:</p>\n\n<pre><code>Epoch 1/10\n19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\nEpoch 2/10\n19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\n....\n</code></pre>\n\n<p>It appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.</p>\n\n<p>[/quote]\nI see the same thing and similar results for the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20129/cloud-gpu-starter-project\">Fomoro</a> model. These appear to be symptoms of overfitting, but adding L2 regularization and other tweaks have not made much of a difference for me.</p>\n\n<p>My other theory is that an image size of (24,32) is just too small. As a human looking at that size of image I don't think I could classify them very well at all. But I've tried (48,64) and it wasn't much better. Larger images start bogging down the computer pretty quickly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116980,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "04/26/2016 19:39:15",
      "content": "<p>[quote=ZFTurbo;113939]</p>\n\n<p>Here is simple solution using CNN to start from:</p>\n\n<p>Ver. 1: <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py</a></p>\n\n<p>Ver. 2 (<strong>UPD 07.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py</a></p>\n\n<p>Ver. 3 (<strong>UPD 09.04</strong>): <a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py</a></p>\n\n<p>[/quote]</p>\n\n<p>Hi, I am using your script as a starter. I was trying to play around with the model. I wanted to ask you 2 things:\n1) I tried adding another Convolutional layer, but I was surprised that log loss on test set did not improve, in fact, it actually got worse. Do you know why? I am attaching my script for you to have a look(turbo_script_v5.py)</p>\n\n<p>2)I tried using a pre-trained model VGG_16. I integrated it with your code, but there seems to be some error regarding the shape of input and output. I just call this model instead of create_model_v1 in your scripts run_keras_cv_drivers.py. Any idea what am I doing wrong here? </p>\n\n<p>Here's the code snippet:</p>\n\n<pre><code>def VCG_16(img_rows, img_cols, color_type=1,weights_path=None):\n        model = Sequential()\n        model.add(ZeroPadding2D((1,1),input_shape=(color_type,img_rows,img_cols)))\n        model.add(Convolution2D(64, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(64, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(ZeroPadding2D((1,1)))\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\n        model.add(MaxPooling2D((2,2), strides=(2,2)))\n\n        model.add(Flatten())\n\n        model.add(Dense(4096, activation='relu'))\n        model.add(Dropout(0.5))\n        model.add(Dense(4096, activation='relu'))\n        model.add(Dropout(0.5))\n        model.add(Dense(10, activation='softmax'))\n\n        if weights_path:\n            model.load_weights(weights_path)\n\n        sgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)\n        model.compile(optimizer=sgd, loss='categorical_crossentropy')\n\n        return model\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116981,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/26/2016 19:47:08",
      "content": "<p>Try this from one of my experiments: </p>\n\n<pre><code>  def create_model_v2(img_rows, img_cols, color_type=1):\n    model = Sequential()\n\n    # 1 block 48x64\n    model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    # model.add(Dropout(0.25))\n\n    # 2 block 24x32\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 3 block 12x16\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 4 block 6x8\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    model.add(Flatten())\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10))\n    model.add(Activation('softmax'))\n\n    sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\n    model.compile(loss='categorical_crossentropy', optimizer=sgd)\n    return model\n</code></pre>\n\n<p>or this:</p>\n\n<pre><code>def create_model_v4(img_rows, img_cols, color_type=1):\n    model = Sequential()\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Flatten())\n    model.add(Dense(128, activation='sigmoid', init='he_normal'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10, activation='softmax', init='he_normal'))\n    model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\n    return model\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116985,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "04/26/2016 19:59:20",
      "content": "<p>[quote=ZFTurbo;116981]</p>\n\n<p>Try this from one of my experiments: </p>\n\n<pre><code>  def create_model_v2(img_rows, img_cols, color_type=1):\n    model = Sequential()\n\n    # 1 block 48x64\n    model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    # model.add(Dropout(0.25))\n\n    # 2 block 24x32\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 3 block 12x16\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 4 block 6x8\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    model.add(Flatten())\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10))\n    model.add(Activation('softmax'))\n\n    sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\n    model.compile(loss='categorical_crossentropy', optimizer=sgd)\n    return model\n</code></pre>\n\n<p>or this:</p>\n\n<pre><code>def create_model_v4(img_rows, img_cols, color_type=1):\n    model = Sequential()\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Flatten())\n    model.add(Dense(128, activation='sigmoid', init='he_normal'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10, activation='softmax', init='he_normal'))\n    model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\n    return model\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>Thanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116988,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "04/26/2016 20:11:11",
      "content": "<p>[quote=V.AbhijayArora;116985]</p>\n\n<p>Thanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? </p>\n\n<p>[/quote]</p>\n\n<p>Not yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117923,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "05/01/2016 19:20:50",
      "content": "<p>[quote=ZFTurbo;116988]</p>\n\n<p>[quote=V.AbhijayArora;116985]</p>\n\n<p>Thanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? </p>\n\n<p>[/quote]</p>\n\n<p>Not yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.</p>\n\n<p>[/quote]\nWhat image size works the best? I have tried (48 x 48 x1) or (64 x 48 x 1) or (32 x 24 x 1). Have you tried increasing the number of epochs?I'm not able to get a better score than 1.22 on LB. Any tips?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117983,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "05/02/2016 08:13:34",
      "content": "<p><strong>Abhijay Arora</strong>, it's just the matter of experiments. In current code 32x24 grayscale is better than 64x48 or RGB color mode. I'll plan to post my current code a little bit later, which allows to obtain around 0.85 on LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118264,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "05/03/2016 04:19:13",
      "content": "<p>[quote=ZFTurbo;116981]</p>\n\n<p>Try this from one of my experiments: </p>\n\n<pre><code>  def create_model_v2(img_rows, img_cols, color_type=1):\n    model = Sequential()\n\n    # 1 block 48x64\n    model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(128, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    # model.add(Dropout(0.25))\n\n    # 2 block 24x32\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(256, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 3 block 12x16\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    # 4 block 6x8\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(ZeroPadding2D((1, 1)))\n    model.add(Convolution2D(512, 3, 3, activation='relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n\n    model.add(Flatten())\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(4096, activation='relu'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10))\n    model.add(Activation('softmax'))\n\n    sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\n    model.compile(loss='categorical_crossentropy', optimizer=sgd)\n    return model\n</code></pre>\n\n<p>or this:</p>\n\n<pre><code>def create_model_v4(img_rows, img_cols, color_type=1):\n    model = Sequential()\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\n    model.add(BatchNormalization())\n    model.add(Activation('relu'))\n    model.add(Flatten())\n    model.add(Dense(128, activation='sigmoid', init='he_normal'))\n    model.add(Dropout(0.5))\n    model.add(Dense(10, activation='softmax', init='he_normal'))\n    model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\n    return model\n</code></pre>\n\n<p>[/quote]</p>\n\n<p>What score did these achieve on LB? The first one took 12 hours to train on my laptop.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118398,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "05/03/2016 14:20:15",
      "content": "<p>I posted latest version of my Keras code (1st topic updated):\n<a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py</a></p>\n\n<p>I used some ideas from this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne</a></p>\n\n<ol>\n<li>Code randomly rotate images +-10 degrees</li>\n<li>Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.</li>\n<li>Code uses 64x64 pixel grayscale images</li>\n<li>CNN is actually simple enough to be run on ordinary computer in reasonable time</li>\n<li>I added some useful callback functions:\nEarlyStopping - to stop early after loss stop decreasing\nModelCheckpoint - save best weights and restore them for minimum loss after &quot;fit&quot; ends. Some kind of XGBoost's best_ntree_limit</li>\n</ol>\n\n<p>Notes:</p>\n\n<ol>\n<li>Crossfold score is actually much lower than leaderboard one </li>\n<li>In most cases best loss achieved right after first epoch </li>\n<li><p>I feel like this line: </p>\n\n<p>train_data = train_data.reshape(train_data.shape[0], 1,\n    img_rows, img_cols) </p></li>\n</ol>\n\n<p>should be replaced with some other function. It would be good if someone propose best replacement.</p>\n\n<ol start=\"4\">\n<li>&quot;batch_size&quot; quite strongly affects learning process </li>\n<li>It seems KFold = 26 is best for this problem.</li>\n</ol>\n\n<p>Running &quot;as is&quot; from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118794,
      "author_name": "abhijayvuyyuru",
      "author_url": "",
      "post_date": "05/05/2016 11:13:17",
      "content": "<p>[quote=ZFTurbo;118398]</p>\n\n<p>I posted latest version of my Keras code (1st topic updated):\n<a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py</a></p>\n\n<p>I used some ideas from this post:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne</a></p>\n\n<ol>\n<li>Code randomly rotate images +-10 degrees</li>\n<li>Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.</li>\n<li>Code uses 64x64 pixel grayscale images</li>\n<li>CNN is actually simple enough to be run on ordinary computer in reasonable time</li>\n<li>I added some useful callback functions:\nEarlyStopping - to stop early after loss stop decreasing\nModelCheckpoint - save best weights and restore them for minimum loss after &quot;fit&quot; ends. Some kind of XGBoost's best_ntree_limit</li>\n</ol>\n\n<p>Notes:</p>\n\n<ol>\n<li>Crossfold score is actually much lower than leaderboard one </li>\n<li>In most cases best loss achieved right after first epoch </li>\n<li><p>I feel like this line: </p>\n\n<p>train_data = train_data.reshape(train_data.shape[0], 1,\n    img_rows, img_cols) </p></li>\n</ol>\n\n<p>should be replaced with some other function. It would be good if someone propose best replacement.</p>\n\n<ol start=\"4\">\n<li>&quot;batch_size&quot; quite strongly affects learning process </li>\n<li>It seems KFold = 26 is best for this problem.</li>\n</ol>\n\n<p>Running &quot;as is&quot; from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.</p>\n\n<p>[/quote]\n1) Did you try increasing the number of convolutional layers? I added a 32 x 32 convolutional layer and a 64 X 64 convolutional layer only to get worse results(LB: 0.88)\n2) Why don't you try adding more dense layers? </p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129029,
      "author_name": "arschenchen",
      "author_url": "",
      "post_date": "07/26/2016 00:47:07",
      "content": "<p>Thanks for sharing! It really helped me to get started.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129338,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/28/2016 22:01:21",
      "content": "<p>I've added my code to run pretrained VGG16 Neural Net:\n<a href=\"https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py\">https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py</a></p>\n\n<p>This code allows to get <strong>0.20646</strong> on LB.</p>\n\n<p>Requirements:</p>\n\n<ol>\n<li>16 GB of RAM (with some swap at preprocessing peaks). Code splits &quot;test&quot; in 5 chunks, so it allowed to reduce memory usage.</li>\n<li>Powerful NVIDIA GPU. It takes around a day on 980Ti 6GB.</li>\n<li><strong>Important</strong>: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.</li>\n<li>Weights: <a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3</a></li>\n</ol>\n\n<p>Notes:</p>\n\n<ol>\n<li>Code doesn't use data augmentation</li>\n<li>Code uses totally random test split (doesn't use split by drivers). So don't look at local validation loss, it will be around ~0.01 at the end.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129350,
      "author_name": "hchandaria",
      "author_url": "",
      "post_date": "07/29/2016 01:17:02",
      "content": "<p>@ZFTurbo :  just wanted to confirm that during training of the model the data is split based on drivers ( was looking at your code for function  run_cross_validation_create_models) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129362,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/29/2016 06:59:04",
      "content": "<p><strong>Hetal Chandaria</strong> it initially splitted by drivers, but later shuffled (look for the following code):</p>\n\n<pre><code># Shuffle experiment START !!!\nperm = permutation(len(train_target))\ntrain_data = train_data[perm]\ntrain_target = train_target[perm]\n# Shuffle experiment END !!!\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129369,
      "author_name": "ferris",
      "author_url": "",
      "post_date": "07/29/2016 08:28:49",
      "content": "<p>Thanks for sharing!</p>\n\n<p>So the key is version of keras? ToT</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129370,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/29/2016 08:32:34",
      "content": "<p><strong>Ferris</strong>, yes. At least this code works much better on outdated version.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129483,
      "author_name": "liviux",
      "author_url": "",
      "post_date": "07/30/2016 09:28:26",
      "content": "<p>Thanks for the efforts, very educative!</p>\n\n<p>One question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129484,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "07/30/2016 09:33:08",
      "content": "<p>[quote=liviu;129483]</p>\n\n<p>Thanks for the efforts, very educative!</p>\n\n<p>One question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day). </p>\n\n<p>[/quote]</p>\n\n<p>I suspect not. I think you'll find that anything scoring 0.5 or better would take literally weeks to train on a CPU-based setup, and you would need to have chosen the meta-params correctly. I got down to 0.89 on a CPU rig after a few iterations, and that took 2.5 days training time. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129485,
      "author_name": "liviux",
      "author_url": "",
      "post_date": "07/30/2016 09:48:27",
      "content": "<p>Not what I was hoping to hear, but it gives me the green light to move on :P. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129519,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "07/30/2016 16:18:09",
      "content": "<p>Hi, @ZFTurbo,\nThank you for your nice code sharing. Is there any reason for not split by driver this time?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129520,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/30/2016 16:37:38",
      "content": "<p><strong>Li Li</strong>, for some reason score while splitting by drivers was worse.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129531,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "07/30/2016 18:25:23",
      "content": "<p>[quote=ZFTurbo;129338]\n 3. <strong>Important</strong>: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.\n[/quote]</p>\n\n<p>Can you post score on latest Keras version you tried? I tried on 1.0.6 and got 0.23 with 8 folds with some modifications (GAP instead of Dense Layers for memory optimization). I will try to debug your code with new Keras version, but wander if you have ideas why its different and what score you got.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129534,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/30/2016 19:02:40",
      "content": "<p><strong>kyv</strong>, It was run by my teammates. And they had problems with my code. I see 0.29 and 0.27 in list (it was calculated on 1.0.* version).</p>\n\n<p>I don't have ideas, except that learning process is somehow different in 0.2.0 and 1.0.*. I probably rerun this code on latest version after contest ends to check the difference myself.</p>\n\n<p>I also notice difference in loss in Nerve contest. When I debug Kernel at local machine, loss was much different from version on Kaggle servers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129642,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "08/01/2016 08:21:34",
      "content": "<p>Could be because of different cudnn or theano versions. You also have to clean theano cache once in a while - a for sure if you change something like cuda drivers, cudnn or theano version.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129644,
      "author_name": "keiku322",
      "author_url": "",
      "post_date": "08/01/2016 08:46:00",
      "content": "<p>@kyv I'm sorry for not knowing. What is &quot;GAP instead of Dense Layers for memory optimization&quot;?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129737,
      "author_name": "tianzhou",
      "author_url": "",
      "post_date": "08/02/2016 01:13:48",
      "content": "<p>[quote=Keiku;129644]</p>\n\n<p>@kyv I'm sorry for not knowing. What is &quot;GAP instead of Dense Layers for memory optimization&quot;?</p>\n\n<p>[/quote]\nI think that's global average pooling layer. You can find this concept in paper:Network In Network</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129740,
      "author_name": "keiku322",
      "author_url": "",
      "post_date": "08/02/2016 01:27:15",
      "content": "<p>@Tian Zhou Thank you for telling me about.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129759,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "08/02/2016 05:35:28",
      "content": "<p>[quote=Keiku;129644]</p>\n\n<p>@kyv I'm sorry for not knowing. What is &quot;GAP instead of Dense Layers for memory optimization&quot;?</p>\n\n<p>[/quote]\nYes its AveragePooling layer with stride equal to image size and channels equal to number of classes followed by softmax .\nBy doing this you reduce model size by 80% (as most parameters are in dense layers) and usually you don't sacrifice model accuracy. New models often uses this approach. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129763,
      "author_name": "keiku322",
      "author_url": "",
      "post_date": "08/02/2016 06:07:52",
      "content": "<p>@kyv Thank you for your advice. I will try with new Keras version for future reference, too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129766,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "08/02/2016 07:12:52",
      "content": "<p><strong>kyv</strong>, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129770,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "08/02/2016 08:05:30",
      "content": "<p>[quote=ZFTurbo;129766]</p>\n\n<p><strong>kyv</strong>, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )\n[/quote]</p>\n\n<p>I replaced last dense layers with</p>\n\n<pre><code> for _ in range(conv_layers):\n    model.add(Convolution2D(nb_filter, nb_cols, nb_cols, &quot;relu&quot;)) \nmodel.add(AveragePooling2D((7, 7)))\nmodel.add(Flatten())\nmodel.add(Dropout(dropout))\nmodel.add(Dense(10, W_regularizer=REG, b_regularizer=REG))\nmodel.add(Activation(&quot;softmax&quot;))\n</code></pre>\n\n<p>conv_layers=3 and nb_filter=50 to 100.\nIt was comparable with original vgg16 on fine-tuning convergence speed and eventually I stopped using dense layers as it was limiting me in batch size. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129773,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "08/02/2016 08:12:37",
      "content": "<p><strong>kyv</strong>, thank you. Does it decrease complexity of computations? Or allows to save memory? I see there are pretty much new layers with many filters. Can you clarify it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129874,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "08/02/2016 17:15:26",
      "content": "<p>ZFTurbo,\nI got initial config from here\n<a href=\"https://github.com/tdeboissiere/VGG16CAM-keras/blob/master/VGGCAM-keras.py\">https://github.com/tdeboissiere/VGG16CAM-keras/blob/master/VGGCAM-keras.py</a> - so it can be just Conv, Avgpool + flatten+softmax. \nand made top classifier more complex just because I trained it separately with parameter optimizations to see how far can I get without fine-tuning (78% accuracy) and then combined with main vgg16.</p>\n\n<p>My model size was 62mb compared to 550mb. It was slightly faster 500 - 600sec per epoch compared to 700 - 800 with full vgg16 and increased batch_size from 8 to 16 (haven't tried more)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "113939": "Here is simple solution using CNN to start from:\r\n\r\nVer. 1: https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\r\n\r\nVer. 2 (**UPD 07.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nVer. 3 (**UPD 09.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\nVer. 4 (**UPD 03.05**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\r\n\r\nVer. 5 (**UPD 29.07**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py - pretrained VGG16 Net.",
    "113963": "The training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.",
    "114014": "[quote=Mike;113963]\r\nThe training set contains a lot of similar images (photos of the same driver with several seconds interval), so score predicted on the validation subset is much better.\r\n[/quote]\r\n\r\nYou are right. Need to find out the way how to deal with it. )",
    "114041": "How did you come up with (128, 96) for resized images?",
    "114046": "[quote=Khanh;114041]\r\nHow did you come up with (128, 96) for resized images?\r\n[/quote]\r\n\r\nI needed to decrease training time consumption. So I choose reasonable picture size where I still can classify images by eyes. Actually in current model reducing pictures even more to (64, 48) gives better LB result. It still requires more experiments with parameters and layers tuning.",
    "114059": "May I ask how long does it take to train the model for (128, 96)  and on what kind of setup?",
    "114062": "Processor: Intel(R) Core(TM) i7-2600K CPU @ 3.40GHz (8 CPUs), ~3.4GHz\r\n\r\nMemory: 8192MB RAM\r\n\r\nGPU: NVIDIA GeForce GTX 560 Ti 1GB\r\n\r\nRequires around 10-15 minutes overall in GPU mode. Half of the time is image reading.",
    "114129": "ZFTurbo, are you using exactly what you have on Github to get 1.3 on LB? Im not doing as well locally. I just wanted to be sure that I can produce the expected result before I start messing with things",
    "114132": "No. Code provided on GitHub, will be around 2.20. \r\n\r\nChange: nb_epoch = 1, img_rows, img_cols = 48, 64 to achieve better results. \r\n\r\nAnd as I said in the first post: local and leaderboard score is totally different.",
    "114139": "Cool thanks, 2.2 is in the neighborhood of what I'm getting, what should I expect if I make the above changes? Thanks this is helpful for benchmarking.",
    "114151": "**DrewWham**: it'll be around ~2.0. The next steps to optimize solution is to go for cross-validation technique, also even more reduction of initial images made solution better for some purpose. Looks like small resolution make CNN focus on overall picture, than on driver clothes or something. ) My current Keras code with cross-validation:\r\n\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nAllows to reach ~1.4 on leaderboard. It also pretty fast, requires around 15 minutes on GPU.\r\n\r\n**Question**: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?",
    "114154": "[quote=ZFTurbo;114151]\r\n\r\n**Question**: what is the best way to combine K predictions for test data? Is just arithmetic mean always OK, or it's better to use something else?\r\n\r\n[/quote]\r\nThank you very much for your code. Usually geometric mean works better for logloss like metrics. And you could also try stacking  http://mlwave.com/kaggle-ensembling-guide/",
    "114183": "This is weird, I implemented a very similar model before I found this thread and while I was using some slightly different parameters and larger images, my LB score is waaay off from my logloss (anywhere from 7-15).  Saw this thread and thought maybe just a bad model, and tried running your  run_keras_cv.py and my lb score was ~9.",
    "114188": "There are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.\r\n\r\nUsing the newly added \"driver_imgs_list.csv\" should make it easier to validate models.",
    "114215": "[quote=Ryan Pream;114188]\r\n\r\nThere are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.\r\n\r\nUsing the newly added \"driver_imgs_list.csv\" should make it easier to validate models.\r\n\r\n[/quote]\r\n\r\n\r\nYeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:\r\n\r\n    def select_subset_of_driver(percentage_split=.2):\r\n        \"\"\"\r\n        \"\"\"\r\n        driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\r\n        all_ids = list(driver_df.subject.unique())\r\n        # for testing\r\n        np.random.seed(69)\r\n        np.random.shuffle(all_ids)\r\n        valid_amount = int(len(all_ids) * percentage_split)\r\n        train_x_drivers = all_ids[valid_amount:]\r\n        test_x_drivers = all_ids[:valid_amount]\r\n        print('-' * 50)\r\n        print('using subset of drivers as validation')\r\n        print('using following ids for training: ', train_x_drivers)\r\n        print('using following ids for testing: ', test_x_drivers)\r\n        print('-' * 50)\r\n        test_imgs = driver_df[\r\n            driver_df['subject'].isin(test_x_drivers)]['img'].values\r\n        train_imgs = driver_df[\r\n            driver_df['subject'].isin(train_x_drivers)]['img'].values\r\n    \r\n        return train_imgs, test_imgs\r\n\r\n    X, y, train_ids = load_train_data()\r\n    train_imgs, validate_imgs = select_subset_of_driver()\r\n    # select indices of validate and train data\r\n    validate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\r\n    train_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\r\n    # validate subset\r\n    X_validate = X[validate_idx]\r\n    y_validate = y[validate_idx]\r\n    validate_ids = train_ids[validate_idx]\r\n    # train subset\r\n    X = X[train_idx]\r\n    y = y[train_idx]\r\n    train_ids = train_ids[train_idx]\r\n\r\n\r\nMy hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on",
    "114220": "Hi hassiktir\r\nYou mentioned that trying to run run_keras_cv.py produced strange results. Can you say how you bypassed this as I get results around 4 when running it for 10 epochs\r\nthnx",
    "114232": "[quote=hassiktir;114215]\r\n\r\n[quote=Ryan Pream;114188]\r\n\r\nThere are 28 drivers and 22,424 images in the training data.  With ~800 images of each driver it is rather easy to train a model that does great on unseen images of known drivers, but is very poor on unseen images of unknown drivers.\r\n\r\nUsing the newly added \"driver_imgs_list.csv\" should make it easier to validate models.\r\n\r\n[/quote]\r\n\r\n\r\nYeesh, trying to figure out a good way to do this (and figured out my error was due to using sample_submission id's which dont line up with the way the test data was loading), but can anyone suggest something better than this:\r\n\r\n    def select_subset_of_driver(percentage_split=.2):\r\n        \"\"\"\r\n        \"\"\"\r\n        driver_df = pd.read_csv(base_path + 'driver_imgs_list.csv')\r\n        all_ids = list(driver_df.subject.unique())\r\n        # for testing\r\n        np.random.seed(69)\r\n        np.random.shuffle(all_ids)\r\n        valid_amount = int(len(all_ids) * percentage_split)\r\n        train_x_drivers = all_ids[valid_amount:]\r\n        test_x_drivers = all_ids[:valid_amount]\r\n        print('-' * 50)\r\n        print('using subset of drivers as validation')\r\n        print('using following ids for training: ', train_x_drivers)\r\n        print('using following ids for testing: ', test_x_drivers)\r\n        print('-' * 50)\r\n        test_imgs = driver_df[\r\n            driver_df['subject'].isin(test_x_drivers)]['img'].values\r\n        train_imgs = driver_df[\r\n            driver_df['subject'].isin(train_x_drivers)]['img'].values\r\n    \r\n        return train_imgs, test_imgs\r\n\r\n    X, y, train_ids = load_train_data()\r\n    train_imgs, validate_imgs = select_subset_of_driver()\r\n    # select indices of validate and train data\r\n    validate_idx = np.in1d(train_ids, validate_imgs).nonzero()[0]\r\n    train_idx = np.in1d(train_ids, train_imgs).nonzero()[0]\r\n    # validate subset\r\n    X_validate = X[validate_idx]\r\n    y_validate = y[validate_idx]\r\n    validate_ids = train_ids[validate_idx]\r\n    # train subset\r\n    X = X[train_idx]\r\n    y = y[train_idx]\r\n    train_ids = train_ids[train_idx]\r\n\r\n\r\nMy hope was to select 20% of the drivers (even though there aren't equal amounts of photos or classifications etc for each) and then use those later to validate on\r\n\r\n[/quote]\r\n\r\nYou probably need the LeavePLabelOut function (http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.LeavePLabelOut.html)",
    "114234": "*There are 28 drivers*\r\n\r\nActually there are 26 drivers. )",
    "114339": "I created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.\r\n\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\nNow it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:\r\n\r\n 1. Is there any docs/papers/faqs for dummies how to construct the\r\n    CNN models? Most of the docs concentrate on MNIST, which is not the\r\n    case.\r\n 2. I tried to follow VGG-16 model scheme for construction of\r\n    CNNs, but adding second conv/pool block actually make LOSS function\r\n    worse. After tuning of filter number and convolution kernel it's\r\n    sometimes became better. How to find out the optimal number of\r\n    convolution/pooling/dense layers? Does it somehow depends on input\r\n    image width/height? Is there some intuitive predictions which CNN\r\n    model will be better for given task? \r\n 3. On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\r\n    quess. And it doesn't improve with epoch number. Otherwise on simple\r\n    models while train loss decrease, valid loss either jump, or\r\n    continue increasing with each epoch (overfit?). How to make it\r\n    decrease with each step? I find out that Dropout layers sometimes\r\n    help with overfitting. May be some other tricks exists? \r\n 4. Is there any way in Keras to feed different picture areas to different CNNs\r\n    with later merging them in one bigger CNN at some stage? How do you\r\n    think will it make model better? \r\n 5. As I can see \"fit\" on CNN works\r\n    totally unpredictable comapring to XGBoost. Should successfull\r\n    models use many epochs? Is there some CNNs, which used for some\r\n    reallife problems to predict something, with only 1 epoch? What\r\n    \"optimizer\" is the best (I tried adadelta and SGD)? \r\n 6. How to find out which initial picture size is optimal as input for CNN? Should\r\n    we use gray or full RGB for this problem?",
    "114345": "[quote=ZFTurbo;114339]\r\n\r\nI created next code version with cross validation based on driver ID. Leaderboard score stays the same, since I didn't change the CNN model. But loss value for validation now is reflect real expected value on test data.\r\n\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\nNow it's time to tune the model. Since I actually didn't have much experience with CNN, I have some questions. Probably experienced users can answer them:\r\n\r\n 1. Is there any docs/papers/faqs for dummies how to construct the\r\n    CNN models? Most of the docs concentrate on MNIST, which is not the\r\n    case.\r\n 2. I tried to follow VGG-16 model scheme for construction of\r\n    CNNs, but adding second conv/pool block actually make LOSS function\r\n    worse. After tuning of filter number and convolution kernel it's\r\n    sometimes became better. How to find out the optimal number of\r\n    convolution/pooling/dense layers? Does it somehow depends on input\r\n    image width/height? Is there some intuitive predictions which CNN\r\n    model will be better for given task? \r\n 3. On complicated models (like VGG) LOSS function stuck at ~2.3 value, which equals to random\r\n    quess. And it doesn't improve with epoch number. Otherwise on simple\r\n    models while train loss decrease, valid loss either jump, or\r\n    continue increasing with each epoch (overfit?). How to make it\r\n    decrease with each step? I find out that Dropout layers sometimes\r\n    help with overfitting. May be some other tricks exists? \r\n 4. Is there any way in Keras to feed different picture areas to different CNNs\r\n    with later merging them in one bigger CNN at some stage? How do you\r\n    think will it make model better? \r\n 5. As I can see \"fit\" on CNN works\r\n    totally unpredictable comapring to XGBoost. Should successfull\r\n    models use many epochs? Is there some CNNs, which used for some\r\n    reallife problems to predict something, with only 1 epoch? What\r\n    \"optimizer\" is the best (I tried adadelta and SGD)? \r\n 6. How to find out which initial picture size is optimal as input for CNN? Should\r\n    we use gray or full RGB for this problem?\r\n\r\n[/quote]\r\n\r\nHere are my personal opinions regarding your questions:\r\n\r\n1, Another benchmark contest in Computer Vision is ILSVRC (http://www.image-net.org/). The most famous paper is this one (http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf).\r\n\r\n2, It is tricky to fine-tune a NN. We may need to explorer more structures.\r\n\r\n3, One possibility is using lower learning rate. You may replace the optimizer with SGD (http://keras.io/optimizers/). You may also find and save the optimal model by using ModelCheckpoint and EarlyStopping (http://keras.io/callbacks/).\r\n\r\n4, You may use ImageDataGenerator to manipulate the images a little bit (http://keras.io/preprocessing/image/). I don't think dividing the whole image to several smaller images could help in this competition.\r\n\r\n5, The same as question 3. I prefer to use SGD. batch_size also makes a difference.\r\n\r\n6, If the computing power is not a constraint, I prefer to use color images. The image should not be too small. At least, humans should be able to identify the categories of the images.",
    "114348": "[quote=ZFTurbo;113939]\r\n\r\nHere is simple solution using CNN to start from:\r\n\r\nVer. 1: https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\r\n\r\nVer. 2 (**UPD 07.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nVer. 3 (**UPD 09.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n[/quote]\r\n\r\nSince I've benefited from your code a bit, I think it would be fair to give you some tips:\r\n\r\n1.) You don't need opencv for image processing, there's scipy.misc.imread, imresize  \r\nSaves on memory.\r\n\r\n2.) You lose a lot of information by going greyscale\r\n\r\n3.) Mean normalization\r\n\r\n4.) Hyper-parameter (i.e. learning rate) tuning is very important, the default ones are really bad\r\n\r\n5.) 1 epoch is not enough for model to generalize well",
    "114420": "I dont understand how you separate the validation set from the trainset.\r\n\r\nWould you clarify?\r\nI wanted to pass a percentage parameter as well, to get the validation set..\r\n\r\nThis is rather weird for me!\r\n\r\ndef load(...):\r\n\r\n    def copy_selected_drivers(train_data, train_target, driver_id, driver_list):\r\n            data = []\r\n            target = []\r\n            index = []\r\n            for i in range(len(driver_id)):\r\n                if driver_id[i] in driver_list:\r\n                    data.append(train_data[i])\r\n                    target.append(train_target[i])\r\n                    index.append(i)\r\n            data = np.array(data, dtype=np.float32)\r\n            target = np.array(target, dtype=np.float32)\r\n            index = np.array(index, dtype=np.uint32)\r\n            return data, target, index\r\n    \r\n        train_data, train_target, driver_id, unique_drivers = read_and_normalize_train_data(img_rows, img_cols, color_type_global)\r\n        test_data, test_id = read_and_normalize_test_data(img_rows, img_cols, color_type_global)\r\n    \r\n        unique_list_train = ['p002', 'p012', 'p014', 'p015', 'p016', 'p021', 'p022', 'p024',\r\n                             'p026', 'p035', 'p039', 'p041', 'p042', 'p045', 'p047', 'p049',\r\n                             'p050', 'p051', 'p052', 'p056', 'p061', 'p064', 'p066', 'p072',\r\n                             'p075']\r\n        X_train, Y_train, train_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_train)\r\n    \r\n        unique_list_valid = ['p081']\r\n        X_valid, Y_valid, test_index = copy_selected_drivers(train_data, train_target, driver_id, unique_list_valid)\r\n    \r\n        print('Split train: ', len(X_train), len(Y_train))\r\n        print('Split valid: ', len(X_valid), len(Y_valid))\r\n        print('Train drivers: ', unique_list_train)\r\n        print('Test drivers: ', unique_list_valid)\r\n\r\n    return X_train, Y_train, train_index, X_valid, Y_valid, test_index, test_data, test_id",
    "114422": "**Andre lopes**, I split not by images, but by drivers. There are hundreds of images of the same driver. There are only 26 drivers in train set. \r\n\r\nIn the version of code you provide. 25 drivers used for train and 1 driver for validation. Check \"*driver_imgs_list.csv*\"",
    "114423": "I see. Any special reason for doing that?",
    "114427": "[quote=Andre lopes;114423]\r\nI see. Any special reason for doing that?\r\n[/quote]\r\n\r\nThe reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.",
    "114433": "[quote=ZFTurbo;114427]\r\n\r\n[quote=Andre lopes;114423]\r\nI see. Any special reason for doing that?\r\n[/quote]\r\n\r\nThe reason is following: If you will add images of the same driver in train and validation set then CNN will try to learn also the driver properties, like color of T-shirt for example. Since in test set drivers are different from train set, CNN will have different behaviors on train set and test set. So you can't predict loss and accuracy of CNN. To avoid this we emulate test set by completely excluding some drivers from train set and put them in validation set only.\r\n\r\n[/quote]\r\n\r\nThere is a *huge* difference and improvement in accuracy of CV metrics when splitting validation by driver. It's still far from perfect, but with ~3 drivers split out for CV (using 8-fold CV by driver) I'm getting within 20% of my LB score (1.5 CV compared to 1.3 LB). Whilst with a random split of e.g. 25% of images - I get CV loss of 0.1 and accuracy of 98%+ for the same meta-params, which is nonsense. \r\n\r\nAs far as I can see a random split for CV provides next to no useful information - you may as well train with everything and just fit to the leaderboard score. Although still a bad idea, I don't think that would be a disaster here - i.e. I think it will be hard to badly overfit the public LB - except for the limited number of times you can test per day.",
    "114518": "Hi, I am new to Keras, I like the idea to manually select the validation set. However, I receive similar results when I don't use the for loop and just set a constant validation set using the last 20% of the data:\r\n\r\n        model.fit(X_train_in, Y_train, batch_size=32, nb_epoch=20, show_accuracy=True, verbose=1, validation_split=0.2)\r\n\r\nAs I am not very familiar with Keras, are the weights actually updated within the for loop, or is it not rather the case that they are reset every time and it is just using the weights from the last iteration?\r\n\r\nThanks for the clarification.",
    "115765": "I think there is a problem in using `np.reshape()` versus `np.transpose()` when loading the images. The current code is\r\n\r\n    train_data = train_data.reshape(train_data.shape[0], color_type, img_rows, img_cols)\r\n\r\nI think it should be\r\n\r\n    train_data = train_data.transpose((0, 3, 1, 2))\r\n\r\nReshape scrambles the indices with the results shown in the attached image.\r\n\r\nThe code I used to create those images is in the `load_train` method:\r\n\r\n    img_raw = train_data[100, ...]  # shape = (color_type, img_rows, img_cols)\r\n    img = np.zeros((img_rows, img_cols, color_type), dtype=np.uint8)\r\n    img[...,0] = img_raw[0]\r\n    img[...,1] = img_raw[1]\r\n    img[...,2] = img_raw[2]\r\n    img = cv2.resize(img, (640, 480))\r\n    cv2.imwrite('reshape.jpg', img)\r\n    cv2.imshow('reshape', img)\r\n    cv2.waitKey(0)\r\n    cv2.destroyAllWindows()",
    "115810": "gauss256, this fix only needed for colored images. It won't work for gray_scale images (I checked image looks correct). So my current fix is the following:\r\n\r\n    if color_type == 1:\r\n        train_data = train_data.reshape(train_data.shape[0], 1, img_rows, img_cols)\r\n    else:\r\n        train_data = train_data.transpose((0, 3, 1, 2))\r\n\r\nAnd the same for test:\r\n\r\n    if color_type == 1:\r\n        test_data = test_data.reshape(test_data.shape[0], 1, img_rows, img_cols)\r\n    else:\r\n        test_data = test_data.transpose((0, 3, 1, 2))\r\n\r\nI made the changes on GITHUB as well.",
    "115814": "BTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.",
    "115838": "[quote=ZFTurbo;115814]\r\n\r\nBTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.\r\n\r\n[/quote]\r\n\r\nI am seeing this sort of thing too, especially with certain drivers in the CV set. However, I think it is expected behaviour, even though it is not wanted. The model is predicting incorrect class *very confidently* on the CV set for some proportion of the images. \r\n\r\nSo I don't think this is a bug, instead the challenge is to find parameters (or maybe different overall approaches) which are less vulnerable to the effect.",
    "115862": "[quote=ZFTurbo;115814]\r\n\r\nBTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.\r\n\r\n[/quote]\r\nAre you sure you're not overfitting? Sure looks like it, but I didn't check the code.",
    "115893": "[quote=ZFTurbo;115814]\r\n\r\nBTW: The main problem with my latest CV code based on drivers is the following. Whenever I do it doesn't converge with next epochs. It looks like this:\r\n\r\n    Epoch 1/10\r\n    19005/19005 [==============================] - 189s - loss: 1.3394 - val_loss: 1.9018\r\n    Epoch 2/10\r\n    19005/19005 [==============================] - 175s - loss: 0.3714 - val_loss: 2.6110\r\n    ....\r\n\r\nIt appears on different models and on different optimizers. Looks like the bug somewhere in the code. It would be good if someone will find what cause this.\r\n\r\n[/quote]\r\nI see the same thing and similar results for the [Fomoro][1] model. These appear to be symptoms of overfitting, but adding L2 regularization and other tweaks have not made much of a difference for me.\r\n\r\nMy other theory is that an image size of (24,32) is just too small. As a human looking at that size of image I don't think I could classify them very well at all. But I've tried (48,64) and it wasn't much better. Larger images start bogging down the computer pretty quickly.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20129/cloud-gpu-starter-project",
    "116980": "[quote=ZFTurbo;113939]\r\n\r\nHere is simple solution using CNN to start from:\r\n\r\nVer. 1: https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_simple.py\r\n\r\nVer. 2 (**UPD 07.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv.py\r\n\r\nVer. 3 (**UPD 09.04**): https://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers.py\r\n\r\n\r\n[/quote]\r\n\r\nHi, I am using your script as a starter. I was trying to play around with the model. I wanted to ask you 2 things:\r\n1) I tried adding another Convolutional layer, but I was surprised that log loss on test set did not improve, in fact, it actually got worse. Do you know why? I am attaching my script for you to have a look(turbo_script_v5.py)\r\n\r\n2)I tried using a pre-trained model VGG_16. I integrated it with your code, but there seems to be some error regarding the shape of input and output. I just call this model instead of create_model_v1 in your scripts run_keras_cv_drivers.py. Any idea what am I doing wrong here? \r\n\r\nHere's the code snippet:\r\n\r\n    def VCG_16(img_rows, img_cols, color_type=1,weights_path=None):\r\n    \t\tmodel = Sequential()\r\n    \t\tmodel.add(ZeroPadding2D((1,1),input_shape=(color_type,img_rows,img_cols)))\r\n    \t\tmodel.add(Convolution2D(64, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(64, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(128, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(128, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(256, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(256, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(256, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(ZeroPadding2D((1,1)))\r\n    \t\tmodel.add(Convolution2D(512, 3, 3, activation='relu'))\r\n    \t\tmodel.add(MaxPooling2D((2,2), strides=(2,2)))\r\n    \r\n    \t\tmodel.add(Flatten())\r\n    \r\n    \t\tmodel.add(Dense(4096, activation='relu'))\r\n    \t\tmodel.add(Dropout(0.5))\r\n    \t\tmodel.add(Dense(4096, activation='relu'))\r\n    \t\tmodel.add(Dropout(0.5))\r\n    \t\tmodel.add(Dense(10, activation='softmax'))\r\n    \r\n    \t\tif weights_path:\r\n    \t\t    model.load_weights(weights_path)\r\n    \r\n    \t\tsgd = SGD(lr=0.1, decay=1e-6, momentum=0.9, nesterov=True)\r\n    \t\tmodel.compile(optimizer=sgd, loss='categorical_crossentropy')\r\n    \t\r\n    \t\treturn model",
    "116981": "Try this from one of my experiments: \r\n \r\n\r\n      def create_model_v2(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n    \r\n        # 1 block 48x64\r\n        model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n        # model.add(Dropout(0.25))\r\n    \r\n        # 2 block 24x32\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 3 block 12x16\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 4 block 6x8\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        model.add(Flatten())\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10))\r\n        model.add(Activation('softmax'))\r\n    \r\n        sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\r\n        model.compile(loss='categorical_crossentropy', optimizer=sgd)\r\n        return model\r\n\r\nor this:\r\n\r\n    def create_model_v4(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Flatten())\r\n        model.add(Dense(128, activation='sigmoid', init='he_normal'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10, activation='softmax', init='he_normal'))\r\n        model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\r\n        return model",
    "116985": "[quote=ZFTurbo;116981]\r\n\r\nTry this from one of my experiments: \r\n \r\n\r\n      def create_model_v2(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n    \r\n        # 1 block 48x64\r\n        model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n        # model.add(Dropout(0.25))\r\n    \r\n        # 2 block 24x32\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 3 block 12x16\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 4 block 6x8\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        model.add(Flatten())\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10))\r\n        model.add(Activation('softmax'))\r\n    \r\n        sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\r\n        model.compile(loss='categorical_crossentropy', optimizer=sgd)\r\n        return model\r\n\r\nor this:\r\n\r\n    def create_model_v4(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Flatten())\r\n        model.add(Dense(128, activation='sigmoid', init='he_normal'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10, activation='softmax', init='he_normal'))\r\n        model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\r\n        return model\r\n\r\n[/quote]\r\n\r\nThanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights?",
    "116988": "[quote=V.AbhijayArora;116985]\r\n\r\nThanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? \r\n\r\n[/quote]\r\n\r\nNot yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.",
    "117923": "[quote=ZFTurbo;116988]\r\n\r\n[quote=V.AbhijayArora;116985]\r\n\r\nThanks, will try that out. Are you using RGB or Grayscale images? Did you try any pre trained models such as VGG_16 with loaded weights? \r\n\r\n[/quote]\r\n\r\nNot yet, my current experiments and TOP score doesn't use any CNN. But I'll try PRE-Trained nets later for sure. ) As I can see all TOP solutions now use them.\r\n\r\n[/quote]\r\nWhat image size works the best? I have tried (48 x 48 x1) or (64 x 48 x 1) or (32 x 24 x 1). Have you tried increasing the number of epochs?I'm not able to get a better score than 1.22 on LB. Any tips?",
    "117983": "**Abhijay Arora**, it's just the matter of experiments. In current code 32x24 grayscale is better than 64x48 or RGB color mode. I'll plan to post my current code a little bit later, which allows to obtain around 0.85 on LB.",
    "118264": "[quote=ZFTurbo;116981]\r\n\r\nTry this from one of my experiments: \r\n \r\n\r\n      def create_model_v2(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n    \r\n        # 1 block 48x64\r\n        model.add(ZeroPadding2D((1, 1), input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(128, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n        # model.add(Dropout(0.25))\r\n    \r\n        # 2 block 24x32\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(256, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 3 block 12x16\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        # 4 block 6x8\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(ZeroPadding2D((1, 1)))\r\n        model.add(Convolution2D(512, 3, 3, activation='relu'))\r\n        model.add(MaxPooling2D(pool_size=(2, 2)))\r\n    \r\n        model.add(Flatten())\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(4096, activation='relu'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10))\r\n        model.add(Activation('softmax'))\r\n    \r\n        sgd = SGD(lr=0.05, decay=0, momentum=0, nesterov=True)\r\n        model.compile(loss='categorical_crossentropy', optimizer=sgd)\r\n        return model\r\n\r\nor this:\r\n\r\n    def create_model_v4(img_rows, img_cols, color_type=1):\r\n        model = Sequential()\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal', input_shape=(color_type, img_rows, img_cols)))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(32, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(64, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, subsample=(2, 2), init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Convolution2D(128, 3, 3, border_mode='same', init='he_normal'))\r\n        model.add(BatchNormalization())\r\n        model.add(Activation('relu'))\r\n        model.add(Flatten())\r\n        model.add(Dense(128, activation='sigmoid', init='he_normal'))\r\n        model.add(Dropout(0.5))\r\n        model.add(Dense(10, activation='softmax', init='he_normal'))\r\n        model.compile(Adam(lr=1e-3), loss='categorical_crossentropy')\r\n        return model\r\n\r\n[/quote]\r\n\r\nWhat score did these achieve on LB? The first one took 12 hours to train on my laptop.",
    "118398": "I posted latest version of my Keras code (1st topic updated):\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\r\n\r\nI used some ideas from this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\r\n\r\n1. Code randomly rotate images +-10 degrees\r\n2. Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.\r\n3. Code uses 64x64 pixel grayscale images\r\n4. CNN is actually simple enough to be run on ordinary computer in reasonable time\r\n5. I added some useful callback functions:\r\nEarlyStopping - to stop early after loss stop decreasing\r\nModelCheckpoint - save best weights and restore them for minimum loss after \"fit\" ends. Some kind of XGBoost's best_ntree_limit\r\n\r\nNotes:\r\n\r\n 1. Crossfold score is actually much lower than leaderboard one \r\n 2. In most cases best loss achieved right after first epoch \r\n 3. I feel like this line: \r\n\r\n    train_data = train_data.reshape(train_data.shape[0], 1,\r\n        img_rows, img_cols) \r\n\r\nshould be replaced with some other function. It would be good if someone propose best replacement.\r\n\r\n  4. \"batch_size\" quite strongly affects learning process \r\n  5. It seems KFold = 26 is best for this problem.\r\n\r\nRunning \"as is\" from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.",
    "118794": "[quote=ZFTurbo;118398]\r\n\r\nI posted latest version of my Keras code (1st topic updated):\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/run_keras_cv_drivers_v2.py\r\n\r\nI used some ideas from this post:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20482/getting-started-with-nolearn-lasagne\r\n\r\n1. Code randomly rotate images +-10 degrees\r\n2. Code uses the same CNN structure from mentioned post with Dropout layers after each Conv/Pool layer. This allows to slightly reduce overfit.\r\n3. Code uses 64x64 pixel grayscale images\r\n4. CNN is actually simple enough to be run on ordinary computer in reasonable time\r\n5. I added some useful callback functions:\r\nEarlyStopping - to stop early after loss stop decreasing\r\nModelCheckpoint - save best weights and restore them for minimum loss after \"fit\" ends. Some kind of XGBoost's best_ntree_limit\r\n\r\nNotes:\r\n\r\n 1. Crossfold score is actually much lower than leaderboard one \r\n 2. In most cases best loss achieved right after first epoch \r\n 3. I feel like this line: \r\n\r\n    train_data = train_data.reshape(train_data.shape[0], 1,\r\n        img_rows, img_cols) \r\n\r\nshould be replaced with some other function. It would be good if someone propose best replacement.\r\n\r\n  4. \"batch_size\" quite strongly affects learning process \r\n  5. It seems KFold = 26 is best for this problem.\r\n\r\nRunning \"as is\" from repository will generate submission with validation loss around 0.27 and LB score ~1.03. I was able to generate solutions with 0.85-0.9 score with same code on different parameters. But I wasn't experimented much.\r\n\r\n[/quote]\r\n1) Did you try increasing the number of convolutional layers? I added a 32 x 32 convolutional layer and a 64 X 64 convolutional layer only to get worse results(LB: 0.88)\r\n2) Why don't you try adding more dense layers? \r\n\r\nThanks",
    "129029": "Thanks for sharing! It really helped me to get started.",
    "129338": "I've added my code to run pretrained VGG16 Neural Net:\r\nhttps://github.com/ZFTurbo/KAGGLE_DISTRACTED_DRIVER/blob/master/kaggle_distracted_drivers_vgg16.py\r\n\r\nThis code allows to get **0.20646** on LB.\r\n\r\nRequirements:\r\n\r\n 1. 16 GB of RAM (with some swap at preprocessing peaks). Code splits \"test\" in 5 chunks, so it allowed to reduce memory usage.\r\n 2. Powerful NVIDIA GPU. It takes around a day on 980Ti 6GB.\r\n 3. **Important**: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.\r\n 4. Weights: https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n\r\nNotes:\r\n\r\n 1. Code doesn't use data augmentation\r\n 2. Code uses totally random test split (doesn't use split by drivers). So don't look at local validation loss, it will be around ~0.01 at the end.",
    "129350": "ZFTurbo :  just wanted to confirm that during training of the model the data is split based on drivers ( was looking at your code for function  run_cross_validation_create_models)",
    "129362": "**Hetal Chandaria** it initially splitted by drivers, but later shuffled (look for the following code):\r\n\r\n    # Shuffle experiment START !!!\r\n    perm = permutation(len(train_target))\r\n    train_data = train_data[perm]\r\n    train_target = train_target[perm]\r\n    # Shuffle experiment END !!!",
    "129369": "Thanks for sharing!\r\n\r\nSo the key is version of keras? ToT",
    "129370": "**Ferris**, yes. At least this code works much better on outdated version.",
    "129483": "Thanks for the efforts, very educative!\r\n\r\nOne question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day).",
    "129484": "[quote=liviu;129483]\r\n\r\nThanks for the efforts, very educative!\r\n\r\nOne question if anyone has any thoughts on this: Is it possible to be competitive in this competition without using GPU? All CNNs I tried on CPU are too slow to be practical (around 4 epochs/day). \r\n\r\n[/quote]\r\n\r\nI suspect not. I think you'll find that anything scoring 0.5 or better would take literally weeks to train on a CPU-based setup, and you would need to have chosen the meta-params correctly. I got down to 0.89 on a CPU rig after a few iterations, and that took 2.5 days training time.",
    "129485": "Not what I was hoping to hear, but it gives me the green light to move on :P. Thanks!",
    "129519": "Hi, @ZFTurbo,\r\nThank you for your nice code sharing. Is there any reason for not split by driver this time?",
    "129520": "**Li Li**, for some reason score while splitting by drivers was worse.",
    "129531": "[quote=ZFTurbo;129338]\r\n 3. **Important**: You need to use Keras 0.2.0. Latest Keras versions work totally different. I didn't dive into it, but on latest Keras version this code gives much lower score.\r\n[/quote]\r\n\r\nCan you post score on latest Keras version you tried? I tried on 1.0.6 and got 0.23 with 8 folds with some modifications (GAP instead of Dense Layers for memory optimization). I will try to debug your code with new Keras version, but wander if you have ideas why its different and what score you got.",
    "129534": "**kyv**, It was run by my teammates. And they had problems with my code. I see 0.29 and 0.27 in list (it was calculated on 1.0.* version).\r\n\r\nI don't have ideas, except that learning process is somehow different in 0.2.0 and 1.0.*. I probably rerun this code on latest version after contest ends to check the difference myself.\r\n\r\nI also notice difference in loss in Nerve contest. When I debug Kernel at local machine, loss was much different from version on Kaggle servers.",
    "129642": "Could be because of different cudnn or theano versions. You also have to clean theano cache once in a while - a for sure if you change something like cuda drivers, cudnn or theano version.",
    "129644": "kyv I'm sorry for not knowing. What is \"GAP instead of Dense Layers for memory optimization\"?",
    "129737": "[quote=Keiku;129644]\r\n\r\n@kyv I'm sorry for not knowing. What is \"GAP instead of Dense Layers for memory optimization\"?\r\n\r\n[/quote]\r\nI think that's global average pooling layer. You can find this concept in paper:Network In Network",
    "129740": "Tian Zhou Thank you for telling me about.",
    "129759": "[quote=Keiku;129644]\r\n\r\n@kyv I'm sorry for not knowing. What is \"GAP instead of Dense Layers for memory optimization\"?\r\n\r\n[/quote]\r\nYes its AveragePooling layer with stride equal to image size and channels equal to number of classes followed by softmax .\r\nBy doing this you reduce model size by 80% (as most parameters are in dense layers) and usually you don't sacrifice model accuracy. New models often uses this approach.",
    "129763": "kyv Thank you for your advice. I will try with new Keras version for future reference, too.",
    "129766": "**kyv**, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )",
    "129770": "[quote=ZFTurbo;129766]\r\n\r\n**kyv**, I'm curious about GAP. So you just modified pretrained VGG16 replacing all convolution part with some small layers? Does it retrained as good as initial VGG16? May be do you have model and weights for it? )\r\n[/quote]\r\n\r\nI replaced last dense layers with\r\n\r\n     for _ in range(conv_layers):\r\n        model.add(Convolution2D(nb_filter, nb_cols, nb_cols, \"relu\")) \r\n    model.add(AveragePooling2D((7, 7)))\r\n    model.add(Flatten())\r\n    model.add(Dropout(dropout))\r\n    model.add(Dense(10, W_regularizer=REG, b_regularizer=REG))\r\n    model.add(Activation(\"softmax\"))\r\n\r\nconv_layers=3 and nb_filter=50 to 100.\r\nIt was comparable with original vgg16 on fine-tuning convergence speed and eventually I stopped using dense layers as it was limiting me in batch size.",
    "129773": "**kyv**, thank you. Does it decrease complexity of computations? Or allows to save memory? I see there are pretty much new layers with many filters. Can you clarify it?",
    "129874": "ZFTurbo,\r\nI got initial config from here\r\nhttps://github.com/tdeboissiere/VGG16CAM-keras/blob/master/VGGCAM-keras.py - so it can be just Conv, Avgpool + flatten+softmax. \r\nand made top classifier more complex just because I trained it separately with parameter optimizations to see how far can I get without fine-tuning (78% accuracy) and then combined with main vgg16.\r\n\r\nMy model size was 62mb compared to 550mb. It was slightly faster 500 - 600sec per epoch compared to 700 - 800 with full vgg16 and increased batch_size from 8 to 16 (haven't tried more)"
  },
  "source": "meta"
}