{
  "id": 20129,
  "title": "Cloud GPU Starter project",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20129",
  "author_name": "Jim Fleming",
  "post_date": "2016-04-14T16:21:14.853000",
  "votes": 10,
  "comment_count": 19,
  "views": 2863,
  "content": "<p>Hi everyone! I saw a few people asking about GPUs so I put these together:</p>\n\n<ul>\n<li><a href=\"https://github.com/fomorians/distracted-drivers-tf\">https://github.com/fomorians/distracted-drivers-tf</a></li>\n<li><a href=\"https://github.com/fomorians/distracted-drivers-keras\">https://github.com/fomorians/distracted-drivers-keras</a></li>\n</ul>\n\n<p>It uses TensorFlow (or Keras) and setup for training in the cloud using Fomoro. *Disclaimer, Fomoro is something I built.</p>",
  "messages": [
    {
      "id": 114886,
      "postDate": "2016-04-14T16:21:14.853Z",
      "content": "<p>Hi everyone! I saw a few people asking about GPUs so I put these together:</p>\n\n<ul>\n<li><a href=\"https://github.com/fomorians/distracted-drivers-tf\">https://github.com/fomorians/distracted-drivers-tf</a></li>\n<li><a href=\"https://github.com/fomorians/distracted-drivers-keras\">https://github.com/fomorians/distracted-drivers-keras</a></li>\n</ul>\n\n<p>It uses TensorFlow (or Keras) and setup for training in the cloud using Fomoro. *Disclaimer, Fomoro is something I built.</p>",
      "rawMarkdown": "Hi everyone! I saw a few people asking about GPUs so I put these together:\r\n\r\n- https://github.com/fomorians/distracted-drivers-tf\r\n- https://github.com/fomorians/distracted-drivers-keras\r\n\r\nIt uses TensorFlow (or Keras) and setup for training in the cloud using Fomoro. *Disclaimer, Fomoro is something I built.",
      "votes": 9
    },
    {
      "id": 115951,
      "postDate": "2016-04-20T23:58:54.743Z",
      "content": "<p>@asmith26 The default format of <code>imread</code> is something like <code>(WIDTH, HEIGHT, NUM_CHANNELS)</code>. Keras prefers <code>(NUM_CHANNELS, WIDTH, HEIGHT)</code> so we swap the axes.</p>",
      "rawMarkdown": "@asmith26 The default format of `imread` is something like `(WIDTH, HEIGHT, NUM_CHANNELS)`. Keras prefers `(NUM_CHANNELS, WIDTH, HEIGHT)` so we swap the axes.",
      "votes": 1
    },
    {
      "id": 115738,
      "postDate": "2016-04-19T20:25:32.833Z",
      "content": "<p>I believe the problem is in prep_dataset.py. Here's a git diff showing the fix that works for me:</p>\n\n<pre><code>-        paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\n+        #paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\n         driver_ids_group = driver_imgs_grouped.get_group('c{}'.format(j))\n+        paths = base + '/c{}/'.format(j) + driver_ids_group.img\n</code></pre>",
      "rawMarkdown": "I believe the problem is in prep_dataset.py. Here's a git diff showing the fix that works for me:\r\n\r\n    -        paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\r\n    +        #paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\r\n             driver_ids_group = driver_imgs_grouped.get_group('c{}'.format(j))\r\n    +        paths = base + '/c{}/'.format(j) + driver_ids_group.img\r\n\r\n",
      "votes": 1
    },
    {
      "id": 115005,
      "postDate": "2016-04-15T15:11:05.083Z",
      "content": "<p>That's a good question! The biggest difference is that the TF version is using more folds and epochs, which will make a significant difference. Just change 3 =&gt; 8 for max folds and 10 =&gt; 20 for epochs. I'll push that change now to make them more similar.</p>\n\n<p>The models themselves are nearly identical, however, there are a couple of small differences for simplicity:</p>\n\n<ul>\n<li><p>The gain parameter of the initialization. I'm using <a href=\"http://arxiv.org/abs/1502.01852\">He initialization</a> for the weights in both models where the std deviation of the normal distribution is <code>gain * sqrt(1 / fan_in)</code>. Keras, by default, specifies &quot;he_normal&quot; and no gain parameter. Ideally, you want a gain of <code>sqrt(2)</code> for ReLU and <code>1</code> for sigmoid.</p></li>\n<li><p>The TF version is using a truncated normal distribution to reduce dying connections. I'm not sure what Keras is using internally.</p></li>\n</ul>",
      "rawMarkdown": "That's a good question! The biggest difference is that the TF version is using more folds and epochs, which will make a significant difference. Just change 3 => 8 for max folds and 10 => 20 for epochs. I'll push that change now to make them more similar.\r\n\r\nThe models themselves are nearly identical, however, there are a couple of small differences for simplicity:\r\n\r\n- The gain parameter of the initialization. I'm using [He initialization][1] for the weights in both models where the std deviation of the normal distribution is `gain * sqrt(1 / fan_in)`. Keras, by default, specifies \"he_normal\" and no gain parameter. Ideally, you want a gain of `sqrt(2)` for ReLU and `1` for sigmoid.\r\n\r\n- The TF version is using a truncated normal distribution to reduce dying connections. I'm not sure what Keras is using internally.\r\n\r\n  [1]: http://arxiv.org/abs/1502.01852",
      "votes": 2
    },
    {
      "id": 115770,
      "postDate": "2016-04-20T01:08:58.227Z",
      "content": "<p>To answer my own question above, the driver ID problem does seem to explain why my cross-validation scores were so much lower than my leaderboard scores.</p>\n\n<p>After I applied the fix suggested by @SpammySmith, I got 1.46 from cross-validation and 1.24 on the leaderboard. Still a little puzzling why LB &lt; CV, but maybe there's reasons.</p>\n\n<p>More notable to me is that the model appears to be badly overfitting. At least, that's my explanation for why the val_loss climbs steadily with every epoch while the training loss plummets to 0.04 or so.</p>\n\n<p>All of the above is for the Keras version. It's also still a puzzle why the TensorFlow version is substantially more accurate than Keras, when you would expect them to be about the same.</p>",
      "rawMarkdown": "To answer my own question above, the driver ID problem does seem to explain why my cross-validation scores were so much lower than my leaderboard scores.\r\n\r\nAfter I applied the fix suggested by @SpammySmith, I got 1.46 from cross-validation and 1.24 on the leaderboard. Still a little puzzling why LB < CV, but maybe there's reasons.\r\n\r\nMore notable to me is that the model appears to be badly overfitting. At least, that's my explanation for why the val_loss climbs steadily with every epoch while the training loss plummets to 0.04 or so.\r\n\r\nAll of the above is for the Keras version. It's also still a puzzle why the TensorFlow version is substantially more accurate than Keras, when you would expect them to be about the same.\r\n"
    },
    {
      "id": 115735,
      "postDate": "2016-04-19T19:01:23.037Z",
      "content": "<p>Hi @jeblist, email me with what's not working: jim [at] fomoro [dot] com</p>\n\n<p>Thanks @zfturbo, makes sense. That's what I get for rushing it out ;) I'll push a fix sometime today.</p>",
      "rawMarkdown": "Hi @jeblist, email me with what's not working: jim [at] fomoro [dot] com\r\n\r\nThanks @zfturbo, makes sense. That's what I get for rushing it out ;) I'll push a fix sometime today."
    },
    {
      "id": 115726,
      "postDate": "2016-04-19T18:30:19.483Z",
      "content": "<p>help setting it up !!\nI can't start</p>",
      "rawMarkdown": " help setting it up !!\r\nI can't start"
    },
    {
      "id": 115671,
      "postDate": "2016-04-19T15:15:32.977Z",
      "content": "<p>Jim Fleming, my example just to show that two images for the same driver ID has different driver on them.</p>\n\n<p>I insert it in your &quot;main.py&quot; code after line:</p>\n\n<pre><code>for train_index, valid_index in LabelShuffleSplit(driver_indices, n_iter=MAX_FOLDS, test_size=0.2, random_state=67):\n</code></pre>",
      "rawMarkdown": "Jim Fleming, my example just to show that two images for the same driver ID has different driver on them.\r\n\r\nI insert it in your \"main.py\" code after line:\r\n\r\n    for train_index, valid_index in LabelShuffleSplit(driver_indices, n_iter=MAX_FOLDS, test_size=0.2, random_state=67):"
    },
    {
      "id": 115670,
      "postDate": "2016-04-19T15:14:51.260Z",
      "content": "<p>[quote=ZFTurbo;115616]I'm not sure, but you probably have a bug in &quot;driver_indices&quot; in your Keras code. Are you sure they are correct?[/quote]</p>\n\n<p>I wonder if this would explain why my CV score (0.04) is so much lower than my LB score (1.26).</p>",
      "rawMarkdown": "[quote=ZFTurbo;115616]I'm not sure, but you probably have a bug in \"driver_indices\" in your Keras code. Are you sure they are correct?[/quote]\r\n\r\nI wonder if this would explain why my CV score (0.04) is so much lower than my LB score (1.26)."
    },
    {
      "id": 115665,
      "postDate": "2016-04-19T14:54:07.473Z",
      "content": "<p>@zfturbo hmm, I'll look into that. I'm not positive one way or the other. I'm not using open cv (scipy is easier to install). Where is that snippet from?</p>\n\n<p>@inoryy thanks, that makes sense. I figured it was a sensible default. Part of me hoped it would be determined by the activation function. :)</p>",
      "rawMarkdown": "@zfturbo hmm, I'll look into that. I'm not positive one way or the other. I'm not using open cv (scipy is easier to install). Where is that snippet from?\r\n\r\n@inoryy thanks, that makes sense. I figured it was a sensible default. Part of me hoped it would be determined by the activation function. :)"
    },
    {
      "id": 115655,
      "postDate": "2016-04-19T14:23:56.737Z",
      "content": "<p>[quote=Jim Fleming;115005]\n- The gain parameter of the initialization. I'm using [He initialization][1] for the weights in both models where the std deviation of the normal distribution is <code>gain * sqrt(1 / fan_in)</code>. Keras, by default, specifies &quot;he_normal&quot; and no gain parameter. Ideally, you want a gain of <code>sqrt(2)</code> for ReLU and <code>1</code> for sigmoid.\n[/quote]</p>\n\n<p>FYI Keras uses sqrt(2) gain by default for  &quot;he_normal&quot;\n<a href=\"https://github.com/fchollet/keras/blob/master/keras/initializations.py#L66\">https://github.com/fchollet/keras/blob/master/keras/initializations.py#L66</a></p>",
      "rawMarkdown": "[quote=Jim Fleming;115005]\r\n- The gain parameter of the initialization. I'm using [He initialization][1] for the weights in both models where the std deviation of the normal distribution is `gain * sqrt(1 / fan_in)`. Keras, by default, specifies \"he_normal\" and no gain parameter. Ideally, you want a gain of `sqrt(2)` for ReLU and `1` for sigmoid.\r\n[/quote]\r\n\r\nFYI Keras uses sqrt(2) gain by default for  \"he_normal\"\r\nhttps://github.com/fchollet/keras/blob/master/keras/initializations.py#L66"
    },
    {
      "id": 115616,
      "postDate": "2016-04-19T10:39:21.083Z",
      "content": "<p>I'm not sure, but you probably have a bug in &quot;driver_indices&quot; in your Keras code. Are you sure they are correct?</p>\n\n<pre><code>for i in range(len(driver_indices)):\n        print(driver_indices[i])\n        img = np.array(X_train_raw[i, ...]).swapaxes(0, 2)\n        img = cv2.resize(img, (500, 500))\n        cv2.imshow('name', img)\n        cv2.waitKey(0)\n        cv2.destroyAllWindows()\n</code></pre>\n\n<p>There are obviously two different drivers for the same driver ID. </p>",
      "rawMarkdown": "I'm not sure, but you probably have a bug in \"driver_indices\" in your Keras code. Are you sure they are correct?\r\n\r\n    for i in range(len(driver_indices)):\r\n            print(driver_indices[i])\r\n            img = np.array(X_train_raw[i, ...]).swapaxes(0, 2)\r\n            img = cv2.resize(img, (500, 500))\r\n            cv2.imshow('name', img)\r\n            cv2.waitKey(0)\r\n            cv2.destroyAllWindows()\r\n\r\nThere are obviously two different drivers for the same driver ID. "
    },
    {
      "id": 115326,
      "postDate": "2016-04-18T03:44:52.970Z",
      "content": "<p>A geometric mean[0] is used to average the score and predictions from each fold. The function <code>calc_geom</code> calculates the mean score, while <code>calc_geom_arr</code> calculates the mean predictions.</p>\n\n<p>[0] In mathematics, the geometric mean is a type of mean or average, which indicates the central tendency or typical value of a set of numbers by using the product of their values (as opposed to the arithmetic mean which uses their sum). (<a href=\"https://en.wikipedia.org/wiki/Geometric_mean\">https://en.wikipedia.org/wiki/Geometric_mean</a>)</p>",
      "rawMarkdown": "A geometric mean[0] is used to average the score and predictions from each fold. The function `calc_geom` calculates the mean score, while `calc_geom_arr` calculates the mean predictions.\r\n\r\n[0] In mathematics, the geometric mean is a type of mean or average, which indicates the central tendency or typical value of a set of numbers by using the product of their values (as opposed to the arithmetic mean which uses their sum). (https://en.wikipedia.org/wiki/Geometric_mean)"
    },
    {
      "id": 115063,
      "postDate": "2016-04-16T00:04:16.817Z",
      "content": "<p>Interesting. I'd love to know if you find the difference. I'll post here if I find it as well.</p>\n\n<p>I tried to keep the implementations as similar as possible so it's likely a small detail somewhere.</p>",
      "rawMarkdown": "Interesting. I'd love to know if you find the difference. I'll post here if I find it as well.\r\n\r\nI tried to keep the implementations as similar as possible so it's likely a small detail somewhere."
    },
    {
      "id": 115060,
      "postDate": "2016-04-15T23:35:06.927Z",
      "content": "<p>I get 1.26 on the leaderboard with those changes versus 1.04 for the TF version.</p>\n\n<p>I'll dig through the code some more to see if I spot anything else. These differences are a bit of a side issue of course, but I am just learning both Keras and TF and it might be instructive to understand if the differences between them are real.</p>",
      "rawMarkdown": "I get 1.26 on the leaderboard with those changes versus 1.04 for the TF version.\r\n\r\nI'll dig through the code some more to see if I spot anything else. These differences are a bit of a side issue of course, but I am just learning both Keras and TF and it might be instructive to understand if the differences between them are real."
    },
    {
      "id": 114973,
      "postDate": "2016-04-15T06:54:23.740Z",
      "content": "<p>I've run the two versions as-is without any tweaking. They give me quite different results on the leaderboard (1.04 for TensorFlow, 1. 33 for Keras).</p>\n\n<p>Is that expected? Why is the Keras version worse?</p>",
      "rawMarkdown": "I've run the two versions as-is without any tweaking. They give me quite different results on the leaderboard (1.04 for TensorFlow, 1. 33 for Keras).\r\n\r\nIs that expected? Why is the Keras version worse?"
    },
    {
      "id": 114950,
      "postDate": "2016-04-14T23:23:33.570Z",
      "content": "<p>Here's the Keras version! Let me know if you need any help setting it up.</p>\n\n<p><a href=\"https://github.com/fomorians/distracted-drivers-keras\">https://github.com/fomorians/distracted-drivers-keras</a></p>",
      "rawMarkdown": "Here's the Keras version! Let me know if you need any help setting it up.\r\n\r\nhttps://github.com/fomorians/distracted-drivers-keras"
    },
    {
      "id": 114946,
      "postDate": "2016-04-14T22:34:52.423Z",
      "content": "<p>Very nice, thanks!</p>\n\n<p>Looking forward to the Keras version.</p>",
      "rawMarkdown": "Very nice, thanks!\r\n\r\nLooking forward to the Keras version."
    },
    {
      "id": 115940,
      "postDate": "2016-04-20T22:38:45Z",
      "content": "<p>Why do we <strong><code>img=img.swapaxes(2,0)</code></strong> in <em>prep_dataset.py</em>? I imagine this is because Keras requires images in a certain format, but I can't find the relevant documentation (if anyone can find it, please post here thanks).</p>",
      "rawMarkdown": "Why do we **`img=img.swapaxes(2,0)`** in *prep_dataset.py*? I imagine this is because Keras requires images in a certain format, but I can't find the relevant documentation (if anyone can find it, please post here thanks).",
      "isDeleted": true
    },
    {
      "id": 115187,
      "postDate": "2016-04-16T23:42:25.987Z",
      "content": "<p>@Jim Fleming - Thanks also for this ;-) I understand the calculations, but why do you calculate &quot;calc_geom&quot; and &quot;calc_geom_arr&quot;? Many thanks for all your help in advance!!!</p>",
      "rawMarkdown": "@Jim Fleming - Thanks also for this ;-) I understand the calculations, but why do you calculate \"calc_geom\" and \"calc_geom_arr\"? Many thanks for all your help in advance!!!",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 115951,
      "author_name": "Jim Fleming",
      "author_url": "",
      "post_date": "2016-04-20T23:58:54.743000",
      "content": "<p>@asmith26 The default format of <code>imread</code> is something like <code>(WIDTH, HEIGHT, NUM_CHANNELS)</code>. Keras prefers <code>(NUM_CHANNELS, WIDTH, HEIGHT)</code> so we swap the axes.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 115738,
      "author_name": "SpammySmith",
      "author_url": "",
      "post_date": "2016-04-19T20:25:32.833000",
      "content": "<p>I believe the problem is in prep_dataset.py. Here's a git diff showing the fix that works for me:</p>\n\n<pre><code>-        paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\n+        #paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\n         driver_ids_group = driver_imgs_grouped.get_group('c{}'.format(j))\n+        paths = base + '/c{}/'.format(j) + driver_ids_group.img\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 115005,
      "author_name": "Jim Fleming",
      "author_url": "",
      "post_date": "2016-04-15T15:11:05.083000",
      "content": "<p>That's a good question! The biggest difference is that the TF version is using more folds and epochs, which will make a significant difference. Just change 3 =&gt; 8 for max folds and 10 =&gt; 20 for epochs. I'll push that change now to make them more similar.</p>\n\n<p>The models themselves are nearly identical, however, there are a couple of small differences for simplicity:</p>\n\n<ul>\n<li><p>The gain parameter of the initialization. I'm using <a href=\"http://arxiv.org/abs/1502.01852\">He initialization</a> for the weights in both models where the std deviation of the normal distribution is <code>gain * sqrt(1 / fan_in)</code>. Keras, by default, specifies &quot;he_normal&quot; and no gain parameter. Ideally, you want a gain of <code>sqrt(2)</code> for ReLU and <code>1</code> for sigmoid.</p></li>\n<li><p>The TF version is using a truncated normal distribution to reduce dying connections. I'm not sure what Keras is using internally.</p></li>\n</ul>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 115770,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "2016-04-20T01:08:58.227000",
      "content": "<p>To answer my own question above, the driver ID problem does seem to explain why my cross-validation scores were so much lower than my leaderboard scores.</p>\n\n<p>After I applied the fix suggested by @SpammySmith, I got 1.46 from cross-validation and 1.24 on the leaderboard. Still a little puzzling why LB &lt; CV, but maybe there's reasons.</p>\n\n<p>More notable to me is that the model appears to be badly overfitting. At least, that's my explanation for why the val_loss climbs steadily with every epoch while the training loss plummets to 0.04 or so.</p>\n\n<p>All of the above is for the Keras version. It's also still a puzzle why the TensorFlow version is substantially more accurate than Keras, when you would expect them to be about the same.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115735,
      "author_name": "Jim Fleming",
      "author_url": "",
      "post_date": "2016-04-19T19:01:23.037000",
      "content": "<p>Hi @jeblist, email me with what's not working: jim [at] fomoro [dot] com</p>\n\n<p>Thanks @zfturbo, makes sense. That's what I get for rushing it out ;) I'll push a fix sometime today.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115726,
      "author_name": "Jeblist",
      "author_url": "",
      "post_date": "2016-04-19T18:30:19.483000",
      "content": "<p>help setting it up !!\nI can't start</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115671,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2016-04-19T15:15:32.977000",
      "content": "<p>Jim Fleming, my example just to show that two images for the same driver ID has different driver on them.</p>\n\n<p>I insert it in your &quot;main.py&quot; code after line:</p>\n\n<pre><code>for train_index, valid_index in LabelShuffleSplit(driver_indices, n_iter=MAX_FOLDS, test_size=0.2, random_state=67):\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115670,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "2016-04-19T15:14:51.260000",
      "content": "<p>[quote=ZFTurbo;115616]I'm not sure, but you probably have a bug in &quot;driver_indices&quot; in your Keras code. Are you sure they are correct?[/quote]</p>\n\n<p>I wonder if this would explain why my CV score (0.04) is so much lower than my LB score (1.26).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115665,
      "author_name": "Jim Fleming",
      "author_url": "",
      "post_date": "2016-04-19T14:54:07.473000",
      "content": "<p>@zfturbo hmm, I'll look into that. I'm not positive one way or the other. I'm not using open cv (scipy is easier to install). Where is that snippet from?</p>\n\n<p>@inoryy thanks, that makes sense. I figured it was a sensible default. Part of me hoped it would be determined by the activation function. :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115655,
      "author_name": "Roman Ring",
      "author_url": "",
      "post_date": "2016-04-19T14:23:56.737000",
      "content": "<p>[quote=Jim Fleming;115005]\n- The gain parameter of the initialization. I'm using [He initialization][1] for the weights in both models where the std deviation of the normal distribution is <code>gain * sqrt(1 / fan_in)</code>. Keras, by default, specifies &quot;he_normal&quot; and no gain parameter. Ideally, you want a gain of <code>sqrt(2)</code> for ReLU and <code>1</code> for sigmoid.\n[/quote]</p>\n\n<p>FYI Keras uses sqrt(2) gain by default for  &quot;he_normal&quot;\n<a href=\"https://github.com/fchollet/keras/blob/master/keras/initializations.py#L66\">https://github.com/fchollet/keras/blob/master/keras/initializations.py#L66</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115616,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2016-04-19T10:39:21.083000",
      "content": "<p>I'm not sure, but you probably have a bug in &quot;driver_indices&quot; in your Keras code. Are you sure they are correct?</p>\n\n<pre><code>for i in range(len(driver_indices)):\n        print(driver_indices[i])\n        img = np.array(X_train_raw[i, ...]).swapaxes(0, 2)\n        img = cv2.resize(img, (500, 500))\n        cv2.imshow('name', img)\n        cv2.waitKey(0)\n        cv2.destroyAllWindows()\n</code></pre>\n\n<p>There are obviously two different drivers for the same driver ID. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115326,
      "author_name": "Jim Fleming",
      "author_url": "",
      "post_date": "2016-04-18T03:44:52.970000",
      "content": "<p>A geometric mean[0] is used to average the score and predictions from each fold. The function <code>calc_geom</code> calculates the mean score, while <code>calc_geom_arr</code> calculates the mean predictions.</p>\n\n<p>[0] In mathematics, the geometric mean is a type of mean or average, which indicates the central tendency or typical value of a set of numbers by using the product of their values (as opposed to the arithmetic mean which uses their sum). (<a href=\"https://en.wikipedia.org/wiki/Geometric_mean\">https://en.wikipedia.org/wiki/Geometric_mean</a>)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115063,
      "author_name": "Jim Fleming",
      "author_url": "",
      "post_date": "2016-04-16T00:04:16.817000",
      "content": "<p>Interesting. I'd love to know if you find the difference. I'll post here if I find it as well.</p>\n\n<p>I tried to keep the implementations as similar as possible so it's likely a small detail somewhere.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115060,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "2016-04-15T23:35:06.927000",
      "content": "<p>I get 1.26 on the leaderboard with those changes versus 1.04 for the TF version.</p>\n\n<p>I'll dig through the code some more to see if I spot anything else. These differences are a bit of a side issue of course, but I am just learning both Keras and TF and it might be instructive to understand if the differences between them are real.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 114973,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "2016-04-15T06:54:23.740000",
      "content": "<p>I've run the two versions as-is without any tweaking. They give me quite different results on the leaderboard (1.04 for TensorFlow, 1. 33 for Keras).</p>\n\n<p>Is that expected? Why is the Keras version worse?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 114950,
      "author_name": "Jim Fleming",
      "author_url": "",
      "post_date": "2016-04-14T23:23:33.570000",
      "content": "<p>Here's the Keras version! Let me know if you need any help setting it up.</p>\n\n<p><a href=\"https://github.com/fomorians/distracted-drivers-keras\">https://github.com/fomorians/distracted-drivers-keras</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 114946,
      "author_name": "gauss256",
      "author_url": "",
      "post_date": "2016-04-14T22:34:52.423000",
      "content": "<p>Very nice, thanks!</p>\n\n<p>Looking forward to the Keras version.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115940,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-20T22:38:45",
      "content": "<p>Why do we <strong><code>img=img.swapaxes(2,0)</code></strong> in <em>prep_dataset.py</em>? I imagine this is because Keras requires images in a certain format, but I can't find the relevant documentation (if anyone can find it, please post here thanks).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 115187,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-16T23:42:25.987000",
      "content": "<p>@Jim Fleming - Thanks also for this ;-) I understand the calculations, but why do you calculate &quot;calc_geom&quot; and &quot;calc_geom_arr&quot;? Many thanks for all your help in advance!!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "114886": "Hi everyone! I saw a few people asking about GPUs so I put these together:\r\n\r\n- https://github.com/fomorians/distracted-drivers-tf\r\n- https://github.com/fomorians/distracted-drivers-keras\r\n\r\nIt uses TensorFlow (or Keras) and setup for training in the cloud using Fomoro. *Disclaimer, Fomoro is something I built.",
    "115951": "@asmith26 The default format of `imread` is something like `(WIDTH, HEIGHT, NUM_CHANNELS)`. Keras prefers `(NUM_CHANNELS, WIDTH, HEIGHT)` so we swap the axes.",
    "115738": "I believe the problem is in prep_dataset.py. Here's a git diff showing the fix that works for me:\r\n\r\n    -        paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\r\n    +        #paths = glob.glob(os.path.join(base, 'c{}/*.jpg'.format(j)))\r\n             driver_ids_group = driver_imgs_grouped.get_group('c{}'.format(j))\r\n    +        paths = base + '/c{}/'.format(j) + driver_ids_group.img\r\n\r\n",
    "115005": "That's a good question! The biggest difference is that the TF version is using more folds and epochs, which will make a significant difference. Just change 3 => 8 for max folds and 10 => 20 for epochs. I'll push that change now to make them more similar.\r\n\r\nThe models themselves are nearly identical, however, there are a couple of small differences for simplicity:\r\n\r\n- The gain parameter of the initialization. I'm using [He initialization][1] for the weights in both models where the std deviation of the normal distribution is `gain * sqrt(1 / fan_in)`. Keras, by default, specifies \"he_normal\" and no gain parameter. Ideally, you want a gain of `sqrt(2)` for ReLU and `1` for sigmoid.\r\n\r\n- The TF version is using a truncated normal distribution to reduce dying connections. I'm not sure what Keras is using internally.\r\n\r\n  [1]: http://arxiv.org/abs/1502.01852",
    "115770": "To answer my own question above, the driver ID problem does seem to explain why my cross-validation scores were so much lower than my leaderboard scores.\r\n\r\nAfter I applied the fix suggested by @SpammySmith, I got 1.46 from cross-validation and 1.24 on the leaderboard. Still a little puzzling why LB < CV, but maybe there's reasons.\r\n\r\nMore notable to me is that the model appears to be badly overfitting. At least, that's my explanation for why the val_loss climbs steadily with every epoch while the training loss plummets to 0.04 or so.\r\n\r\nAll of the above is for the Keras version. It's also still a puzzle why the TensorFlow version is substantially more accurate than Keras, when you would expect them to be about the same.\r\n",
    "115735": "Hi @jeblist, email me with what's not working: jim [at] fomoro [dot] com\r\n\r\nThanks @zfturbo, makes sense. That's what I get for rushing it out ;) I'll push a fix sometime today.",
    "115726": " help setting it up !!\r\nI can't start",
    "115671": "Jim Fleming, my example just to show that two images for the same driver ID has different driver on them.\r\n\r\nI insert it in your \"main.py\" code after line:\r\n\r\n    for train_index, valid_index in LabelShuffleSplit(driver_indices, n_iter=MAX_FOLDS, test_size=0.2, random_state=67):",
    "115670": "[quote=ZFTurbo;115616]I'm not sure, but you probably have a bug in \"driver_indices\" in your Keras code. Are you sure they are correct?[/quote]\r\n\r\nI wonder if this would explain why my CV score (0.04) is so much lower than my LB score (1.26).",
    "115665": "@zfturbo hmm, I'll look into that. I'm not positive one way or the other. I'm not using open cv (scipy is easier to install). Where is that snippet from?\r\n\r\n@inoryy thanks, that makes sense. I figured it was a sensible default. Part of me hoped it would be determined by the activation function. :)",
    "115655": "[quote=Jim Fleming;115005]\r\n- The gain parameter of the initialization. I'm using [He initialization][1] for the weights in both models where the std deviation of the normal distribution is `gain * sqrt(1 / fan_in)`. Keras, by default, specifies \"he_normal\" and no gain parameter. Ideally, you want a gain of `sqrt(2)` for ReLU and `1` for sigmoid.\r\n[/quote]\r\n\r\nFYI Keras uses sqrt(2) gain by default for  \"he_normal\"\r\nhttps://github.com/fchollet/keras/blob/master/keras/initializations.py#L66",
    "115616": "I'm not sure, but you probably have a bug in \"driver_indices\" in your Keras code. Are you sure they are correct?\r\n\r\n    for i in range(len(driver_indices)):\r\n            print(driver_indices[i])\r\n            img = np.array(X_train_raw[i, ...]).swapaxes(0, 2)\r\n            img = cv2.resize(img, (500, 500))\r\n            cv2.imshow('name', img)\r\n            cv2.waitKey(0)\r\n            cv2.destroyAllWindows()\r\n\r\nThere are obviously two different drivers for the same driver ID. ",
    "115326": "A geometric mean[0] is used to average the score and predictions from each fold. The function `calc_geom` calculates the mean score, while `calc_geom_arr` calculates the mean predictions.\r\n\r\n[0] In mathematics, the geometric mean is a type of mean or average, which indicates the central tendency or typical value of a set of numbers by using the product of their values (as opposed to the arithmetic mean which uses their sum). (https://en.wikipedia.org/wiki/Geometric_mean)",
    "115063": "Interesting. I'd love to know if you find the difference. I'll post here if I find it as well.\r\n\r\nI tried to keep the implementations as similar as possible so it's likely a small detail somewhere.",
    "115060": "I get 1.26 on the leaderboard with those changes versus 1.04 for the TF version.\r\n\r\nI'll dig through the code some more to see if I spot anything else. These differences are a bit of a side issue of course, but I am just learning both Keras and TF and it might be instructive to understand if the differences between them are real.",
    "114973": "I've run the two versions as-is without any tweaking. They give me quite different results on the leaderboard (1.04 for TensorFlow, 1. 33 for Keras).\r\n\r\nIs that expected? Why is the Keras version worse?",
    "114950": "Here's the Keras version! Let me know if you need any help setting it up.\r\n\r\nhttps://github.com/fomorians/distracted-drivers-keras",
    "114946": "Very nice, thanks!\r\n\r\nLooking forward to the Keras version.",
    "115940": "Why do we **`img=img.swapaxes(2,0)`** in *prep_dataset.py*? I imagine this is because Keras requires images in a certain format, but I can't find the relevant documentation (if anyone can find it, please post here thanks).",
    "115187": "@Jim Fleming - Thanks also for this ;-) I understand the calculations, but why do you calculate \"calc_geom\" and \"calc_geom_arr\"? Many thanks for all your help in advance!!!"
  }
}