{
  "id": 22551,
  "title": "A beginner's ? - Benefits of & method to create multi-dim array of all images (in grayscale)",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22551",
  "author_name": "",
  "post_date": "2016-07-28T23:18:08.253Z",
  "votes": null,
  "comment_count": 6,
  "views": 484,
  "content": "<p>I'm a beginner to all of this and decided to 'dive into the deep end of the pool' by tackling this project so your kind suggestions and comments are greatly appreciated.</p>\n\n<p>Because all the images are the same dimension and because I have assumed that a grayscale version of each image is sufficient to solve this challenge, I wanted to create an array that contained the value of the grayscale at each pixel's location. </p>\n\n<p>I started with just one image and was able to create the array using this code:</p>\n\n<pre><code>im = Image.open(&quot;[drive location]BW driver photo for testing.jpg&quot;)\npix = im.load()\ndata = numpy.asarray(im)\n</code></pre>\n\n<p>Hoping that I could use some form of an iterative process to create the array, I created a new folder and renamed all the image files so that they are in numeric sequence from 1 - x but I still couldn't figure out which array create function to use or how to embed an iteration statement within it. (or if that is even a good idea)</p>\n\n<p>Please let me know if there is a way to create the array by combining functions and/or run iterations to add each image as a new dimension to the array.</p>\n\n<p><em>Are there benefits to generating such an array?</em></p>\n\n<p>My strategy was to run the following steps:</p>\n\n<ol>\n<li><p>Identify meaningful zones - such as driver's head / face and the steering wheel area - and get their coordinates within each specific image.  </p></li>\n<li><p>Evaluate the pixel values within those selected zones as well as within other grid areas - for example, I could find the outer bounds of the driver's side window and loosely infer from this where their head is and where the steering wheel is.  </p></li>\n<li><p>For the known images, compare the preselected image grid sections among all images for just that one driver and find out if there are variations that correlate highly with the known condition.  I don't want to compare different drivers yet because something like short sleeve vs. long sleeve shirt impacts ability to discern that both hands are on the wheel.</p></li>\n<li><p>I'd then like to find a way of assigning each test image to a particular driver (assuming there are multiple images per driver) so that I could limit the analysis to within-driver variations.  I would search for something that is uniquely characterized to each driver in the training images, such as a combination of pixel values from areas within the vehicle that don't change such as the bottom edge of passenger seat. This of course assumes that there's a high correlation of some portion of the image within drivers but not across drivers.</p></li>\n</ol>\n\n<p>I'm sure that I'm not including a lot of additional steps even within this beginning attempt and I welcome any top-level suggestions or comments on my overall approach.</p>\n\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "129345",
      "postDate": "07/28/2016 23:18:08",
      "content": "<p>I'm a beginner to all of this and decided to 'dive into the deep end of the pool' by tackling this project so your kind suggestions and comments are greatly appreciated.</p>\n\n<p>Because all the images are the same dimension and because I have assumed that a grayscale version of each image is sufficient to solve this challenge, I wanted to create an array that contained the value of the grayscale at each pixel's location. </p>\n\n<p>I started with just one image and was able to create the array using this code:</p>\n\n<pre><code>im = Image.open(&quot;[drive location]BW driver photo for testing.jpg&quot;)\npix = im.load()\ndata = numpy.asarray(im)\n</code></pre>\n\n<p>Hoping that I could use some form of an iterative process to create the array, I created a new folder and renamed all the image files so that they are in numeric sequence from 1 - x but I still couldn't figure out which array create function to use or how to embed an iteration statement within it. (or if that is even a good idea)</p>\n\n<p>Please let me know if there is a way to create the array by combining functions and/or run iterations to add each image as a new dimension to the array.</p>\n\n<p><em>Are there benefits to generating such an array?</em></p>\n\n<p>My strategy was to run the following steps:</p>\n\n<ol>\n<li><p>Identify meaningful zones - such as driver's head / face and the steering wheel area - and get their coordinates within each specific image.  </p></li>\n<li><p>Evaluate the pixel values within those selected zones as well as within other grid areas - for example, I could find the outer bounds of the driver's side window and loosely infer from this where their head is and where the steering wheel is.  </p></li>\n<li><p>For the known images, compare the preselected image grid sections among all images for just that one driver and find out if there are variations that correlate highly with the known condition.  I don't want to compare different drivers yet because something like short sleeve vs. long sleeve shirt impacts ability to discern that both hands are on the wheel.</p></li>\n<li><p>I'd then like to find a way of assigning each test image to a particular driver (assuming there are multiple images per driver) so that I could limit the analysis to within-driver variations.  I would search for something that is uniquely characterized to each driver in the training images, such as a combination of pixel values from areas within the vehicle that don't change such as the bottom edge of passenger seat. This of course assumes that there's a high correlation of some portion of the image within drivers but not across drivers.</p></li>\n</ol>\n\n<p>I'm sure that I'm not including a lot of additional steps even within this beginning attempt and I welcome any top-level suggestions or comments on my overall approach.</p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "I'm a beginner to all of this and decided to 'dive into the deep end of the pool' by tackling this project so your kind suggestions and comments are greatly appreciated.\r\n\r\nBecause all the images are the same dimension and because I have assumed that a grayscale version of each image is sufficient to solve this challenge, I wanted to create an array that contained the value of the grayscale at each pixel's location. \r\n\r\nI started with just one image and was able to create the array using this code:\r\n \r\n    im = Image.open(\"[drive location]BW driver photo for testing.jpg\")\r\n    pix = im.load()\r\n    data = numpy.asarray(im)\r\n\r\nHoping that I could use some form of an iterative process to create the array, I created a new folder and renamed all the image files so that they are in numeric sequence from 1 - x but I still couldn't figure out which array create function to use or how to embed an iteration statement within it. (or if that is even a good idea)\r\n\r\nPlease let me know if there is a way to create the array by combining functions and/or run iterations to add each image as a new dimension to the array.\r\n\r\n*Are there benefits to generating such an array?*\r\n\r\nMy strategy was to run the following steps:\r\n\r\n  1. Identify meaningful zones - such as driver's head / face and the steering wheel area - and get their coordinates within each specific image.  \r\n\r\n  2.  Evaluate the pixel values within those selected zones as well as within other grid areas - for example, I could find the outer bounds of the driver's side window and loosely infer from this where their head is and where the steering wheel is.  \r\n\r\n  3. For the known images, compare the preselected image grid sections among all images for just that one driver and find out if there are variations that correlate highly with the known condition.  I don't want to compare different drivers yet because something like short sleeve vs. long sleeve shirt impacts ability to discern that both hands are on the wheel.\r\n\r\n  4. I'd then like to find a way of assigning each test image to a particular driver (assuming there are multiple images per driver) so that I could limit the analysis to within-driver variations.  I would search for something that is uniquely characterized to each driver in the training images, such as a combination of pixel values from areas within the vehicle that don't change such as the bottom edge of passenger seat. This of course assumes that there's a high correlation of some portion of the image within drivers but not across drivers.\r\n\r\nI'm sure that I'm not including a lot of additional steps even within this beginning attempt and I welcome any top-level suggestions or comments on my overall approach.\r\n\r\nThank you!",
      "votes": null
    },
    {
      "id": "129402",
      "postDate": "07/29/2016 14:03:50",
      "content": "<p>Hi David,</p>\n\n<p>You should look at scripts in the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/kernels\">Kernels</a> section of the Dashboard. There are many examples that read the image files and convert them into a multidimensional array. The Keras (sample) script by ZFTurbo is a great example.</p>",
      "rawMarkdown": "Hi David,\r\n\r\nYou should look at scripts in the [Kernels][1] section of the Dashboard. There are many examples that read the image files and convert them into a multidimensional array. The Keras (sample) script by ZFTurbo is a great example.\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/kernels",
      "votes": null
    },
    {
      "id": "129418",
      "postDate": "07/29/2016 15:19:37",
      "content": "<p>Jeff,</p>\n\n<p>Thanks very much for your suggestion. </p>\n\n<p>I looked at ZFTurbo's code and it seems like it doesn't restrict the analysis to a particular portion of the image. </p>\n\n<p>Do you think it is worthwhile to first test against selected zones like the ones I've noted in the image below?  My assumption is that within each driver's image there would be some correlation that might exist between the driver's condition (reaching, drinking, etc) and the values within the different image sections. In essence, I was thinking that I could eliminate a lot of 'noise' by reducing the analysis to subsections of the image.</p>\n\n<p>To do this, I was hoping that I could create a variable to denote the position of 'driver's_head' for example within each image (which would now be each dimension in the array) and set the variable so that it was always the same size (eg 200 x 125) <em>but</em> the position of the area might change on each dimension. For example, on dimension one it could be at (76, 212) and on dimension two it might be at (114, 300). Could this be done by creating a completely new array for each of these image sections and I would then populate this new array with just that portion of each image that I've identified?</p>\n\n<p>I welcome any thoughts or comments.</p>\n\n<p>Thank you,</p>\n\n<p>David</p>\n\n<p><img src=\"http://nebula.wsimg.com/d8fec30170a573caa1554b3fd82d91c4?AccessKeyId=DF840A9E748B071D5D1E&disposition=0&alloworigin=1\" alt=\"Grids of image to evaluate\" title></p>",
      "rawMarkdown": "Jeff,\r\n\r\nThanks very much for your suggestion. \r\n\r\nI looked at ZFTurbo's code and it seems like it doesn't restrict the analysis to a particular portion of the image. \r\n\r\n\r\nDo you think it is worthwhile to first test against selected zones like the ones I've noted in the image below?  My assumption is that within each driver's image there would be some correlation that might exist between the driver's condition (reaching, drinking, etc) and the values within the different image sections. In essence, I was thinking that I could eliminate a lot of 'noise' by reducing the analysis to subsections of the image.\r\n\r\nTo do this, I was hoping that I could create a variable to denote the position of 'driver's_head' for example within each image (which would now be each dimension in the array) and set the variable so that it was always the same size (eg 200 x 125) *but* the position of the area might change on each dimension. For example, on dimension one it could be at (76, 212) and on dimension two it might be at (114, 300). Could this be done by creating a completely new array for each of these image sections and I would then populate this new array with just that portion of each image that I've identified?\r\n\r\nI welcome any thoughts or comments.\r\n\r\nThank you,\r\n\r\nDavid\r\n\r\n![Grids of image to evaluate][1]\r\n\r\n\r\n  [1]: http://nebula.wsimg.com/d8fec30170a573caa1554b3fd82d91c4?AccessKeyId=DF840A9E748B071D5D1E&disposition=0&alloworigin=1",
      "votes": null
    },
    {
      "id": "129430",
      "postDate": "07/29/2016 17:10:00",
      "content": "<p><strong>David J Stier</strong>, I think the main advantages of modern neural nets, that they found all these dependencies by itself during training. And you don't need to spend much time to find it by additional code. Nevertheless your additional image partition can increase final accuracy for some CNNs. I used OpenCV train cascade in early stage of competition and it gives improvement on LB.</p>",
      "rawMarkdown": "**David J Stier**, I think the main advantages of modern neural nets, that they found all these dependencies by itself during training. And you don't need to spend much time to find it by additional code. Nevertheless your additional image partition can increase final accuracy for some CNNs. I used OpenCV train cascade in early stage of competition and it gives improvement on LB.",
      "votes": null
    },
    {
      "id": "129433",
      "postDate": "07/29/2016 17:43:12",
      "content": "<p>Thanks very much for the suggestion!</p>\n\n<p>If you have a moment to reply to a more general question I'd greatly appreciate it....</p>\n\n<p>Is there a benefit of first looking at intra-driver images (just the images from one particular driver and see how they vary amongst the conditions being learned) before expanding the training to all drivers' images, or is it better to train against all drivers' images at the same time?  I assumed that because we know that the driver is the same within certain images, that this added information would improve the accuracy of the training model.</p>\n\n<p>Many thanks!</p>\n\n<p>David</p>",
      "rawMarkdown": "Thanks very much for the suggestion!\r\n\r\nIf you have a moment to reply to a more general question I'd greatly appreciate it....\r\n\r\nIs there a benefit of first looking at intra-driver images (just the images from one particular driver and see how they vary amongst the conditions being learned) before expanding the training to all drivers' images, or is it better to train against all drivers' images at the same time?  I assumed that because we know that the driver is the same within certain images, that this added information would improve the accuracy of the training model.\r\n\r\nMany thanks!\r\n\r\nDavid",
      "votes": null
    },
    {
      "id": "129438",
      "postDate": "07/29/2016 18:18:11",
      "content": "<p>It depends on way you want to create your model. I don't know what tool or method you want to use.</p>\n\n<p>I think the right way is to split all images in 2 parts by drivers. And then use the first part for training and second part for validation. For example, there are 26 drivers in train data set. So you can choose 24 for training and other 2 for validation. More complicated way is to use <a href=\"https://en.wikipedia.org/wiki/Cross-validation_(statistics)#k-fold_cross-validation\">KFold cross validation</a>.</p>\n\n<p>Split by drivers is important because test set contains of drivers which is not present in train dataset.</p>",
      "rawMarkdown": "It depends on way you want to create your model. I don't know what tool or method you want to use.\r\n\r\nI think the right way is to split all images in 2 parts by drivers. And then use the first part for training and second part for validation. For example, there are 26 drivers in train data set. So you can choose 24 for training and other 2 for validation. More complicated way is to use [KFold cross validation][1].\r\n\r\nSplit by drivers is important because test set contains of drivers which is not present in train dataset.\r\n\r\n  [1]: https://en.wikipedia.org/wiki/Cross-validation_(statistics)#k-fold_cross-validation",
      "votes": null
    },
    {
      "id": "129440",
      "postDate": "07/29/2016 18:26:28",
      "content": "<p>Thank you for your thoughtful reply. </p>\n\n<p>I appreciate your suggestion to keep out 2 images for validation. </p>\n\n<p>To be honest, I'm not yet sure how I'll build the model as I'm just starting out in Python.</p>\n\n<p>Thank you for taking the time to share your insights.</p>\n\n<p>Regards,</p>\n\n<p>David</p>",
      "rawMarkdown": "Thank you for your thoughtful reply. \r\n\r\nI appreciate your suggestion to keep out 2 images for validation. \r\n\r\nTo be honest, I'm not yet sure how I'll build the model as I'm just starting out in Python.\r\n\r\nThank you for taking the time to share your insights.\r\n\r\nRegards,\r\n\r\nDavid",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129402,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "07/29/2016 14:03:50",
      "content": "<p>Hi David,</p>\n\n<p>You should look at scripts in the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/kernels\">Kernels</a> section of the Dashboard. There are many examples that read the image files and convert them into a multidimensional array. The Keras (sample) script by ZFTurbo is a great example.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129418,
      "author_name": "davidstier",
      "author_url": "",
      "post_date": "07/29/2016 15:19:37",
      "content": "<p>Jeff,</p>\n\n<p>Thanks very much for your suggestion. </p>\n\n<p>I looked at ZFTurbo's code and it seems like it doesn't restrict the analysis to a particular portion of the image. </p>\n\n<p>Do you think it is worthwhile to first test against selected zones like the ones I've noted in the image below?  My assumption is that within each driver's image there would be some correlation that might exist between the driver's condition (reaching, drinking, etc) and the values within the different image sections. In essence, I was thinking that I could eliminate a lot of 'noise' by reducing the analysis to subsections of the image.</p>\n\n<p>To do this, I was hoping that I could create a variable to denote the position of 'driver's_head' for example within each image (which would now be each dimension in the array) and set the variable so that it was always the same size (eg 200 x 125) <em>but</em> the position of the area might change on each dimension. For example, on dimension one it could be at (76, 212) and on dimension two it might be at (114, 300). Could this be done by creating a completely new array for each of these image sections and I would then populate this new array with just that portion of each image that I've identified?</p>\n\n<p>I welcome any thoughts or comments.</p>\n\n<p>Thank you,</p>\n\n<p>David</p>\n\n<p><img src=\"http://nebula.wsimg.com/d8fec30170a573caa1554b3fd82d91c4?AccessKeyId=DF840A9E748B071D5D1E&disposition=0&alloworigin=1\" alt=\"Grids of image to evaluate\" title></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129430,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/29/2016 17:10:00",
      "content": "<p><strong>David J Stier</strong>, I think the main advantages of modern neural nets, that they found all these dependencies by itself during training. And you don't need to spend much time to find it by additional code. Nevertheless your additional image partition can increase final accuracy for some CNNs. I used OpenCV train cascade in early stage of competition and it gives improvement on LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129433,
      "author_name": "davidstier",
      "author_url": "",
      "post_date": "07/29/2016 17:43:12",
      "content": "<p>Thanks very much for the suggestion!</p>\n\n<p>If you have a moment to reply to a more general question I'd greatly appreciate it....</p>\n\n<p>Is there a benefit of first looking at intra-driver images (just the images from one particular driver and see how they vary amongst the conditions being learned) before expanding the training to all drivers' images, or is it better to train against all drivers' images at the same time?  I assumed that because we know that the driver is the same within certain images, that this added information would improve the accuracy of the training model.</p>\n\n<p>Many thanks!</p>\n\n<p>David</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129438,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "07/29/2016 18:18:11",
      "content": "<p>It depends on way you want to create your model. I don't know what tool or method you want to use.</p>\n\n<p>I think the right way is to split all images in 2 parts by drivers. And then use the first part for training and second part for validation. For example, there are 26 drivers in train data set. So you can choose 24 for training and other 2 for validation. More complicated way is to use <a href=\"https://en.wikipedia.org/wiki/Cross-validation_(statistics)#k-fold_cross-validation\">KFold cross validation</a>.</p>\n\n<p>Split by drivers is important because test set contains of drivers which is not present in train dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129440,
      "author_name": "davidstier",
      "author_url": "",
      "post_date": "07/29/2016 18:26:28",
      "content": "<p>Thank you for your thoughtful reply. </p>\n\n<p>I appreciate your suggestion to keep out 2 images for validation. </p>\n\n<p>To be honest, I'm not yet sure how I'll build the model as I'm just starting out in Python.</p>\n\n<p>Thank you for taking the time to share your insights.</p>\n\n<p>Regards,</p>\n\n<p>David</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129345": "I'm a beginner to all of this and decided to 'dive into the deep end of the pool' by tackling this project so your kind suggestions and comments are greatly appreciated.\r\n\r\nBecause all the images are the same dimension and because I have assumed that a grayscale version of each image is sufficient to solve this challenge, I wanted to create an array that contained the value of the grayscale at each pixel's location. \r\n\r\nI started with just one image and was able to create the array using this code:\r\n \r\n    im = Image.open(\"[drive location]BW driver photo for testing.jpg\")\r\n    pix = im.load()\r\n    data = numpy.asarray(im)\r\n\r\nHoping that I could use some form of an iterative process to create the array, I created a new folder and renamed all the image files so that they are in numeric sequence from 1 - x but I still couldn't figure out which array create function to use or how to embed an iteration statement within it. (or if that is even a good idea)\r\n\r\nPlease let me know if there is a way to create the array by combining functions and/or run iterations to add each image as a new dimension to the array.\r\n\r\n*Are there benefits to generating such an array?*\r\n\r\nMy strategy was to run the following steps:\r\n\r\n  1. Identify meaningful zones - such as driver's head / face and the steering wheel area - and get their coordinates within each specific image.  \r\n\r\n  2.  Evaluate the pixel values within those selected zones as well as within other grid areas - for example, I could find the outer bounds of the driver's side window and loosely infer from this where their head is and where the steering wheel is.  \r\n\r\n  3. For the known images, compare the preselected image grid sections among all images for just that one driver and find out if there are variations that correlate highly with the known condition.  I don't want to compare different drivers yet because something like short sleeve vs. long sleeve shirt impacts ability to discern that both hands are on the wheel.\r\n\r\n  4. I'd then like to find a way of assigning each test image to a particular driver (assuming there are multiple images per driver) so that I could limit the analysis to within-driver variations.  I would search for something that is uniquely characterized to each driver in the training images, such as a combination of pixel values from areas within the vehicle that don't change such as the bottom edge of passenger seat. This of course assumes that there's a high correlation of some portion of the image within drivers but not across drivers.\r\n\r\nI'm sure that I'm not including a lot of additional steps even within this beginning attempt and I welcome any top-level suggestions or comments on my overall approach.\r\n\r\nThank you!",
    "129402": "Hi David,\r\n\r\nYou should look at scripts in the [Kernels][1] section of the Dashboard. There are many examples that read the image files and convert them into a multidimensional array. The Keras (sample) script by ZFTurbo is a great example.\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/kernels",
    "129418": "Jeff,\r\n\r\nThanks very much for your suggestion. \r\n\r\nI looked at ZFTurbo's code and it seems like it doesn't restrict the analysis to a particular portion of the image. \r\n\r\n\r\nDo you think it is worthwhile to first test against selected zones like the ones I've noted in the image below?  My assumption is that within each driver's image there would be some correlation that might exist between the driver's condition (reaching, drinking, etc) and the values within the different image sections. In essence, I was thinking that I could eliminate a lot of 'noise' by reducing the analysis to subsections of the image.\r\n\r\nTo do this, I was hoping that I could create a variable to denote the position of 'driver's_head' for example within each image (which would now be each dimension in the array) and set the variable so that it was always the same size (eg 200 x 125) *but* the position of the area might change on each dimension. For example, on dimension one it could be at (76, 212) and on dimension two it might be at (114, 300). Could this be done by creating a completely new array for each of these image sections and I would then populate this new array with just that portion of each image that I've identified?\r\n\r\nI welcome any thoughts or comments.\r\n\r\nThank you,\r\n\r\nDavid\r\n\r\n![Grids of image to evaluate][1]\r\n\r\n\r\n  [1]: http://nebula.wsimg.com/d8fec30170a573caa1554b3fd82d91c4?AccessKeyId=DF840A9E748B071D5D1E&disposition=0&alloworigin=1",
    "129430": "**David J Stier**, I think the main advantages of modern neural nets, that they found all these dependencies by itself during training. And you don't need to spend much time to find it by additional code. Nevertheless your additional image partition can increase final accuracy for some CNNs. I used OpenCV train cascade in early stage of competition and it gives improvement on LB.",
    "129433": "Thanks very much for the suggestion!\r\n\r\nIf you have a moment to reply to a more general question I'd greatly appreciate it....\r\n\r\nIs there a benefit of first looking at intra-driver images (just the images from one particular driver and see how they vary amongst the conditions being learned) before expanding the training to all drivers' images, or is it better to train against all drivers' images at the same time?  I assumed that because we know that the driver is the same within certain images, that this added information would improve the accuracy of the training model.\r\n\r\nMany thanks!\r\n\r\nDavid",
    "129438": "It depends on way you want to create your model. I don't know what tool or method you want to use.\r\n\r\nI think the right way is to split all images in 2 parts by drivers. And then use the first part for training and second part for validation. For example, there are 26 drivers in train data set. So you can choose 24 for training and other 2 for validation. More complicated way is to use [KFold cross validation][1].\r\n\r\nSplit by drivers is important because test set contains of drivers which is not present in train dataset.\r\n\r\n  [1]: https://en.wikipedia.org/wiki/Cross-validation_(statistics)#k-fold_cross-validation",
    "129440": "Thank you for your thoughtful reply. \r\n\r\nI appreciate your suggestion to keep out 2 images for validation. \r\n\r\nTo be honest, I'm not yet sure how I'll build the model as I'm just starting out in Python.\r\n\r\nThank you for taking the time to share your insights.\r\n\r\nRegards,\r\n\r\nDavid"
  },
  "source": "meta"
}