{
  "id": 12842,
  "title": "Starter code for leaderboard score of ~0.46.",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/12842",
  "author_name": "",
  "post_date": "2015-03-16T23:22:01.660Z",
  "votes": 47,
  "comment_count": 76,
  "views": 34494,
  "content": "<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>",
  "messages": [
    {
      "id": "66564",
      "postDate": "03/16/2015 23:22:01",
      "content": "<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66589",
      "postDate": "03/17/2015 00:43:55",
      "content": "<p>Thank you. Could you share the cross validation kappa score of this approach?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66601",
      "postDate": "03/17/2015 01:17:26",
      "content": "<p>Hi RCarson,</p>\n<p>I have not yet run full cross-validation on this model yet to get a final kappa score. &nbsp;My goal was to get something done quickly, but running cross validation and tuning the intermediate parameters would be the next step and likely give a much better result. &nbsp;</p>\n<p>My holdout validation RMSE for the regression problem at the end to predict&nbsp;the level&nbsp;was around 0.63 if that helps.</p>\n<p>-- Hoyt</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66607",
      "postDate": "03/17/2015 01:36:13",
      "content": "<p>Great! my model gets rmse is 1.8 and LB 0.09. Thank you.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66626",
      "postDate": "03/17/2015 03:50:25",
      "content": "<p>Thanks for the great shared code!</p>\n<p>Just one thing which I hope you could clarify (thanks in advance): I noticed that your code used the&nbsp;Graphlab Create, which is not fully open source. Are you using any non-open source elements of the Graphlab Create in your code or not?</p>\n<p>Thanks!</p>\n<p>Best wishes,</p>\n<p>Shize</p>\n<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66654",
      "postDate": "03/17/2015 07:31:20",
      "content": "<p>Hello&nbsp;Shize,</p>\n<p>That's a fair question. I used <a href=\"https://dato.com/products/create/\">Dato's Graphlab Create</a>&nbsp;to prototype different models to quickly get at a good solution, which is what I've posted. However, everything that I ended up using is either open source or can be substituted with another open source package. The SFrames is open sourced in dato-core, the GLC deep learning toolkit generates config files that are compatible with CXXNet, and the boosted tree regression can be translated to use xgboost.</p>\n<p>Thanks!</p>\n<p>-- Hoyt</p>\n<p>[quote=Shize Su;66626]</p>\n<p>Thanks for the great shared code!</p>\n<p>Just one thing which I hope you could clarify (thanks in advance): I noticed that your code used the&nbsp;Graphlab Create, which is not fully open source. Are you using any non-open source elements of the Graphlab Create in your code or not?</p>\n<p>Thanks!</p>\n<p>Best wishes,</p>\n<p>Shize</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67948",
      "postDate": "03/24/2015 11:50:00",
      "content": "<p>Hello,</p>\n<p>Running the&nbsp;create_image_sframes.py script I get the following errors:</p>\n<p>PROGRESS: Read 163807 images in 745.984 secs speed: 0 file/sec<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/test/18344_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-stretch/train/4400_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-contrast-2/train/6210_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/train/27213_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-sat-1/test/8671_left.jpeg</p>\n\n\n<p>any idea where this might come from?</p>\n<p>thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67987",
      "postDate": "03/24/2015 15:34:17",
      "content": "<p>[quote=knarfben;67948]</p>\n<p>Hello,</p>\n<p>Running the&nbsp;create_image_sframes.py script I get the following errors:</p>\n<p>PROGRESS: Read 163807 images in 745.984 secs speed: 0 file/sec<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/test/18344_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-stretch/train/4400_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-contrast-2/train/6210_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/train/27213_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-sat-1/test/8671_left.jpeg</p>\n<p>any idea where this might come from?</p>\n<p>thanks.</p>\n<p>[/quote]</p>\n\n<p>ok.... not enough disk space, had to change&nbsp;GRAPHLAB_CACHE_FILE_LOCATIONS</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67989",
      "postDate": "03/24/2015 15:45:20",
      "content": "<p>Thank you Hoyt! it takes my computer a whole week to run this benchmark and it gets 0.44. In the following months I'm going to figure out what is going on and replace all the Dato tool with open source tool. It will be a great learning process!</p>\n<p>Again, Dato is really too good for kaggle. Hope you could publish more analysis, ipython notebook style tutorial, to help us learn your thought process, besides generating full submission. I think most kagglers will be more happy to use Dato as an analysis tool. It is indeed really powerful. Great work!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68033",
      "postDate": "03/24/2015 17:58:21",
      "content": "<p>@rcarson I ran this code for only the original dataset no hue, sat, ... processing and got 0.045. I was just curious what I could get in a day or so. Now, I am running the full set and expect me near you on the LB sometime this week :)</p>\n<p>As for the code, there is no magic to it. I actually think the network is so primitive (much more shallow (5 vs 2 conv layers)&nbsp;than the starter code based on cxxnet on the plankton competition though the fully connected layer has one more hidden layer).</p>\n<p>I think, the almost brute-force style&nbsp;pre-processing of color spaces, trying to balance the dataset, and extracting features from almost all kinds of class combinations are responsible for the higher score. This is evident in my score of only 0.045 on the original dataset (although balanced) vs. 0.44 or 0.46 for all this brute force. Overall, the data pre processing is not systematic rather brute-force and the network architecture is too shallow.</p>\n<p>Mind you although graphlab create net could be optimized, under the hood it is still cxxnet.&nbsp;</p>\n<p>It is a good starting place regardless.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68057",
      "postDate": "03/24/2015 19:33:59",
      "content": "<p>Thank you Hoyt. A kind reminder, multiple accounts are against Kaggle's rule. Luckily your new account doesn't enter this contest, so don't! and I suggest you delete the new account. You certainly don't need that. This happens to others who are new to kaggle. </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68061",
      "postDate": "03/24/2015 19:41:37",
      "content": "<p>@rcarson -- Thank you for pointing this out. &nbsp;I accidently logged in with the wrong google+ account and didn't even notice! &nbsp;I closed that one and will just stick with this one.&nbsp;</p>\n<p>In case the above post gets deleted, here is the text:</p>\n<p>---------------------------------------</p>\n<p>Hello @deep and @rcarson,</p>\n<p>I'm really glad you found my code helpful. I will try to make a notebook regarding my thought process here. I'm at work now, so I can't write much, but I'll try to fill in more when I get home.</p>\n<p>I agree with @deep -- this method is pretty brute force. The guiding principle was to create more observations that varied in the ways you want the NN to ignore. There are also a lot of different things that could be tuned and improved, especially with the NN architectures. I'm pretty new to neural nets, so I'll try playing around with it in the way you suggested.</p>\n<p>I think the other thing to realize here is that this is an ordinal regression problem rather than a classification problem. This motivated using the boosted regression trees over the activation levels in the final layers of the NN at the end to produce the result rather than using the straight NN to classify the images.</p>\n<p>@rcarson -- FYI, a lot of the Dato stuff is open source -- see https://github.com/dato-code/Dato-Core.</p>\n<p>@deep -- I have also found that sometimes what I'm doing here would overfit the training data, resulting in poorer predictions. E.g., it seemed like the NN overfit for the 0,1,2 vs 3,4 case based on the training / validation difference, and leaving that one out of the final regression problem resulted in a slightly better score. Just one thing I found.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68570",
      "postDate": "03/27/2015 13:25:48",
      "content": "<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Hello Hoyt</p>\n\n<p>I ran into the following issue while running&nbsp;create_image_sframes.py:</p>\n<p><code>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>Get roughly equal class representation by duplicating the different levels.<br>Do a poor mans random shuffle<br>Unable to reach server for 3 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 4 consecutive pings. Server is considered dead. Please exit and restart.<br>Traceback (most recent call last):<br> File &quot;/home/ubuntu/drd/create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 4988, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 39, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.<br>Unable to reach server for 5 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 6 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 7 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 8 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 9 consecutive pings. Server is considered dead. Please exit and restart.<br>[INFO] Stopping the server connection.<br></code></p>\n<p>Any idea where this might come from? I ran the script many times and always end up with this error, although previous communication with the server were ok....</p>\n\n<p>thanks&nbsp;</p>\n<p>Frank</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68602",
      "postDate": "03/27/2015 18:02:14",
      "content": "<p>[quote=knarfben;68570]</p>\n<p>Hello Hoyt</p>\n<p>I ran into the following issue while running&nbsp;create_image_sframes.py:</p>\n<p><code>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>Get roughly equal class representation by duplicating the different levels.<br>Do a poor mans random shuffle<br>Unable to reach server for 3 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 4 consecutive pings. Server is considered dead. Please exit and restart.<br>Traceback (most recent call last):<br> File &quot;/home/ubuntu/drd/create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 4988, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 39, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.<br>Unable to reach server for 5 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 6 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 7 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 8 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 9 consecutive pings. Server is considered dead. Please exit and restart.<br>[INFO] Stopping the server connection.<br></code></p>\n<p>Any idea where this might come from? I ran the script many times and always end up with this error, although previous communication with the server were ok....</p>\n<p>thanks&nbsp;</p>\n<p>Frank</p>\n<p>[/quote]</p>\n<p>Hello Frank, I also hit this error a couple of times -- hence the &quot;if not os.path.exists(...)&quot; line at the top, but running it a few times seems to get through it. &nbsp;Sorting the SFrame with the images in it seems really expensive, and all the disk IO is causing something to time out. I filed a bug report on this, so hopefully it will get fixed in the next Graphlab create version.</p>\n<p>I'm also working on doing all of the operations up to this point with filenames instead of the actual images, which would be much faster. &nbsp;</p>\n<p>Does it complete any of the rounds before crashing, or does it simply crash on the first one every time? &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68603",
      "postDate": "03/27/2015 18:07:58",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68903",
      "postDate": "03/29/2015 21:32:30",
      "content": "<p>The g2.2xlarge instance in AWS comes with 60GB of hard disk space. Could you let me know how do you increase this space to accommodate the dataset we are dealing with here?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68904",
      "postDate": "03/29/2015 21:39:20",
      "content": "<p>Create and mount a volume of sufficient size. I think I used 300GB but that was more than I needed.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68914",
      "postDate": "03/29/2015 22:39:55",
      "content": "<p>Did you create a EBS volume? I am new to AWS. If you store something in the 60GB instance store that comes with g2.2xlarge, does that data get deleted when you stop your instance? **If so, why would anyone ever use that store**? Second question I have is that when I log into my machine, how do I know if I am accessing the instance store or the EBS? i.e., given a path like /home/ubuntu how to know is this directory is under instance store or EBS?</p>\n\n<p>References:</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/AmazonEBS.html</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-using-volumes.html</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Storage.html</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/InstanceStorage.html</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68915",
      "postDate": "03/29/2015 22:57:29",
      "content": "<p>Yes it's an ebs volume. The data in the root device will disappear by default unless you select an option to make it persist when you create the AMI. The idea is to use it for software rather than data.&nbsp;</p>\n<p>To create a volume and attach it to a running instance, select create volume at the volumes screen, and then right click to and select attach volume. It will probably say it's attaching to /dev/sdf but in reality it will be on&nbsp;/dev/xvdf on ubuntu. From your instance, execute</p>\n\n<p><code>sudo mkfs -t ext4&nbsp;/dev/xvdf &nbsp;# (only the first time you use the volume).</code></p>\n<p><code></code><code>sudo mkdir /data # &nbsp;(replace /data with wherever you want to mount the volume).</code></p>\n<p><code>sudo mount /dev/xvdf /data</code><code></code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68947",
      "postDate": "03/30/2015 00:23:41",
      "content": "<p>hey i have one more question. i created a ebs volume and mounted it to /data but am unable to save anything in it because i lack the permissions.&nbsp;</p>\n<p>ubuntu:~$ ls -all /data<br>total 24<br>drwxr-xr-x 3 root root 4096 Mar 30 00:04 .<br>drwxr-xr-x 24 root root 4096 Mar 30 00:12 ..</p>\n<p>what can i do so that root becomes ubuntu? whoami gives ubuntu.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68963",
      "postDate": "03/30/2015 02:29:50",
      "content": "<p>sudo chown ubuntu /data</p>\n<p><code>chmod if necessary</code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68980",
      "postDate": "03/30/2015 05:28:02",
      "content": "<p>thanks. that worked.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68986",
      "postDate": "03/30/2015 07:05:40",
      "content": "<p>@Hoyt, just finished running your btb code. I like the intuition of pulling all the features and use GBT for classification instead of just straight softmax from the CNN.</p>\n<p>As you mentioned, the last two comparisons for 0-1-2 vs. 3-4 and 0-1-2-3 vs 4 overfit a lot (training: 0.91 and 0.95 vs. validation around 0.64 and 0.65) for this shallow network. This tells me that the network training params aren't quite optimal. For instance, adding regularization and adaptive learning rate could help.</p>\n<p>Great work overall and thanks for sharing!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69072",
      "postDate": "03/30/2015 20:00:08",
      "content": "<p>@Hoyt, can you say a few sentences about why you decided to go with just two convolutional layers and a stride of 4? I'm mostly trying to understand why two conv layers might be considered sufficient and whether a stride of 4 (&gt;1) was mostly selected to speed things up.</p>\n<p>This thread has been really instructive to me - thanks to all of you nice folks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69074",
      "postDate": "03/30/2015 20:16:46",
      "content": "<p>@small yellow duck --</p>\n<p>To be honest, I'm by no means the expert in this thread on deep learning -- others probably have better advice to share, but I would be happy to tell you what I was thinking.&nbsp;</p>\n<p>I based the architecture off of the configuration for imagenet, but I cut down a number of the parameters and the number of convolution layers. &nbsp;The main reason for only choosing 2 convolution layers was that I didn't feel this problem needed as much location invariance as the imagenet task, which had more variance in where the relevant features were located. &nbsp;After the processing by the imagemagick convert tool, a lot of relevant features appear in similar&nbsp;places in the image, so I didn't think as many convolution layers would be needed to capture it.</p>\n<p>As for the stride, I found that smaller strides generated out-of-memory errors on my GPU (a GTX 780), so that's the only reason for a stride of 4.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69703",
      "postDate": "04/05/2015 10:01:15",
      "content": "<p>Hello Hoyt</p>\n<p>Thanks for sharing ~~</p>\n<p>However, I ran into the following issue while running create_nn_model.py:</p>\n<p>File &quot;create_nn_model.py&quot;, line 85, in &lt;module&gt;<br> <strong>mean_image = X_train[&quot;image&quot;].mean()</strong></p>\n<p>graphlab.toolkits._main.ToolkitError: Cannot perform sum or average over images of different sizes. <strong>Found images of total size (ie. width * height * channels) of both 196608 and 65536.</strong> Please use graplab.image_analysis.resize() to make images a uniform size.</p>\n<p>Any idea where this might come from?&nbsp;</p>\n<p>thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69772",
      "postDate": "04/06/2015 06:29:35",
      "content": "<p>Since 196608 = 65536*3, it seems to me you are mixing 256x256 RGB (3 channels) and 256x256 grey level (1 channel) images, while the NN expects images with the same size AND number of channels.</p>\n<p>Just a guess...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69782",
      "postDate": "04/06/2015 08:58:59",
      "content": "<p>hi&nbsp;</p>\n<p>how to start this project</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70063",
      "postDate": "04/09/2015 00:30:36",
      "content": "<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70065",
      "postDate": "04/09/2015 01:07:31",
      "content": "<p>Better than 3 days before contest end (like in Plankton).</p>\n<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70066",
      "postDate": "04/09/2015 01:13:57",
      "content": "<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>\n<p>It may be the top ~7% currently but as others have mentioned it's not overly advanced and surely anyone who spent the length of the competition on their work would've bested it anyways.&nbsp; Additionally, I think we can all agree that openly discussing models that detect diseases has benefits outside of the competition alone :-)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70071",
      "postDate": "04/09/2015 02:30:17",
      "content": "<p>[quote=J Kolb;70066]</p>\n<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>\n<p>It may be the top ~7% currently but as others have mentioned it's not overly advanced and surely anyone who spent the length of the competition on their work would've bested it anyways.&nbsp; Additionally, I think we can all agree that openly discussing models that detect diseases has benefits outside of the competition alone :-)</p>\n<p>[/quote]</p>\n<p>Just saying, there is currently a user (not me) that is at 0.38 that has 59 submissions. &nbsp;Assuming that &quot;surely anyone who spent the length of the competition on their work would've bested it anyways&quot; is an unfair assumption in my estimation. &nbsp;Maybe that user did they best they possibly could and their work was original and may have been good enough to squeak into the top 25%. &nbsp;Seeing a score posted in the forums that trumps your work has to be extremely disheartening. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70072",
      "postDate": "04/09/2015 02:38:11",
      "content": "<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70078",
      "postDate": "04/09/2015 03:49:40",
      "content": "<p>[quote=James King;70072]</p>\n<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>\n<p>[/quote]</p>\n<p>I can understand the debate over how it affects people's scores, but this analogy doesn't really apply because if it's solely about working on the puzzle, no one makes you open the thread, much less go to his github, set up Graphlab, and do a couple days of processing and machine learning.&nbsp; He stated in the title what was in his thread.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70079",
      "postDate": "04/09/2015 03:57:56",
      "content": "<p>John,</p>\n<p>Assuming the converse is true, and the given person would not be able to beat the score of .46 after four months of work, it gives them the ability to peek into what other people are doing, if they wish, and not spend those four months without obtaining that result.&nbsp; They don't have to, however, and can still try things on their own, but at least they have that option.&nbsp; Also, it's well before the deadline so it's not like people are getting upended last minute.&nbsp; I'm not going to go further into whether it is disheartening vs. enlightening because I think this has already been beaten like a dead horse and there is obviously cases where each is true.&nbsp; I do want to also reiterate that this is an active research topic.&nbsp; Would you ask a research group not to publish their results because another research group is still working on the same problem?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70083",
      "postDate": "04/09/2015 04:33:19",
      "content": "<p>[quote=J Kolb;70079]</p>\n<p>John,</p>\n<p>Assuming the converse is true, and the given person would not be able to beat the score of .46 after four months of work, it gives them the ability to peek into what other people are doing, if they wish, and not spend those four months without obtaining that result.&nbsp; They don't have to, however, and can still try things on their own, but at least they have that option.&nbsp; Also, it's well before the deadline so it's not like people are getting upended last minute.&nbsp; I'm not going to go farther into whether it is disheartening vs. enlightening because I think this has already been beaten like a dead horse and there is obviously cases where each is true.&nbsp; I do want to also reiterate that this is an active research topic.&nbsp; Would you ask a research group not to publish their results because another research group is still working on the same problem?</p>\n<p>[/quote]</p>\n<p>My objections have nothing to do with publishing results. At the heart of scientific discovery is collaboration. It&#8217;s what makes this site and science in general, pretty awesome. My objections are with a company whoring out their products at the expense of others hard work. Users have posted successful models, plugging the same company&#8217;s tools, in this, the Otto challenge, and as inversion mentioned, one 3 days before the deadline in the Plankton competition, which easily put you in the top 25%. I am aware I have no results on this site to speak of and am not an influential member of this community, so if my objections are unwarranted, I will certainly acquiesce but I must remind those reading of the Law of Unintended Consequences; I had never heard of Dato until a few weeks ago. Maybe that was a good thing for them &#8230;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70084",
      "postDate": "04/09/2015 04:34:41",
      "content": "<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>\n<p>Hi John, thanks for your question, and hopefully I can give you a satisfactory answer. &nbsp;<br><br>I posted the code and write-up of my&nbsp;insights into the problem in the hope that others would be able to build on my&nbsp;ideas, incorporate their own, and hopefully come out ahead of either approach individually. &nbsp;I've been a&nbsp;researcher in machine learning for over 9 years -- first in my PhD&nbsp;program, and now at Dato -- and most of my research advances have benefitted greatly&nbsp;from open collaboration and an open&nbsp;exchange of ideas.&nbsp;&nbsp;As a result, I'm always free to share what insights I have in the hopes someone can build on them. &nbsp;I know this isn't necessarily academia, but this problem&nbsp;is still an open research problem, and there are several research papers posted elsewhere in the forum&nbsp;describing techniques that&nbsp;allegedly get a much better kappa score than my method. &nbsp;<br><br>I simply don't have the free time needed to win this contest, so that's not my aim; rather, I see it as an interesting and very difficult research problem that I want to feel like I can contribute to.&nbsp;&nbsp;And, from experience, I'll bet the winning idea is likely going to be a combination of openly shared ideas, published research, and original insights&nbsp;-- but it definitely won't be entirely&nbsp;original ideas.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70085",
      "postDate": "04/09/2015 04:51:48",
      "content": "<p>[quote=Hoyt Koepke;68061]</p>\n<p>I think the other thing to realize here is that this is an ordinal regression problem rather than a classification problem. This motivated using the boosted regression trees over the activation levels in the final layers of the NN at the end to produce the result rather than using the straight NN to classify the images.</p>\n<p>[/quote]</p>\n<p>I'm not sure I understand your thought process here.&nbsp; Wouldn't it be possible to use the Neural net for a regression problem and then round the result to the nearest value of 0, 1, 2, 3, or 4 instead of doing a one-vs-all classification? And then you wouldn't use the boosted trees at all.&nbsp; What is the benefit of using the trees then?&nbsp; Does using the NN for regression not work well in practice in this situation?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70086",
      "postDate": "04/09/2015 05:00:15",
      "content": "<p>[quote=John Martinez;70083]</p>\n<p>[quote=J Kolb;70079]</p>\n<p>John,</p>\n<p>Assuming the converse is true, and the given person would not be able to beat the score of .46 after four months of work, it gives them the ability to peek into what other people are doing, if they wish, and not spend those four months without obtaining that result.&nbsp; They don't have to, however, and can still try things on their own, but at least they have that option.&nbsp; Also, it's well before the deadline so it's not like people are getting upended last minute.&nbsp; I'm not going to go farther into whether it is disheartening vs. enlightening because I think this has already been beaten like a dead horse and there is obviously cases where each is true.&nbsp; I do want to also reiterate that this is an active research topic.&nbsp; Would you ask a research group not to publish their results because another research group is still working on the same problem?</p>\n<p>[/quote]</p>\n<p>My objections have nothing to do with publishing results. At the heart of scientific discovery is collaboration. It&#8217;s what makes this site and science in general, pretty awesome. My objections are with a company whoring out their products at the expense of others hard work. Users have posted successful models, plugging the same company&#8217;s tools, in this, the Otto challenge, and as inversion mentioned, one 3 days before the deadline in the Plankton competition, which easily put you in the top 25%. I am aware I have no results on this site to speak of and am not an influential member of this community, so if my objections are unwarranted, I will certainly acquiesce but I must remind those reading of the Law of Unintended Consequences; I had never heard of Dato until a few weeks ago. Maybe that was a good thing for them &#8230;</p>\n<p>[/quote]</p>\n<p>Hi John,&nbsp;</p>\n<p>I hope you don't think that poorly of Dato :-).&nbsp;&nbsp;We're a fairly new startup company originally formed in part to support Graphlab, an open source academic project, and now we have expanded to provide other&nbsp;tools for machine learning research and general data science. &nbsp;Because of our academic roots,&nbsp;we've made sure our stuff here is either open source or completely free for academic/non-commercial use to benefit research-oriented endeavors like this. &nbsp;As mentioned earlier, all of my code can either use our&nbsp;open source offering or be translated&nbsp;into other open source tools. &nbsp;I used it since it's (obviously) what I'm most familiar with, so I find it easier to quickly prototype ideas that way. &nbsp;&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70088",
      "postDate": "04/09/2015 05:06:51",
      "content": "<p>[quote=J Kolb;70085]</p>\n<p>[quote=Hoyt Koepke;68061]</p>\n<p>I think the other thing to realize here is that this is an ordinal regression problem rather than a classification problem. This motivated using the boosted regression trees over the activation levels in the final layers of the NN at the end to produce the result rather than using the straight NN to classify the images.</p>\n<p>[/quote]</p>\n<p>I'm not sure I understand your thought process here.&nbsp; Wouldn't it be possible to use the Neural net for a regression problem and then round the result to the nearest value of 0, 1, 2, 3, or 4 instead of doing a one-vs-all classification? And then you wouldn't use the boosted trees at all.&nbsp; What is the benefit of using the trees then?&nbsp; Does using the NN for regression not work well in practice in this situation?</p>\n<p>[/quote]</p>\n<p>Hi J Kolb,</p>\n<p>It's likely that you could get NN regression to work in this context, but I haven't really tried. &nbsp;The main reason is 2-fold. &nbsp;First, I had several NN models at the end, and this presented a quick and dirty way to ensemble them together. &nbsp;(I tried other classifiers and other tricks, but this was the only one that worked). &nbsp;Second,&nbsp;it took me a day or so to train each of the neural nets, but only a minute or two to train the boosted tree model at the end. &nbsp;By playing around with the boosted tree parameters, I was able to get a much better final score without spending 8 hours between each attempt.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70177",
      "postDate": "04/10/2015 06:53:49",
      "content": "<p>[quote=zyy1990;69703]</p>\n<p>Hello Hoyt</p>\n<p>Thanks for sharing ~~</p>\n<p>However, I ran into the following issue while running create_nn_model.py:</p>\n<p>File &quot;create_nn_model.py&quot;, line 85, in &lt;module&gt;<br> <strong>mean_image = X_train[&quot;image&quot;].mean()</strong></p>\n<p>graphlab.toolkits._main.ToolkitError: Cannot perform sum or average over images of different sizes. <strong>Found images of total size (ie. width * height * channels) of both 196608 and 65536.</strong> Please use graplab.image_analysis.resize() to make images a uniform size.</p>\n<p>Any idea where this might come from?&nbsp;</p>\n<p>thanks</p>\n<p>[/quote]</p>\n<p>zyy1990: Have&nbsp;you solved this problem?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "70186",
      "postDate": "04/10/2015 07:29:29",
      "content": "<p>[quote=vnneo;70177]</p>\n<p>[quote=zyy1990;69703]</p>\n<p>Hello Hoyt</p>\n<p>Thanks for sharing ~~</p>\n<p>However, I ran into the following issue while running create_nn_model.py:</p>\n<p>File &quot;create_nn_model.py&quot;, line 85, in &lt;module&gt;<br> <strong>mean_image = X_train[&quot;image&quot;].mean()</strong></p>\n<p>graphlab.toolkits._main.ToolkitError: Cannot perform sum or average over images of different sizes. <strong>Found images of total size (ie. width * height * channels) of both 196608 and 65536.</strong> Please use graplab.image_analysis.resize() to make images a uniform size.</p>\n<p>Any idea where this might come from?&nbsp;</p>\n<p>thanks</p>\n<p>[/quote]</p>\n<p>zyy1990: Have&nbsp;you solved this problem?</p>\n<p>[/quote]</p>\n\n<p>Yes. For some unknown reason, the data i downloaded had many images which were all black, so i picked &nbsp; those images out. Here are the id of those images:</p>\n<p>train/41176_left.jpeg<br>train/34689_left.jpeg<br>train/32253_right.jpeg<br>train/43457_left.jpeg<br>train/1986_left.jpeg<br>test/28544_left.jpeg<br>test/16519_left.jpeg<br>test/38549_left.jpeg<br>test/35433_left.jpeg<br>test/3517_left.jpeg<br>test/35762_left.jpeg</p>\n<p>train/15222_left.jpeg<br>test/16349_left.jpeg</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71221",
      "postDate": "04/10/2015 15:44:27",
      "content": "<p>[quote=James King;70072]</p>\n<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>\n<p>[/quote]</p>\n\n<p>I agree.</p>\n<p>My opinion was that posting this solution was poor sportsmanship. If part of the motivation was to market a company's tools, then it leads me to have a negative opinion of that company as well. Seeing this solution decreased my motivation to participant in this contest.</p>\n<p>Obviously the original poster did not think feel that way, so what we have right now is a difference of opinion. However, we have an opportunity to resolve this question by means of (yes!) data.</p>\n<p>If lots of people agree with me, then as a community we should discourage this kind of posting, and the site administrators should note that it will be bad for Kaggle if people are turned off by the practice and stop participating. If people think this kind of posting is a good idea, then we should encourage people to post full solutions early and often.</p>\n<p>Please voice your opinion. What do you think?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71224",
      "postDate": "04/10/2015 15:56:11",
      "content": "<p>28 people upvoted the OP... so, we've already had a vote of sorts. Please don't conflate self-promotion with code sharing. Although code sharing&nbsp;is not&nbsp;entirely altruistic, I am fundamentally for it (both in general and specifically in Kaggle competitions). I myself was motivated (rather than discouraged) by this thread. I felt stimulated to brainstorm ways to improve upon the OP's score, yet also comforted that if I failed to get a good model during these next few months, that I could fall back on this shared code. A HUGE thank you to the OP.</p>\n\n<p>My strong opinion is that collaboration is beautiful. If you don't want to look at the shared code, don't look at it. Easy solution. Live and let live.</p>\n\n<p>Perhaps, another question we should ask (for another thread), is at what point does it become rude to hijack someone else's thread simply because you have a different opinion?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71227",
      "postDate": "04/10/2015 16:26:40",
      "content": "<p>[quote=Michael 7402;71221]Please voice your opinion. What do you think?[/quote]</p>\n<p>Idem datanewb: please stop this flood. (yeah.. self-defeating demand :-) )</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71229",
      "postDate": "04/10/2015 16:57:53",
      "content": "<p>[quote=datanewb;71224]</p>\n<p>28 people upvoted the OP... so, we've already had a vote of sorts. Please don't conflate self-promotion with code sharing. Although code sharing&nbsp;is not&nbsp;entirely altruistic, I am fundamentally for it (both in general and specifically in Kaggle competitions). I myself was motivated (rather than discouraged) by this thread. I felt stimulated to brainstorm ways to improve upon the OP's score, yet also comforted that if I failed to get a good model during these next few months, that I could fall back on this shared code. A HUGE thank you to the OP.</p>\n<p>My strong opinion is that collaboration is beautiful. If you don't want to look at the shared code, don't look at it. Easy solution. Live and let live.</p>\n<p>Perhaps, another question we should ask (for another thread), is at what point does it become rude to hijack someone else's thread simply because you have a different opinion?</p>\n<p>[/quote]</p>\n<p>There was no conflation of self-promotion and code-sharing. &nbsp;The evidence is in the forums of those 3 competitions (and maybe others), as I stated. &nbsp;I completely agree, and would assume that Kaggle does as well, that collaboration is a beautiful thing, which is why we are allowed to form&nbsp;Teams.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71248",
      "postDate": "04/10/2015 20:51:45",
      "content": "<p>[quote=datanewb;71224]</p>\n<p>if I failed to get a good model during these next few months, that I could fall back on this shared code.</p>\n<p>[/quote]</p>\n<p>At the detriment of everyone you passed over on the leader board. Just by copying and pasting code. Great. Now the LB is no longer reflective of what you are capable of.</p>\n<p>This issue could be completely resolved if the Dato folks posted models without optimized hyperparameters, and perhaps even stating what the model is&nbsp;<em>capable</em> of. And let their users work at it a bit.</p>\n<p>In fact, they'd be doing their users a favor.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71249",
      "postDate": "04/10/2015 21:08:44",
      "content": "<p>This is a hard one ----yes or no to.</p>\n<p>I follow a good many Kaggle Forums and to be open I have learned quite a few things by looking into everything everyone has posted and have learned new ways to compete in this environment.</p>\n<p>Being a Novice here (or anywhere for that matter) for this type of competition I initially had no idea how I was going to use what I know about ML to compete.</p>\n<p>I agree with all of you these type of posing are a type of self promotion.... by thanks to a few of them I am now feeling a bit more confident about being able to participate in these competitions and see nothing wrong as long as we are able to learn something from the post.... even if this particular one seem to be very self-promoting.... IMHO</p>\n<p>I do commend you all for coming out against it to keep self-promotion under control&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71382",
      "postDate": "04/13/2015 12:09:41",
      "content": "<p>Hi Hoyt,</p>\n<p>Thank you for sharing!</p>\n<p>Do you have any data, what is the best benchmark with this base solution so far?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71397",
      "postDate": "04/13/2015 17:55:08",
      "content": "<p>Hello Hoyt,</p>\n<p>I've just started the NN training process on my GTX 980 and since you wrote: &quot;and about 320 images / second on a GTX 780&quot; in the readme, I expected about 20% better performance on my configuration.</p>\n<p>I got 280 img/s.</p>\n<p>Do you have any idea what can be the reason?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71399",
      "postDate": "04/13/2015 18:16:20",
      "content": "<p>[quote=George Solymosi;71382]</p>\n<p>Hi Hoyt,</p>\n<p>Thank you for sharing!</p>\n<p>Do you have any data, what is the best benchmark with this base solution so far?</p>\n<p>[/quote]</p>\n<p>Hi George,</p>\n<p>No, I'm not sure at this point. &nbsp;It's a pretty crude approach; there are a lot of parameters to tune, and I've since realized there are a lot of things holding back the accuracy that become obvious when you examine the output solutions. &nbsp;In particular, as others have mentioned, the NN architecture here overfits to some parts of the the training data, it doesn't take into account the unbalanced class sizes at prediction time, etc.&nbsp;&nbsp;I got somewhere close to 0.46 with this solution; some tweaks to boosted tree algorithm (unpublished) boosted&nbsp;the score somewhat. &nbsp; I'm sure that more sophistication in this approach, a better NN architecture, and/or ensembling this model with other approaches is needed to get much better than this.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71415",
      "postDate": "04/13/2015 23:06:20",
      "content": "<p>@hoyt, just some comments on the image processing... I have yet to implement something myself</p>\n<p>a) equalisation is normally only done on the intensity component so that you don't distort the colour balance (ie you don't want to stretch each colour channel separately)</p>\n<p><a href=\"http://www.imagemagick.org/script/command-line-options.php#equalize\">http://www.imagemagick.org/script/command-line-options.php#equalize</a></p>\n\n<p>b) i would think that its better to convert to png rather than jpg (which might introduce new jpeg artifacts at 256x256 size?)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71417",
      "postDate": "04/13/2015 23:11:41",
      "content": "<p>[quote=George Solymosi;71397]</p>\n<p>Hello Hoyt,</p>\n<p>I've just started the NN training process on my GTX 980 and since you wrote: &quot;and about 320 images / second on a GTX 780&quot; in the readme, I expected about 20% better performance on my configuration.</p>\n<p>I got 280 img/s.</p>\n<p>Do you have any idea what can be the reason?</p>\n<p>[/quote]</p>\n<p>Hey George,&nbsp;</p>\n<p>Sorry -- I have no idea. &nbsp;I suspect it also has to do with the other configurations of your system. The only thing I can think of is that it's IO bound in moving data on and off the card -- maybe it's mounted in the x8 PCIe slot instead of the x16 PCIe slot (although that shouldn't make a difference). &nbsp;</p>\n<p>The batch_size parameter controls how many images are loaded onto the graphics card at a time for training. &nbsp;It's possible that adjusting this parameter will help with performance.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71418",
      "postDate": "04/13/2015 23:16:15",
      "content": "<p>[quote=Sean;71415]</p>\n<p>@hoyt, just some comments on the image processing... I have yet to implement something myself</p>\n<p>a) equalisation is normally only done on the intensity component so that you don't distort the colour balance (ie you don't want to stretch each colour channel separately)</p>\n<p><a href=\"http://www.imagemagick.org/script/command-line-options.php#equalize\">http://www.imagemagick.org/script/command-line-options.php#equalize</a></p>\n<p>b) i would think that its better to convert to png rather than jpg (which might introduce new jpeg artifacts at 256x256 size?)</p>\n<p>[/quote]</p>\n<p>Hey Sean,&nbsp;</p>\n<p>Thanks for your comments! &nbsp;I called equalize mainly so that the 3 different channels of the NN would be on the same scale, although I'm not sure that would have any significant effect. &nbsp;And yes, png might be better. &nbsp;I'll give that a shot for my latest iteration.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72504",
      "postDate": "04/19/2015 19:26:31",
      "content": "<p>We can fight with each other, but if we can quicker get closer with a bit help of something like Hoyt's shortcut to a better solution for help to prevent and cure people's illness, then I don't think that, that is unfair...</p>\n<p>Consider it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72719",
      "postDate": "04/20/2015 19:06:49",
      "content": "<p>[quote=Hoyt Koepke;66564]</p>\n<p>Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</p>\n<p>[/quote]</p>\n<p>I'm a bit puzzled though, you say you convert images to 256x256 and at the same time&nbsp;you say you avoid scaling the images. Could you kindly clarify this part?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72986",
      "postDate": "04/21/2015 17:58:09",
      "content": "<p>[quote=K0stIa;72719]</p>\n<p>I'm a bit puzzled though, you say you convert images to 256x256 and at the same time&nbsp;you say you avoid scaling the images. Could you kindly clarify this part?</p>\n<p>[/quote]</p>\n<p>Sorry, I could have been clearer there. &nbsp;What I meant is that I made sure that the aspect ratio of the images and the location of the eyes in the picture stayed the same so that the resulting NN did not have to handle scale and location invariance along with learning to classify the images. &nbsp;</p>\n\n<p>Hope that helps!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "73007",
      "postDate": "04/21/2015 18:50:03",
      "content": "<p>You do realize that the beauty of CNN's that scaling and translation does really cause any performance penalties of the CNN's..... However make sure you don't scale away the useful information ;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "73184",
      "postDate": "04/22/2015 15:24:24",
      "content": "<p>Thanks Dato Inc / Hoyt for trying to democratize Big Data / Deep Learning algorithms. I was curious what pretrained filters the GraphLab's deeplearning algorithm make use of ? Is it Gabor Filters or BSD Caffe or something else ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74244",
      "postDate": "04/23/2015 04:23:32",
      "content": "<p>Hi Hoyt,</p>\n<p>I'm&nbsp;a little confused on the performance you mentioned. 300 images/s tells me each epoch takes around 2 min, and 5-10 epochs for each 2-class model comes out to be 20 min of training. How does this translate to about a day (24h) worth of training? Would you mind clarifying? I'm using theano so I'm curious how my performance compares to your cxxnet implementation.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74358",
      "postDate": "04/23/2015 18:12:36",
      "content": "<p>[quote=Michael George Hart;73007]</p>\n<p>You do realize that the beauty of CNN's that scaling and translation does really cause any performance penalties of the CNN's..... However make sure you don't scale away the useful information ;)</p>\n<p>[/quote]</p>\n<p>Hi Michael,&nbsp;</p>\n<p>Yes, true, though in general, when you add more convolution/pooling layers in order to get closer to true scale invariance, you end up needing more data in order to get it to train effectively. &nbsp;I opted for having much fewer convolution layers than is needed for true translation invariance, since it seemed like location in the image could actually be informative. &nbsp;I think there may be something to this approach, as I've been trying to train with deeper architectures since then, but haven't really got significantly better results yet.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74360",
      "postDate": "04/23/2015 18:15:58",
      "content": "<p>[quote=saurk;73184]</p>\n<p>Thanks Dato Inc / Hoyt for trying to democratize Big Data / Deep Learning algorithms. I was curious what pretrained filters the GraphLab's deeplearning algorithm make use of ? Is it Gabor Filters or BSD Caffe or something else ?</p>\n<p>[/quote]</p>\n<p>Hello Saurk,</p>\n<p>Thanks! &nbsp;Under the hood, graphlab create doesn't use any pretrained filters. However, the filters in the convolution layers will often end up resembling Gabor filters, but they are learned directly from the data during training. &nbsp;</p>\n<p>I think Caffe will end up doing the same thing...</p>\n<p>-- Hoyt</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74361",
      "postDate": "04/23/2015 18:16:22",
      "content": "<p>[quote=jstaker7;74244]</p>\n<p>Hi Hoyt,</p>\n<p>I'm&nbsp;a little confused on the performance you mentioned. 300 images/s tells me each epoch takes around 2 min, and 5-10 epochs for each 2-class model comes out to be 20 min of training. How does this translate to about a day (24h) worth of training? Would you mind clarifying? I'm using theano so I'm curious how my performance compares to your cxxnet implementation.</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74364",
      "postDate": "04/23/2015 18:30:44",
      "content": "<p>[quote=jstaker7;74244]</p>\n<p>Hi Hoyt,</p>\n<p>I'm&nbsp;a little confused on the performance you mentioned. 300 images/s tells me each epoch takes around 2 min, and 5-10 epochs for each 2-class model comes out to be 20 min of training. How does this translate to about a day (24h) worth of training? Would you mind clarifying? I'm using theano so I'm curious how my performance compares to your cxxnet implementation.</p>\n<p>[/quote]</p>\n<p>Hello jstaker,</p>\n<p>The code I use brute-force creates a bunch of different versions of each image, as well as duplicating some of them to&nbsp;balance the classes in the training set. &nbsp;Thus there end up being hundreds of thousands of images trained in each epoch instead of the original 35,000 or so. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74417",
      "postDate": "04/24/2015 00:44:22",
      "content": "<p>[quote=Hoyt Koepke;74358]</p>\n<p>... I opted for having much fewer convolution layers than is needed for true translation invariance, since it seemed like location in the image could actually be informative...</p>\n<p>[/quote]</p>\n\n<p>This challenges my understanding of CNNs. Isn't just feature detection translation invariant? I thought one of the great advantages of CNNs is that it can capture spacial relationships. Am I understanding your comment correctly?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74445",
      "postDate": "04/24/2015 06:36:51",
      "content": "<p>[quote=jstaker7;74417]</p>\n<p>[quote=Hoyt Koepke;74358]</p>\n<p>... I opted for having much fewer convolution layers than is needed for true translation invariance, since it seemed like location in the image could actually be informative...</p>\n<p>[/quote]</p>\n<p>This challenges my understanding of CNNs. Isn't just feature detection translation invariant? I thought one of the great advantages of CNNs is that it can capture spacial relationships. Am I understanding your comment correctly?</p>\n<p>[/quote]</p>\n<p>CNNs can learn not only translation invariance, but even scale and rotation invariances, but for this to happen, those characteristics need to be derived from the training set, i.e. the independence between (position, rotation, scale) of the features and class label needs to be present in the dataset. If, for some reason, some examples contain some features which are positioned constantly in the same general area of the image, even the translation invariance will go away,&nbsp;because the CNN would 'think' in this case that the position of those features should affect the decision in some way, i.e. the position of the feature and the class label would not be independent and identically distributed (i.i.d.) any more. In conclusion, CNNs CAN learn a lot of invariances, but in order for them to learn, those invariances need to be present in the training set, and this is&nbsp;why data augmentation is done.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75145",
      "postDate": "04/27/2015 23:47:21",
      "content": "<p>I'm not an expert on CNNs, but my understanding is that the translation invariance comes purely from the pooling layers. &nbsp;Thus CNNs are only truly translation invariant if there are enough pooling layers that sufficiently significant features detected on one side of the image can affect the outcome of any of the nodes before the final fully connected layers. &nbsp;In my original net, I didn't have enough of these layers to achieve this, so it then relied on the fully connected layers at the end to do more of the work.</p>\n<p>Similarly, scale and rotation invariances come from the number of layers and having enough channels in the convolution layers to adequately fit the range of features present in the data.&nbsp;</p>\n<p>The problem, however, is that the more channels and layers you add, the more parameters need to be trained, and thus the more data you need to train the net without overfitting. &nbsp;Ideally, you want a complicated enough structure to capture all relevant features, but a simple enough structure that it doesn't overfit (though dropout and regularization can help with that).&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75242",
      "postDate": "04/28/2015 13:38:09",
      "content": "<p>Yes pooling helps with the invariance ... A type of averaging mechanism.....</p>\n<p>however it is the Convolution Layer, a, more realistic, model of the visual system, that makes i variances .... Get the spacing right on the convolution will be more useful for this problem than the adding more pooling. ... Most of the work has already done for us by GoogLeNet, AlexNet etc... The trick here is to modify those convolution layers and perhaps tweeting the pooling layers to meet the needs of this particular problem....</p>\n<p>i imagine many people already see this as a solution... That is why I am only using it as a starting point to see how far I can get in the rankings, because the competition will come to who can tune the existing CNN's for a solution to this challenge....&nbsp;</p>\n<p>To win this challenge someone will have to step outside the CNN box since it is such an abvious solution... ;)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "77297",
      "postDate": "05/06/2015 11:02:26",
      "content": "<p>Kaggle has made their <strong>position clear</strong> on this in the competition&nbsp;rules I think: &quot;<em>It's okay to share code if made available to all participants on the forums.&quot;</em></p>\n<p>Why would publishing any competition entry bad or discouraging?</p>\n<p>- it's absolutely *<strong>not</strong>* the same as someone coming and telling you the 'answer to a puzzle' as you say.</p>\n<p>Here, you&nbsp;are not <strong>forced</strong> to read the forums, let alone study/run the code on github.</p>\n<p>I think that *<strong>anyone</strong>* who reads these forums is doing so *<strong>exactly</strong>* because they want to find out what other people are doing.</p>\n<p>- how is<strong> seeing a solution</strong> that scores better than you discouraging? You already can see in the leaderboard people's scores. Obviously those&nbsp;who&nbsp;scored above you have a solution to go with it, there is no surprise there.</p>\n<p>- if there is a <strong>tool that can help</strong>, and it's allowed by the rules (perhaps), why not use it? You say &quot;seeing this solution decreased by motivation to participant [sic] in this contest&quot;,&nbsp;others would say &quot;seeing this solution <strong>increased&nbsp;</strong>my motivation&quot; as they have something to get started with.</p>\n<p>[quote=Michael 7402;71221]</p>\n<p>[quote=James King;70072]</p>\n<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>\n<p>[/quote]</p>\n<p>I agree.</p>\n<p>My opinion was that posting this solution was poor sportsmanship. If part of the motivation was to market a company's tools, then it leads me to have a negative opinion of that company as well. Seeing this solution decreased my motivation to participant in this contest.</p>\n<p>Obviously the original poster did not think feel that way, so what we have right now is a difference of opinion. However, we have an opportunity to resolve this question by means of (yes!) data.</p>\n<p>If lots of people agree with me, then as a community we should discourage this kind of posting, and the site administrators should note that it will be bad for Kaggle if people are turned off by the practice and stop participating. If people think this kind of posting is a good idea, then we should encourage people to post full solutions early and often.</p>\n<p>Please voice your opinion. What do you think?</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80553",
      "postDate": "06/01/2015 06:46:22",
      "content": "<p>Hi Hoyt,</p>\n<p>Thanks for the code. During prediction, I am getting different prediction levels for the same image in different runs. Please help me to sort out this. The results are as follows for consecutive runs of same set of images.</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.751127591425 |<br>| 11730_right | 0 | 0.358918072124 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.423049250692 |<br>| 11730_right | 0 | 0.326992818052 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.70960257402 |<br>| 11730_right | 0 | 0.357797767757 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.408557624363 |<br>| 11730_right | 0 | 0.355133798881 |<br>+-------------+-------+----------------+</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "81044",
      "postDate": "06/05/2015 19:35:40",
      "content": "<p>Hi Hoyt,</p>\n<p>Thanks for posting this started code. I am impressed by what Graphlab can do with just a few hundred lines of Python code! I tried to run your Github code but looks like the Graphlab API crashed in sframe sorting - see below and the log file attached. If I comment out the sorting code then it will run fine. Any idea?</p>\n<p>I'm running Ubuntu 14.04 on a desktop Intel Core i7-5820K CPU @ 3.30GHz &#215; 12, with 8GB Memory plus GeForce GTX 970 GPUs.</p>\n<p>=========================</p>\n<p>Traceback (most recent call last):<br> File &quot;create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 5077, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 31, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.</p>\n<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "81046",
      "postDate": "06/05/2015 19:52:09",
      "content": "<p>Hi Rajkumar,</p>\n<p>It turns out that this is a known issue that the NN people here are working on. &nbsp;I'm not sure there's a way around it now. &nbsp;I think if you use the boosted tree regression as a post-processing step on the activation layers, it's a bit more robust.&nbsp;</p>\n<p>-- Hoyt</p>\n\n<p>[quote=Rajkumar;80553]</p>\n<p>Hi Hoyt,</p>\n<p>Thanks for the code. During prediction, I am getting different prediction levels for the same image in different runs. Please help me to sort out this. The results are as follows for consecutive runs of same set of images.</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.751127591425 |<br>| 11730_right | 0 | 0.358918072124 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.423049250692 |<br>| 11730_right | 0 | 0.326992818052 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.70960257402 |<br>| 11730_right | 0 | 0.357797767757 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.408557624363 |<br>| 11730_right | 0 | 0.355133798881 |<br>+-------------+-------+----------------+</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "81048",
      "postDate": "06/05/2015 20:01:48",
      "content": "<p>Hello fisherman,</p>\n<p>In this case, I kept hitting issues on sort running out of memory, etc. as well. &nbsp;I'm not sure of a workaround except to use a machine with more memory. &nbsp;I did talk with the people here at Dato responsible for the sort stuff, and it's getting better. &nbsp;</p>\n<p>The solution I ended up using was to&nbsp;modify the code so that it just used the file paths up until the final part when it generated the sframes, at which point I load them using a lambda function. &nbsp;This makes it significantly faster and the sort is much easier.</p>\n<p>Something like:&nbsp;</p>\n<p><code>X[&quot;path&quot;] = list(chain(*[[abspath(join(root, f)) for f in files<br></code><code><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if f.endswith('jpeg')]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for root, dir_list, files in os.walk(image_path)]))<br></code><code><br><br></code></code></p>\n<p>Then use something like</p>\n<p><code>X[&quot;image&quot;] = X[&quot;path&quot;].apply(lambda p: gl.Image(p)) </code></p>\n<p>to load it at the end.&nbsp;</p>\n<p>Hope something like that helps!&nbsp;</p>\n<p>-- Hoyt</p>\n<p>[quote=fisherman;81044]</p>\n<p>Hi Hoyt,</p>\n<p>Thanks for posting this started code. I am impressed by what Graphlab can do with just a few hundred lines of Python code! I tried to run your Github code but looks like the Graphlab API crashed in sframe sorting - see below and the log file attached. If I comment out the sorting code then it will run fine. Any idea?</p>\n<p>I'm running Ubuntu 14.04 on a desktop Intel Core i7-5820K CPU @ 3.30GHz &#215; 12, with 8GB Memory plus GeForce GTX 970 GPUs.</p>\n<p>=========================</p>\n<p>Traceback (most recent call last):<br> File &quot;create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 5077, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 31, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.</p>\n<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "82722",
      "postDate": "06/25/2015 11:12:43",
      "content": "<p>Hi Hoyt, thanks for the startercode to begin with.</p>\n<p>I'm having an issue when I run&nbsp;<span style=\"line-height: 1.4\">create_image_sframes.py</span></p>\n<p>I get the message below repeatedly</p>\n<p>PROGRESS: Unexpected JPEG decode failure file: /mnt/my-data/kaggle-data/processed/run-normal/test/21697_left.jpeg</p>\n<p>Any ideas off-hand how to fix this?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "82746",
      "postDate": "06/25/2015 17:50:51",
      "content": "<p>Hey Octave1,</p>\n<p>I think it's happening because somehow the shrinking part of the imagemagick script ended up shrinking one of those images to zero. &nbsp;It should be fine if you just delete that image, or rerun the convert command for just that image.&nbsp;</p>\n<p>-- Hoyt</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "84211",
      "postDate": "07/12/2015 15:02:54",
      "content": "<p>thank you Hoyt, I can run following your guides.</p>",
      "rawMarkdown": "thank you Hoyt, I can run following your guides.",
      "votes": null
    },
    {
      "id": "86129",
      "postDate": "07/20/2015 06:58:23",
      "content": "<p>Hi Hoyt,</p>\n\n<p>I tried running your code. I did it only to train <code>which_model = 4</code> i.e. predicting 4-vs-all retinopathy grade.\nMy training and validation accuracy figures were quite low so I checked the features that have been stored in <code>Xty</code> and <code>Xtst</code> in <code>create_nn_model.py</code> and it shows <code>nan</code> values in both <code>scores_4</code> and <code>features_4</code> columns. Any idea about what is it that I did wrong?</p>\n\n<p>Following are the details of the model that was stored:</p>\n\n<pre><code>Class               : NeuralNetClassifier\n\nSchema\n------\nExamples            : 848387\nFeatures            : 1\nTarget column       : class\n\nTraining Summary\n----------------\nTraining accuracy   : 0.4992\nValidation accuracy : 0.5416\nTraining time (sec) : 85850.3034\n</code></pre>\n\n<p>Following are the details of <code>Xf_train</code> from <code>create_submission.py</code> code.</p>\n\n<pre><code>Columns:\n    name    str\n    scores_4    dict\n    level   int\n    features_4  dict\n\nRows: 34424\n\nData:\n+-------------+-------------------------------+-------+\n|     name    |            scores_4           | level |\n+-------------+-------------------------------+-------+\n|  2157_right | {'contrast-2.1': nan, 'con... |   0   |\n|  8649_right | {'contrast-2.1': nan, 'con... |   0   |\n|  497_right  | {'contrast-2.1': nan, 'con... |   0   |\n| 34024_right | {'contrast-2.1': nan, 'con... |   0   |\n|  2613_left  | {'contrast-2.1': nan, 'con... |   2   |\n|  21769_left | {'contrast-2.1': nan, 'con... |   0   |\n| 25197_right | {'contrast-2.1': nan, 'con... |   2   |\n|  21833_left | {'contrast-2.1': nan, 'con... |   1   |\n|  11417_left | {'contrast-2.1': nan, 'con... |   2   |\n|  5096_right | {'contrast-2.1': nan, 'con... |   0   |\n+-------------+-------------------------------+-------+\n+-------------------------------+\n|           features_4          |\n+-------------------------------+\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n+-------------------------------+\n[34424 rows x 4 columns]\nNote: Only the head of the SFrame is printed.\nYou can use print_rows(num_rows=m, num_columns=n) to print more rows and columns.\n</code></pre>",
      "rawMarkdown": "Hi Hoyt,\r\n\r\nI tried running your code. I did it only to train `which_model = 4` i.e. predicting 4-vs-all retinopathy grade.\r\nMy training and validation accuracy figures were quite low so I checked the features that have been stored in `Xty` and `Xtst` in `create_nn_model.py` and it shows `nan` values in both `scores_4` and `features_4` columns. Any idea about what is it that I did wrong?\r\n\r\nFollowing are the details of the model that was stored:\r\n\r\n       \r\n    Class               : NeuralNetClassifier\r\n    \r\n    Schema\r\n    ------\r\n    Examples            : 848387\r\n    Features            : 1\r\n    Target column       : class\r\n    \r\n    Training Summary\r\n    ----------------\r\n    Training accuracy   : 0.4992\r\n    Validation accuracy : 0.5416\r\n    Training time (sec) : 85850.3034\r\n\r\n\r\nFollowing are the details of `Xf_train` from `create_submission.py` code.\r\n\r\n    Columns:\r\n    \tname\tstr\r\n    \tscores_4\tdict\r\n    \tlevel\tint\r\n    \tfeatures_4\tdict\r\n    \r\n    Rows: 34424\r\n    \r\n    Data:\r\n    +-------------+-------------------------------+-------+\r\n    |     name    |            scores_4           | level |\r\n    +-------------+-------------------------------+-------+\r\n    |  2157_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    |  8649_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    |  497_right  | {'contrast-2.1': nan, 'con... |   0   |\r\n    | 34024_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    |  2613_left  | {'contrast-2.1': nan, 'con... |   2   |\r\n    |  21769_left | {'contrast-2.1': nan, 'con... |   0   |\r\n    | 25197_right | {'contrast-2.1': nan, 'con... |   2   |\r\n    |  21833_left | {'contrast-2.1': nan, 'con... |   1   |\r\n    |  11417_left | {'contrast-2.1': nan, 'con... |   2   |\r\n    |  5096_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    +-------------+-------------------------------+-------+\r\n    +-------------------------------+\r\n    |           features_4          |\r\n    +-------------------------------+\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    +-------------------------------+\r\n    [34424 rows x 4 columns]\r\n    Note: Only the head of the SFrame is printed.\r\n    You can use print_rows(num_rows=m, num_columns=n) to print more rows and columns.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 66589,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/17/2015 00:43:55",
      "content": "<p>Thank you. Could you share the cross validation kappa score of this approach?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66601,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "03/17/2015 01:17:26",
      "content": "<p>Hi RCarson,</p>\n<p>I have not yet run full cross-validation on this model yet to get a final kappa score. &nbsp;My goal was to get something done quickly, but running cross validation and tuning the intermediate parameters would be the next step and likely give a much better result. &nbsp;</p>\n<p>My holdout validation RMSE for the regression problem at the end to predict&nbsp;the level&nbsp;was around 0.63 if that helps.</p>\n<p>-- Hoyt</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66607,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/17/2015 01:36:13",
      "content": "<p>Great! my model gets rmse is 1.8 and LB 0.09. Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66626,
      "author_name": "sushize",
      "author_url": "",
      "post_date": "03/17/2015 03:50:25",
      "content": "<p>Thanks for the great shared code!</p>\n<p>Just one thing which I hope you could clarify (thanks in advance): I noticed that your code used the&nbsp;Graphlab Create, which is not fully open source. Are you using any non-open source elements of the Graphlab Create in your code or not?</p>\n<p>Thanks!</p>\n<p>Best wishes,</p>\n<p>Shize</p>\n<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66654,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "03/17/2015 07:31:20",
      "content": "<p>Hello&nbsp;Shize,</p>\n<p>That's a fair question. I used <a href=\"https://dato.com/products/create/\">Dato's Graphlab Create</a>&nbsp;to prototype different models to quickly get at a good solution, which is what I've posted. However, everything that I ended up using is either open source or can be substituted with another open source package. The SFrames is open sourced in dato-core, the GLC deep learning toolkit generates config files that are compatible with CXXNet, and the boosted tree regression can be translated to use xgboost.</p>\n<p>Thanks!</p>\n<p>-- Hoyt</p>\n<p>[quote=Shize Su;66626]</p>\n<p>Thanks for the great shared code!</p>\n<p>Just one thing which I hope you could clarify (thanks in advance): I noticed that your code used the&nbsp;Graphlab Create, which is not fully open source. Are you using any non-open source elements of the Graphlab Create in your code or not?</p>\n<p>Thanks!</p>\n<p>Best wishes,</p>\n<p>Shize</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67948,
      "author_name": "knarfben",
      "author_url": "",
      "post_date": "03/24/2015 11:50:00",
      "content": "<p>Hello,</p>\n<p>Running the&nbsp;create_image_sframes.py script I get the following errors:</p>\n<p>PROGRESS: Read 163807 images in 745.984 secs speed: 0 file/sec<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/test/18344_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-stretch/train/4400_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-contrast-2/train/6210_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/train/27213_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-sat-1/test/8671_left.jpeg</p>\n\n\n<p>any idea where this might come from?</p>\n<p>thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67987,
      "author_name": "knarfben",
      "author_url": "",
      "post_date": "03/24/2015 15:34:17",
      "content": "<p>[quote=knarfben;67948]</p>\n<p>Hello,</p>\n<p>Running the&nbsp;create_image_sframes.py script I get the following errors:</p>\n<p>PROGRESS: Read 163807 images in 745.984 secs speed: 0 file/sec<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/test/18344_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-stretch/train/4400_right.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-contrast-2/train/6210_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-hue-2/train/27213_left.jpeg<br>PROGRESS: Unknown error reading image file: /data2/DRD/processed/run-sat-1/test/8671_left.jpeg</p>\n<p>any idea where this might come from?</p>\n<p>thanks.</p>\n<p>[/quote]</p>\n\n<p>ok.... not enough disk space, had to change&nbsp;GRAPHLAB_CACHE_FILE_LOCATIONS</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67989,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/24/2015 15:45:20",
      "content": "<p>Thank you Hoyt! it takes my computer a whole week to run this benchmark and it gets 0.44. In the following months I'm going to figure out what is going on and replace all the Dato tool with open source tool. It will be a great learning process!</p>\n<p>Again, Dato is really too good for kaggle. Hope you could publish more analysis, ipython notebook style tutorial, to help us learn your thought process, besides generating full submission. I think most kagglers will be more happy to use Dato as an analysis tool. It is indeed really powerful. Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68033,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "03/24/2015 17:58:21",
      "content": "<p>@rcarson I ran this code for only the original dataset no hue, sat, ... processing and got 0.045. I was just curious what I could get in a day or so. Now, I am running the full set and expect me near you on the LB sometime this week :)</p>\n<p>As for the code, there is no magic to it. I actually think the network is so primitive (much more shallow (5 vs 2 conv layers)&nbsp;than the starter code based on cxxnet on the plankton competition though the fully connected layer has one more hidden layer).</p>\n<p>I think, the almost brute-force style&nbsp;pre-processing of color spaces, trying to balance the dataset, and extracting features from almost all kinds of class combinations are responsible for the higher score. This is evident in my score of only 0.045 on the original dataset (although balanced) vs. 0.44 or 0.46 for all this brute force. Overall, the data pre processing is not systematic rather brute-force and the network architecture is too shallow.</p>\n<p>Mind you although graphlab create net could be optimized, under the hood it is still cxxnet.&nbsp;</p>\n<p>It is a good starting place regardless.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68057,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/24/2015 19:33:59",
      "content": "<p>Thank you Hoyt. A kind reminder, multiple accounts are against Kaggle's rule. Luckily your new account doesn't enter this contest, so don't! and I suggest you delete the new account. You certainly don't need that. This happens to others who are new to kaggle. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68061,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "03/24/2015 19:41:37",
      "content": "<p>@rcarson -- Thank you for pointing this out. &nbsp;I accidently logged in with the wrong google+ account and didn't even notice! &nbsp;I closed that one and will just stick with this one.&nbsp;</p>\n<p>In case the above post gets deleted, here is the text:</p>\n<p>---------------------------------------</p>\n<p>Hello @deep and @rcarson,</p>\n<p>I'm really glad you found my code helpful. I will try to make a notebook regarding my thought process here. I'm at work now, so I can't write much, but I'll try to fill in more when I get home.</p>\n<p>I agree with @deep -- this method is pretty brute force. The guiding principle was to create more observations that varied in the ways you want the NN to ignore. There are also a lot of different things that could be tuned and improved, especially with the NN architectures. I'm pretty new to neural nets, so I'll try playing around with it in the way you suggested.</p>\n<p>I think the other thing to realize here is that this is an ordinal regression problem rather than a classification problem. This motivated using the boosted regression trees over the activation levels in the final layers of the NN at the end to produce the result rather than using the straight NN to classify the images.</p>\n<p>@rcarson -- FYI, a lot of the Dato stuff is open source -- see https://github.com/dato-code/Dato-Core.</p>\n<p>@deep -- I have also found that sometimes what I'm doing here would overfit the training data, resulting in poorer predictions. E.g., it seemed like the NN overfit for the 0,1,2 vs 3,4 case based on the training / validation difference, and leaving that one out of the final regression problem resulted in a slightly better score. Just one thing I found.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68570,
      "author_name": "knarfben",
      "author_url": "",
      "post_date": "03/27/2015 13:25:48",
      "content": "<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>\n\n<p>Hello Hoyt</p>\n\n<p>I ran into the following issue while running&nbsp;create_image_sframes.py:</p>\n<p><code>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>Get roughly equal class representation by duplicating the different levels.<br>Do a poor mans random shuffle<br>Unable to reach server for 3 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 4 consecutive pings. Server is considered dead. Please exit and restart.<br>Traceback (most recent call last):<br> File &quot;/home/ubuntu/drd/create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 4988, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 39, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.<br>Unable to reach server for 5 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 6 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 7 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 8 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 9 consecutive pings. Server is considered dead. Please exit and restart.<br>[INFO] Stopping the server connection.<br></code></p>\n<p>Any idea where this might come from? I ran the script many times and always end up with this error, although previous communication with the server were ok....</p>\n\n<p>thanks&nbsp;</p>\n<p>Frank</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68602,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "03/27/2015 18:02:14",
      "content": "<p>[quote=knarfben;68570]</p>\n<p>Hello Hoyt</p>\n<p>I ran into the following issue while running&nbsp;create_image_sframes.py:</p>\n<p><code>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>[INFO] Starting new HTTP connection (1): d3dou3jr4nway6.cloudfront.net<br>Get roughly equal class representation by duplicating the different levels.<br>Do a poor mans random shuffle<br>Unable to reach server for 3 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 4 consecutive pings. Server is considered dead. Please exit and restart.<br>Traceback (most recent call last):<br> File &quot;/home/ubuntu/drd/create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 4988, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/usr/local/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 39, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.<br>Unable to reach server for 5 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 6 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 7 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 8 consecutive pings. Server is considered dead. Please exit and restart.<br>Unable to reach server for 9 consecutive pings. Server is considered dead. Please exit and restart.<br>[INFO] Stopping the server connection.<br></code></p>\n<p>Any idea where this might come from? I ran the script many times and always end up with this error, although previous communication with the server were ok....</p>\n<p>thanks&nbsp;</p>\n<p>Frank</p>\n<p>[/quote]</p>\n<p>Hello Frank, I also hit this error a couple of times -- hence the &quot;if not os.path.exists(...)&quot; line at the top, but running it a few times seems to get through it. &nbsp;Sorting the SFrame with the images in it seems really expensive, and all the disk IO is causing something to time out. I filed a bug report on this, so hopefully it will get fixed in the next Graphlab create version.</p>\n<p>I'm also working on doing all of the operations up to this point with filenames instead of the actual images, which would be much faster. &nbsp;</p>\n<p>Does it complete any of the rounds before crashing, or does it simply crash on the first one every time? &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68603,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "03/27/2015 18:07:58",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 68903,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "03/29/2015 21:32:30",
      "content": "<p>The g2.2xlarge instance in AWS comes with 60GB of hard disk space. Could you let me know how do you increase this space to accommodate the dataset we are dealing with here?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68904,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "03/29/2015 21:39:20",
      "content": "<p>Create and mount a volume of sufficient size. I think I used 300GB but that was more than I needed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68914,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "03/29/2015 22:39:55",
      "content": "<p>Did you create a EBS volume? I am new to AWS. If you store something in the 60GB instance store that comes with g2.2xlarge, does that data get deleted when you stop your instance? **If so, why would anyone ever use that store**? Second question I have is that when I log into my machine, how do I know if I am accessing the instance store or the EBS? i.e., given a path like /home/ubuntu how to know is this directory is under instance store or EBS?</p>\n\n<p>References:</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/AmazonEBS.html</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-using-volumes.html</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Storage.html</p>\n<p>https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/InstanceStorage.html</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68915,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "03/29/2015 22:57:29",
      "content": "<p>Yes it's an ebs volume. The data in the root device will disappear by default unless you select an option to make it persist when you create the AMI. The idea is to use it for software rather than data.&nbsp;</p>\n<p>To create a volume and attach it to a running instance, select create volume at the volumes screen, and then right click to and select attach volume. It will probably say it's attaching to /dev/sdf but in reality it will be on&nbsp;/dev/xvdf on ubuntu. From your instance, execute</p>\n\n<p><code>sudo mkfs -t ext4&nbsp;/dev/xvdf &nbsp;# (only the first time you use the volume).</code></p>\n<p><code></code><code>sudo mkdir /data # &nbsp;(replace /data with wherever you want to mount the volume).</code></p>\n<p><code>sudo mount /dev/xvdf /data</code><code></code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68947,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "03/30/2015 00:23:41",
      "content": "<p>hey i have one more question. i created a ebs volume and mounted it to /data but am unable to save anything in it because i lack the permissions.&nbsp;</p>\n<p>ubuntu:~$ ls -all /data<br>total 24<br>drwxr-xr-x 3 root root 4096 Mar 30 00:04 .<br>drwxr-xr-x 24 root root 4096 Mar 30 00:12 ..</p>\n<p>what can i do so that root becomes ubuntu? whoami gives ubuntu.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68963,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "03/30/2015 02:29:50",
      "content": "<p>sudo chown ubuntu /data</p>\n<p><code>chmod if necessary</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68980,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "03/30/2015 05:28:02",
      "content": "<p>thanks. that worked.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68986,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "03/30/2015 07:05:40",
      "content": "<p>@Hoyt, just finished running your btb code. I like the intuition of pulling all the features and use GBT for classification instead of just straight softmax from the CNN.</p>\n<p>As you mentioned, the last two comparisons for 0-1-2 vs. 3-4 and 0-1-2-3 vs 4 overfit a lot (training: 0.91 and 0.95 vs. validation around 0.64 and 0.65) for this shallow network. This tells me that the network training params aren't quite optimal. For instance, adding regularization and adaptive learning rate could help.</p>\n<p>Great work overall and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69072,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "03/30/2015 20:00:08",
      "content": "<p>@Hoyt, can you say a few sentences about why you decided to go with just two convolutional layers and a stride of 4? I'm mostly trying to understand why two conv layers might be considered sufficient and whether a stride of 4 (&gt;1) was mostly selected to speed things up.</p>\n<p>This thread has been really instructive to me - thanks to all of you nice folks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69074,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "03/30/2015 20:16:46",
      "content": "<p>@small yellow duck --</p>\n<p>To be honest, I'm by no means the expert in this thread on deep learning -- others probably have better advice to share, but I would be happy to tell you what I was thinking.&nbsp;</p>\n<p>I based the architecture off of the configuration for imagenet, but I cut down a number of the parameters and the number of convolution layers. &nbsp;The main reason for only choosing 2 convolution layers was that I didn't feel this problem needed as much location invariance as the imagenet task, which had more variance in where the relevant features were located. &nbsp;After the processing by the imagemagick convert tool, a lot of relevant features appear in similar&nbsp;places in the image, so I didn't think as many convolution layers would be needed to capture it.</p>\n<p>As for the stride, I found that smaller strides generated out-of-memory errors on my GPU (a GTX 780), so that's the only reason for a stride of 4.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69703,
      "author_name": "danielzhang1990",
      "author_url": "",
      "post_date": "04/05/2015 10:01:15",
      "content": "<p>Hello Hoyt</p>\n<p>Thanks for sharing ~~</p>\n<p>However, I ran into the following issue while running create_nn_model.py:</p>\n<p>File &quot;create_nn_model.py&quot;, line 85, in &lt;module&gt;<br> <strong>mean_image = X_train[&quot;image&quot;].mean()</strong></p>\n<p>graphlab.toolkits._main.ToolkitError: Cannot perform sum or average over images of different sizes. <strong>Found images of total size (ie. width * height * channels) of both 196608 and 65536.</strong> Please use graplab.image_analysis.resize() to make images a uniform size.</p>\n<p>Any idea where this might come from?&nbsp;</p>\n<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69772,
      "author_name": "ajulian",
      "author_url": "",
      "post_date": "04/06/2015 06:29:35",
      "content": "<p>Since 196608 = 65536*3, it seems to me you are mixing 256x256 RGB (3 channels) and 256x256 grey level (1 channel) images, while the NN expects images with the same size AND number of channels.</p>\n<p>Just a guess...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69782,
      "author_name": "pavanikrishna0",
      "author_url": "",
      "post_date": "04/06/2015 08:58:59",
      "content": "<p>hi&nbsp;</p>\n<p>how to start this project</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70063,
      "author_name": "gobrewers14",
      "author_url": "",
      "post_date": "04/09/2015 00:30:36",
      "content": "<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70065,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/09/2015 01:07:31",
      "content": "<p>Better than 3 days before contest end (like in Plankton).</p>\n<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70066,
      "author_name": "jessekolb",
      "author_url": "",
      "post_date": "04/09/2015 01:13:57",
      "content": "<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>\n<p>It may be the top ~7% currently but as others have mentioned it's not overly advanced and surely anyone who spent the length of the competition on their work would've bested it anyways.&nbsp; Additionally, I think we can all agree that openly discussing models that detect diseases has benefits outside of the competition alone :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70071,
      "author_name": "gobrewers14",
      "author_url": "",
      "post_date": "04/09/2015 02:30:17",
      "content": "<p>[quote=J Kolb;70066]</p>\n<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>\n<p>It may be the top ~7% currently but as others have mentioned it's not overly advanced and surely anyone who spent the length of the competition on their work would've bested it anyways.&nbsp; Additionally, I think we can all agree that openly discussing models that detect diseases has benefits outside of the competition alone :-)</p>\n<p>[/quote]</p>\n<p>Just saying, there is currently a user (not me) that is at 0.38 that has 59 submissions. &nbsp;Assuming that &quot;surely anyone who spent the length of the competition on their work would've bested it anyways&quot; is an unfair assumption in my estimation. &nbsp;Maybe that user did they best they possibly could and their work was original and may have been good enough to squeak into the top 25%. &nbsp;Seeing a score posted in the forums that trumps your work has to be extremely disheartening. &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70072,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "04/09/2015 02:38:11",
      "content": "<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70078,
      "author_name": "jessekolb",
      "author_url": "",
      "post_date": "04/09/2015 03:49:40",
      "content": "<p>[quote=James King;70072]</p>\n<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>\n<p>[/quote]</p>\n<p>I can understand the debate over how it affects people's scores, but this analogy doesn't really apply because if it's solely about working on the puzzle, no one makes you open the thread, much less go to his github, set up Graphlab, and do a couple days of processing and machine learning.&nbsp; He stated in the title what was in his thread.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70079,
      "author_name": "jessekolb",
      "author_url": "",
      "post_date": "04/09/2015 03:57:56",
      "content": "<p>John,</p>\n<p>Assuming the converse is true, and the given person would not be able to beat the score of .46 after four months of work, it gives them the ability to peek into what other people are doing, if they wish, and not spend those four months without obtaining that result.&nbsp; They don't have to, however, and can still try things on their own, but at least they have that option.&nbsp; Also, it's well before the deadline so it's not like people are getting upended last minute.&nbsp; I'm not going to go further into whether it is disheartening vs. enlightening because I think this has already been beaten like a dead horse and there is obviously cases where each is true.&nbsp; I do want to also reiterate that this is an active research topic.&nbsp; Would you ask a research group not to publish their results because another research group is still working on the same problem?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70083,
      "author_name": "gobrewers14",
      "author_url": "",
      "post_date": "04/09/2015 04:33:19",
      "content": "<p>[quote=J Kolb;70079]</p>\n<p>John,</p>\n<p>Assuming the converse is true, and the given person would not be able to beat the score of .46 after four months of work, it gives them the ability to peek into what other people are doing, if they wish, and not spend those four months without obtaining that result.&nbsp; They don't have to, however, and can still try things on their own, but at least they have that option.&nbsp; Also, it's well before the deadline so it's not like people are getting upended last minute.&nbsp; I'm not going to go farther into whether it is disheartening vs. enlightening because I think this has already been beaten like a dead horse and there is obviously cases where each is true.&nbsp; I do want to also reiterate that this is an active research topic.&nbsp; Would you ask a research group not to publish their results because another research group is still working on the same problem?</p>\n<p>[/quote]</p>\n<p>My objections have nothing to do with publishing results. At the heart of scientific discovery is collaboration. It&#8217;s what makes this site and science in general, pretty awesome. My objections are with a company whoring out their products at the expense of others hard work. Users have posted successful models, plugging the same company&#8217;s tools, in this, the Otto challenge, and as inversion mentioned, one 3 days before the deadline in the Plankton competition, which easily put you in the top 25%. I am aware I have no results on this site to speak of and am not an influential member of this community, so if my objections are unwarranted, I will certainly acquiesce but I must remind those reading of the Law of Unintended Consequences; I had never heard of Dato until a few weeks ago. Maybe that was a good thing for them &#8230;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70084,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/09/2015 04:34:41",
      "content": "<p>[quote=John Martinez;70063]</p>\n<p>Another Dato &quot;benchmark&quot;. &nbsp;What is the purpose of posting code that puts people in the top ~7% with 3 months left in the competition (this is a sincere question)?</p>\n<p>[/quote]</p>\n<p>Hi John, thanks for your question, and hopefully I can give you a satisfactory answer. &nbsp;<br><br>I posted the code and write-up of my&nbsp;insights into the problem in the hope that others would be able to build on my&nbsp;ideas, incorporate their own, and hopefully come out ahead of either approach individually. &nbsp;I've been a&nbsp;researcher in machine learning for over 9 years -- first in my PhD&nbsp;program, and now at Dato -- and most of my research advances have benefitted greatly&nbsp;from open collaboration and an open&nbsp;exchange of ideas.&nbsp;&nbsp;As a result, I'm always free to share what insights I have in the hopes someone can build on them. &nbsp;I know this isn't necessarily academia, but this problem&nbsp;is still an open research problem, and there are several research papers posted elsewhere in the forum&nbsp;describing techniques that&nbsp;allegedly get a much better kappa score than my method. &nbsp;<br><br>I simply don't have the free time needed to win this contest, so that's not my aim; rather, I see it as an interesting and very difficult research problem that I want to feel like I can contribute to.&nbsp;&nbsp;And, from experience, I'll bet the winning idea is likely going to be a combination of openly shared ideas, published research, and original insights&nbsp;-- but it definitely won't be entirely&nbsp;original ideas.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70085,
      "author_name": "jessekolb",
      "author_url": "",
      "post_date": "04/09/2015 04:51:48",
      "content": "<p>[quote=Hoyt Koepke;68061]</p>\n<p>I think the other thing to realize here is that this is an ordinal regression problem rather than a classification problem. This motivated using the boosted regression trees over the activation levels in the final layers of the NN at the end to produce the result rather than using the straight NN to classify the images.</p>\n<p>[/quote]</p>\n<p>I'm not sure I understand your thought process here.&nbsp; Wouldn't it be possible to use the Neural net for a regression problem and then round the result to the nearest value of 0, 1, 2, 3, or 4 instead of doing a one-vs-all classification? And then you wouldn't use the boosted trees at all.&nbsp; What is the benefit of using the trees then?&nbsp; Does using the NN for regression not work well in practice in this situation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70086,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/09/2015 05:00:15",
      "content": "<p>[quote=John Martinez;70083]</p>\n<p>[quote=J Kolb;70079]</p>\n<p>John,</p>\n<p>Assuming the converse is true, and the given person would not be able to beat the score of .46 after four months of work, it gives them the ability to peek into what other people are doing, if they wish, and not spend those four months without obtaining that result.&nbsp; They don't have to, however, and can still try things on their own, but at least they have that option.&nbsp; Also, it's well before the deadline so it's not like people are getting upended last minute.&nbsp; I'm not going to go farther into whether it is disheartening vs. enlightening because I think this has already been beaten like a dead horse and there is obviously cases where each is true.&nbsp; I do want to also reiterate that this is an active research topic.&nbsp; Would you ask a research group not to publish their results because another research group is still working on the same problem?</p>\n<p>[/quote]</p>\n<p>My objections have nothing to do with publishing results. At the heart of scientific discovery is collaboration. It&#8217;s what makes this site and science in general, pretty awesome. My objections are with a company whoring out their products at the expense of others hard work. Users have posted successful models, plugging the same company&#8217;s tools, in this, the Otto challenge, and as inversion mentioned, one 3 days before the deadline in the Plankton competition, which easily put you in the top 25%. I am aware I have no results on this site to speak of and am not an influential member of this community, so if my objections are unwarranted, I will certainly acquiesce but I must remind those reading of the Law of Unintended Consequences; I had never heard of Dato until a few weeks ago. Maybe that was a good thing for them &#8230;</p>\n<p>[/quote]</p>\n<p>Hi John,&nbsp;</p>\n<p>I hope you don't think that poorly of Dato :-).&nbsp;&nbsp;We're a fairly new startup company originally formed in part to support Graphlab, an open source academic project, and now we have expanded to provide other&nbsp;tools for machine learning research and general data science. &nbsp;Because of our academic roots,&nbsp;we've made sure our stuff here is either open source or completely free for academic/non-commercial use to benefit research-oriented endeavors like this. &nbsp;As mentioned earlier, all of my code can either use our&nbsp;open source offering or be translated&nbsp;into other open source tools. &nbsp;I used it since it's (obviously) what I'm most familiar with, so I find it easier to quickly prototype ideas that way. &nbsp;&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70088,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/09/2015 05:06:51",
      "content": "<p>[quote=J Kolb;70085]</p>\n<p>[quote=Hoyt Koepke;68061]</p>\n<p>I think the other thing to realize here is that this is an ordinal regression problem rather than a classification problem. This motivated using the boosted regression trees over the activation levels in the final layers of the NN at the end to produce the result rather than using the straight NN to classify the images.</p>\n<p>[/quote]</p>\n<p>I'm not sure I understand your thought process here.&nbsp; Wouldn't it be possible to use the Neural net for a regression problem and then round the result to the nearest value of 0, 1, 2, 3, or 4 instead of doing a one-vs-all classification? And then you wouldn't use the boosted trees at all.&nbsp; What is the benefit of using the trees then?&nbsp; Does using the NN for regression not work well in practice in this situation?</p>\n<p>[/quote]</p>\n<p>Hi J Kolb,</p>\n<p>It's likely that you could get NN regression to work in this context, but I haven't really tried. &nbsp;The main reason is 2-fold. &nbsp;First, I had several NN models at the end, and this presented a quick and dirty way to ensemble them together. &nbsp;(I tried other classifiers and other tricks, but this was the only one that worked). &nbsp;Second,&nbsp;it took me a day or so to train each of the neural nets, but only a minute or two to train the boosted tree model at the end. &nbsp;By playing around with the boosted tree parameters, I was able to get a much better final score without spending 8 hours between each attempt.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70177,
      "author_name": "nthanhtam",
      "author_url": "",
      "post_date": "04/10/2015 06:53:49",
      "content": "<p>[quote=zyy1990;69703]</p>\n<p>Hello Hoyt</p>\n<p>Thanks for sharing ~~</p>\n<p>However, I ran into the following issue while running create_nn_model.py:</p>\n<p>File &quot;create_nn_model.py&quot;, line 85, in &lt;module&gt;<br> <strong>mean_image = X_train[&quot;image&quot;].mean()</strong></p>\n<p>graphlab.toolkits._main.ToolkitError: Cannot perform sum or average over images of different sizes. <strong>Found images of total size (ie. width * height * channels) of both 196608 and 65536.</strong> Please use graplab.image_analysis.resize() to make images a uniform size.</p>\n<p>Any idea where this might come from?&nbsp;</p>\n<p>thanks</p>\n<p>[/quote]</p>\n<p>zyy1990: Have&nbsp;you solved this problem?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 70186,
      "author_name": "danielzhang1990",
      "author_url": "",
      "post_date": "04/10/2015 07:29:29",
      "content": "<p>[quote=vnneo;70177]</p>\n<p>[quote=zyy1990;69703]</p>\n<p>Hello Hoyt</p>\n<p>Thanks for sharing ~~</p>\n<p>However, I ran into the following issue while running create_nn_model.py:</p>\n<p>File &quot;create_nn_model.py&quot;, line 85, in &lt;module&gt;<br> <strong>mean_image = X_train[&quot;image&quot;].mean()</strong></p>\n<p>graphlab.toolkits._main.ToolkitError: Cannot perform sum or average over images of different sizes. <strong>Found images of total size (ie. width * height * channels) of both 196608 and 65536.</strong> Please use graplab.image_analysis.resize() to make images a uniform size.</p>\n<p>Any idea where this might come from?&nbsp;</p>\n<p>thanks</p>\n<p>[/quote]</p>\n<p>zyy1990: Have&nbsp;you solved this problem?</p>\n<p>[/quote]</p>\n\n<p>Yes. For some unknown reason, the data i downloaded had many images which were all black, so i picked &nbsp; those images out. Here are the id of those images:</p>\n<p>train/41176_left.jpeg<br>train/34689_left.jpeg<br>train/32253_right.jpeg<br>train/43457_left.jpeg<br>train/1986_left.jpeg<br>test/28544_left.jpeg<br>test/16519_left.jpeg<br>test/38549_left.jpeg<br>test/35433_left.jpeg<br>test/3517_left.jpeg<br>test/35762_left.jpeg</p>\n<p>train/15222_left.jpeg<br>test/16349_left.jpeg</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71221,
      "author_name": "budiansky",
      "author_url": "",
      "post_date": "04/10/2015 15:44:27",
      "content": "<p>[quote=James King;70072]</p>\n<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>\n<p>[/quote]</p>\n\n<p>I agree.</p>\n<p>My opinion was that posting this solution was poor sportsmanship. If part of the motivation was to market a company's tools, then it leads me to have a negative opinion of that company as well. Seeing this solution decreased my motivation to participant in this contest.</p>\n<p>Obviously the original poster did not think feel that way, so what we have right now is a difference of opinion. However, we have an opportunity to resolve this question by means of (yes!) data.</p>\n<p>If lots of people agree with me, then as a community we should discourage this kind of posting, and the site administrators should note that it will be bad for Kaggle if people are turned off by the practice and stop participating. If people think this kind of posting is a good idea, then we should encourage people to post full solutions early and often.</p>\n<p>Please voice your opinion. What do you think?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71224,
      "author_name": "jbools",
      "author_url": "",
      "post_date": "04/10/2015 15:56:11",
      "content": "<p>28 people upvoted the OP... so, we've already had a vote of sorts. Please don't conflate self-promotion with code sharing. Although code sharing&nbsp;is not&nbsp;entirely altruistic, I am fundamentally for it (both in general and specifically in Kaggle competitions). I myself was motivated (rather than discouraged) by this thread. I felt stimulated to brainstorm ways to improve upon the OP's score, yet also comforted that if I failed to get a good model during these next few months, that I could fall back on this shared code. A HUGE thank you to the OP.</p>\n\n<p>My strong opinion is that collaboration is beautiful. If you don't want to look at the shared code, don't look at it. Easy solution. Live and let live.</p>\n\n<p>Perhaps, another question we should ask (for another thread), is at what point does it become rude to hijack someone else's thread simply because you have a different opinion?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71227,
      "author_name": "woolsey",
      "author_url": "",
      "post_date": "04/10/2015 16:26:40",
      "content": "<p>[quote=Michael 7402;71221]Please voice your opinion. What do you think?[/quote]</p>\n<p>Idem datanewb: please stop this flood. (yeah.. self-defeating demand :-) )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71229,
      "author_name": "gobrewers14",
      "author_url": "",
      "post_date": "04/10/2015 16:57:53",
      "content": "<p>[quote=datanewb;71224]</p>\n<p>28 people upvoted the OP... so, we've already had a vote of sorts. Please don't conflate self-promotion with code sharing. Although code sharing&nbsp;is not&nbsp;entirely altruistic, I am fundamentally for it (both in general and specifically in Kaggle competitions). I myself was motivated (rather than discouraged) by this thread. I felt stimulated to brainstorm ways to improve upon the OP's score, yet also comforted that if I failed to get a good model during these next few months, that I could fall back on this shared code. A HUGE thank you to the OP.</p>\n<p>My strong opinion is that collaboration is beautiful. If you don't want to look at the shared code, don't look at it. Easy solution. Live and let live.</p>\n<p>Perhaps, another question we should ask (for another thread), is at what point does it become rude to hijack someone else's thread simply because you have a different opinion?</p>\n<p>[/quote]</p>\n<p>There was no conflation of self-promotion and code-sharing. &nbsp;The evidence is in the forums of those 3 competitions (and maybe others), as I stated. &nbsp;I completely agree, and would assume that Kaggle does as well, that collaboration is a beautiful thing, which is why we are allowed to form&nbsp;Teams.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71248,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/10/2015 20:51:45",
      "content": "<p>[quote=datanewb;71224]</p>\n<p>if I failed to get a good model during these next few months, that I could fall back on this shared code.</p>\n<p>[/quote]</p>\n<p>At the detriment of everyone you passed over on the leader board. Just by copying and pasting code. Great. Now the LB is no longer reflective of what you are capable of.</p>\n<p>This issue could be completely resolved if the Dato folks posted models without optimized hyperparameters, and perhaps even stating what the model is&nbsp;<em>capable</em> of. And let their users work at it a bit.</p>\n<p>In fact, they'd be doing their users a favor.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71249,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/10/2015 21:08:44",
      "content": "<p>This is a hard one ----yes or no to.</p>\n<p>I follow a good many Kaggle Forums and to be open I have learned quite a few things by looking into everything everyone has posted and have learned new ways to compete in this environment.</p>\n<p>Being a Novice here (or anywhere for that matter) for this type of competition I initially had no idea how I was going to use what I know about ML to compete.</p>\n<p>I agree with all of you these type of posing are a type of self promotion.... by thanks to a few of them I am now feeling a bit more confident about being able to participate in these competitions and see nothing wrong as long as we are able to learn something from the post.... even if this particular one seem to be very self-promoting.... IMHO</p>\n<p>I do commend you all for coming out against it to keep self-promotion under control&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71382,
      "author_name": "",
      "author_url": "",
      "post_date": "04/13/2015 12:09:41",
      "content": "<p>Hi Hoyt,</p>\n<p>Thank you for sharing!</p>\n<p>Do you have any data, what is the best benchmark with this base solution so far?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71397,
      "author_name": "",
      "author_url": "",
      "post_date": "04/13/2015 17:55:08",
      "content": "<p>Hello Hoyt,</p>\n<p>I've just started the NN training process on my GTX 980 and since you wrote: &quot;and about 320 images / second on a GTX 780&quot; in the readme, I expected about 20% better performance on my configuration.</p>\n<p>I got 280 img/s.</p>\n<p>Do you have any idea what can be the reason?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71399,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/13/2015 18:16:20",
      "content": "<p>[quote=George Solymosi;71382]</p>\n<p>Hi Hoyt,</p>\n<p>Thank you for sharing!</p>\n<p>Do you have any data, what is the best benchmark with this base solution so far?</p>\n<p>[/quote]</p>\n<p>Hi George,</p>\n<p>No, I'm not sure at this point. &nbsp;It's a pretty crude approach; there are a lot of parameters to tune, and I've since realized there are a lot of things holding back the accuracy that become obvious when you examine the output solutions. &nbsp;In particular, as others have mentioned, the NN architecture here overfits to some parts of the the training data, it doesn't take into account the unbalanced class sizes at prediction time, etc.&nbsp;&nbsp;I got somewhere close to 0.46 with this solution; some tweaks to boosted tree algorithm (unpublished) boosted&nbsp;the score somewhat. &nbsp; I'm sure that more sophistication in this approach, a better NN architecture, and/or ensembling this model with other approaches is needed to get much better than this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71415,
      "author_name": "seanv507",
      "author_url": "",
      "post_date": "04/13/2015 23:06:20",
      "content": "<p>@hoyt, just some comments on the image processing... I have yet to implement something myself</p>\n<p>a) equalisation is normally only done on the intensity component so that you don't distort the colour balance (ie you don't want to stretch each colour channel separately)</p>\n<p><a href=\"http://www.imagemagick.org/script/command-line-options.php#equalize\">http://www.imagemagick.org/script/command-line-options.php#equalize</a></p>\n\n<p>b) i would think that its better to convert to png rather than jpg (which might introduce new jpeg artifacts at 256x256 size?)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71417,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/13/2015 23:11:41",
      "content": "<p>[quote=George Solymosi;71397]</p>\n<p>Hello Hoyt,</p>\n<p>I've just started the NN training process on my GTX 980 and since you wrote: &quot;and about 320 images / second on a GTX 780&quot; in the readme, I expected about 20% better performance on my configuration.</p>\n<p>I got 280 img/s.</p>\n<p>Do you have any idea what can be the reason?</p>\n<p>[/quote]</p>\n<p>Hey George,&nbsp;</p>\n<p>Sorry -- I have no idea. &nbsp;I suspect it also has to do with the other configurations of your system. The only thing I can think of is that it's IO bound in moving data on and off the card -- maybe it's mounted in the x8 PCIe slot instead of the x16 PCIe slot (although that shouldn't make a difference). &nbsp;</p>\n<p>The batch_size parameter controls how many images are loaded onto the graphics card at a time for training. &nbsp;It's possible that adjusting this parameter will help with performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71418,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/13/2015 23:16:15",
      "content": "<p>[quote=Sean;71415]</p>\n<p>@hoyt, just some comments on the image processing... I have yet to implement something myself</p>\n<p>a) equalisation is normally only done on the intensity component so that you don't distort the colour balance (ie you don't want to stretch each colour channel separately)</p>\n<p><a href=\"http://www.imagemagick.org/script/command-line-options.php#equalize\">http://www.imagemagick.org/script/command-line-options.php#equalize</a></p>\n<p>b) i would think that its better to convert to png rather than jpg (which might introduce new jpeg artifacts at 256x256 size?)</p>\n<p>[/quote]</p>\n<p>Hey Sean,&nbsp;</p>\n<p>Thanks for your comments! &nbsp;I called equalize mainly so that the 3 different channels of the NN would be on the same scale, although I'm not sure that would have any significant effect. &nbsp;And yes, png might be better. &nbsp;I'll give that a shot for my latest iteration.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72504,
      "author_name": "",
      "author_url": "",
      "post_date": "04/19/2015 19:26:31",
      "content": "<p>We can fight with each other, but if we can quicker get closer with a bit help of something like Hoyt's shortcut to a better solution for help to prevent and cure people's illness, then I don't think that, that is unfair...</p>\n<p>Consider it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72719,
      "author_name": "",
      "author_url": "",
      "post_date": "04/20/2015 19:06:49",
      "content": "<p>[quote=Hoyt Koepke;66564]</p>\n<p>Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</p>\n<p>[/quote]</p>\n<p>I'm a bit puzzled though, you say you convert images to 256x256 and at the same time&nbsp;you say you avoid scaling the images. Could you kindly clarify this part?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72986,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/21/2015 17:58:09",
      "content": "<p>[quote=K0stIa;72719]</p>\n<p>I'm a bit puzzled though, you say you convert images to 256x256 and at the same time&nbsp;you say you avoid scaling the images. Could you kindly clarify this part?</p>\n<p>[/quote]</p>\n<p>Sorry, I could have been clearer there. &nbsp;What I meant is that I made sure that the aspect ratio of the images and the location of the eyes in the picture stayed the same so that the resulting NN did not have to handle scale and location invariance along with learning to classify the images. &nbsp;</p>\n\n<p>Hope that helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 73007,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/21/2015 18:50:03",
      "content": "<p>You do realize that the beauty of CNN's that scaling and translation does really cause any performance penalties of the CNN's..... However make sure you don't scale away the useful information ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 73184,
      "author_name": "zsaurk",
      "author_url": "",
      "post_date": "04/22/2015 15:24:24",
      "content": "<p>Thanks Dato Inc / Hoyt for trying to democratize Big Data / Deep Learning algorithms. I was curious what pretrained filters the GraphLab's deeplearning algorithm make use of ? Is it Gabor Filters or BSD Caffe or something else ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74244,
      "author_name": "jstaker7",
      "author_url": "",
      "post_date": "04/23/2015 04:23:32",
      "content": "<p>Hi Hoyt,</p>\n<p>I'm&nbsp;a little confused on the performance you mentioned. 300 images/s tells me each epoch takes around 2 min, and 5-10 epochs for each 2-class model comes out to be 20 min of training. How does this translate to about a day (24h) worth of training? Would you mind clarifying? I'm using theano so I'm curious how my performance compares to your cxxnet implementation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74358,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/23/2015 18:12:36",
      "content": "<p>[quote=Michael George Hart;73007]</p>\n<p>You do realize that the beauty of CNN's that scaling and translation does really cause any performance penalties of the CNN's..... However make sure you don't scale away the useful information ;)</p>\n<p>[/quote]</p>\n<p>Hi Michael,&nbsp;</p>\n<p>Yes, true, though in general, when you add more convolution/pooling layers in order to get closer to true scale invariance, you end up needing more data in order to get it to train effectively. &nbsp;I opted for having much fewer convolution layers than is needed for true translation invariance, since it seemed like location in the image could actually be informative. &nbsp;I think there may be something to this approach, as I've been trying to train with deeper architectures since then, but haven't really got significantly better results yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74360,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/23/2015 18:15:58",
      "content": "<p>[quote=saurk;73184]</p>\n<p>Thanks Dato Inc / Hoyt for trying to democratize Big Data / Deep Learning algorithms. I was curious what pretrained filters the GraphLab's deeplearning algorithm make use of ? Is it Gabor Filters or BSD Caffe or something else ?</p>\n<p>[/quote]</p>\n<p>Hello Saurk,</p>\n<p>Thanks! &nbsp;Under the hood, graphlab create doesn't use any pretrained filters. However, the filters in the convolution layers will often end up resembling Gabor filters, but they are learned directly from the data during training. &nbsp;</p>\n<p>I think Caffe will end up doing the same thing...</p>\n<p>-- Hoyt</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74361,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/23/2015 18:16:22",
      "content": "<p>[quote=jstaker7;74244]</p>\n<p>Hi Hoyt,</p>\n<p>I'm&nbsp;a little confused on the performance you mentioned. 300 images/s tells me each epoch takes around 2 min, and 5-10 epochs for each 2-class model comes out to be 20 min of training. How does this translate to about a day (24h) worth of training? Would you mind clarifying? I'm using theano so I'm curious how my performance compares to your cxxnet implementation.</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74364,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/23/2015 18:30:44",
      "content": "<p>[quote=jstaker7;74244]</p>\n<p>Hi Hoyt,</p>\n<p>I'm&nbsp;a little confused on the performance you mentioned. 300 images/s tells me each epoch takes around 2 min, and 5-10 epochs for each 2-class model comes out to be 20 min of training. How does this translate to about a day (24h) worth of training? Would you mind clarifying? I'm using theano so I'm curious how my performance compares to your cxxnet implementation.</p>\n<p>[/quote]</p>\n<p>Hello jstaker,</p>\n<p>The code I use brute-force creates a bunch of different versions of each image, as well as duplicating some of them to&nbsp;balance the classes in the training set. &nbsp;Thus there end up being hundreds of thousands of images trained in each epoch instead of the original 35,000 or so. &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74417,
      "author_name": "jstaker7",
      "author_url": "",
      "post_date": "04/24/2015 00:44:22",
      "content": "<p>[quote=Hoyt Koepke;74358]</p>\n<p>... I opted for having much fewer convolution layers than is needed for true translation invariance, since it seemed like location in the image could actually be informative...</p>\n<p>[/quote]</p>\n\n<p>This challenges my understanding of CNNs. Isn't just feature detection translation invariant? I thought one of the great advantages of CNNs is that it can capture spacial relationships. Am I understanding your comment correctly?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74445,
      "author_name": "ionelhosu",
      "author_url": "",
      "post_date": "04/24/2015 06:36:51",
      "content": "<p>[quote=jstaker7;74417]</p>\n<p>[quote=Hoyt Koepke;74358]</p>\n<p>... I opted for having much fewer convolution layers than is needed for true translation invariance, since it seemed like location in the image could actually be informative...</p>\n<p>[/quote]</p>\n<p>This challenges my understanding of CNNs. Isn't just feature detection translation invariant? I thought one of the great advantages of CNNs is that it can capture spacial relationships. Am I understanding your comment correctly?</p>\n<p>[/quote]</p>\n<p>CNNs can learn not only translation invariance, but even scale and rotation invariances, but for this to happen, those characteristics need to be derived from the training set, i.e. the independence between (position, rotation, scale) of the features and class label needs to be present in the dataset. If, for some reason, some examples contain some features which are positioned constantly in the same general area of the image, even the translation invariance will go away,&nbsp;because the CNN would 'think' in this case that the position of those features should affect the decision in some way, i.e. the position of the feature and the class label would not be independent and identically distributed (i.i.d.) any more. In conclusion, CNNs CAN learn a lot of invariances, but in order for them to learn, those invariances need to be present in the training set, and this is&nbsp;why data augmentation is done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75145,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "04/27/2015 23:47:21",
      "content": "<p>I'm not an expert on CNNs, but my understanding is that the translation invariance comes purely from the pooling layers. &nbsp;Thus CNNs are only truly translation invariant if there are enough pooling layers that sufficiently significant features detected on one side of the image can affect the outcome of any of the nodes before the final fully connected layers. &nbsp;In my original net, I didn't have enough of these layers to achieve this, so it then relied on the fully connected layers at the end to do more of the work.</p>\n<p>Similarly, scale and rotation invariances come from the number of layers and having enough channels in the convolution layers to adequately fit the range of features present in the data.&nbsp;</p>\n<p>The problem, however, is that the more channels and layers you add, the more parameters need to be trained, and thus the more data you need to train the net without overfitting. &nbsp;Ideally, you want a complicated enough structure to capture all relevant features, but a simple enough structure that it doesn't overfit (though dropout and regularization can help with that).&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75242,
      "author_name": "spaceman",
      "author_url": "",
      "post_date": "04/28/2015 13:38:09",
      "content": "<p>Yes pooling helps with the invariance ... A type of averaging mechanism.....</p>\n<p>however it is the Convolution Layer, a, more realistic, model of the visual system, that makes i variances .... Get the spacing right on the convolution will be more useful for this problem than the adding more pooling. ... Most of the work has already done for us by GoogLeNet, AlexNet etc... The trick here is to modify those convolution layers and perhaps tweeting the pooling layers to meet the needs of this particular problem....</p>\n<p>i imagine many people already see this as a solution... That is why I am only using it as a starting point to see how far I can get in the rankings, because the competition will come to who can tune the existing CNN's for a solution to this challenge....&nbsp;</p>\n<p>To win this challenge someone will have to step outside the CNN box since it is such an abvious solution... ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 77297,
      "author_name": "donfulanito",
      "author_url": "",
      "post_date": "05/06/2015 11:02:26",
      "content": "<p>Kaggle has made their <strong>position clear</strong> on this in the competition&nbsp;rules I think: &quot;<em>It's okay to share code if made available to all participants on the forums.&quot;</em></p>\n<p>Why would publishing any competition entry bad or discouraging?</p>\n<p>- it's absolutely *<strong>not</strong>* the same as someone coming and telling you the 'answer to a puzzle' as you say.</p>\n<p>Here, you&nbsp;are not <strong>forced</strong> to read the forums, let alone study/run the code on github.</p>\n<p>I think that *<strong>anyone</strong>* who reads these forums is doing so *<strong>exactly</strong>* because they want to find out what other people are doing.</p>\n<p>- how is<strong> seeing a solution</strong> that scores better than you discouraging? You already can see in the leaderboard people's scores. Obviously those&nbsp;who&nbsp;scored above you have a solution to go with it, there is no surprise there.</p>\n<p>- if there is a <strong>tool that can help</strong>, and it's allowed by the rules (perhaps), why not use it? You say &quot;seeing this solution decreased by motivation to participant [sic] in this contest&quot;,&nbsp;others would say &quot;seeing this solution <strong>increased&nbsp;</strong>my motivation&quot; as they have something to get started with.</p>\n<p>[quote=Michael 7402;71221]</p>\n<p>[quote=James King;70072]</p>\n<p>It's like you're working on a puzzle, and you're really absorbed in it and making progress, and then someone comes along and tells you the answer.</p>\n<p>[/quote]</p>\n<p>I agree.</p>\n<p>My opinion was that posting this solution was poor sportsmanship. If part of the motivation was to market a company's tools, then it leads me to have a negative opinion of that company as well. Seeing this solution decreased my motivation to participant in this contest.</p>\n<p>Obviously the original poster did not think feel that way, so what we have right now is a difference of opinion. However, we have an opportunity to resolve this question by means of (yes!) data.</p>\n<p>If lots of people agree with me, then as a community we should discourage this kind of posting, and the site administrators should note that it will be bad for Kaggle if people are turned off by the practice and stop participating. If people think this kind of posting is a good idea, then we should encourage people to post full solutions early and often.</p>\n<p>Please voice your opinion. What do you think?</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80553,
      "author_name": "rajkumarb",
      "author_url": "",
      "post_date": "06/01/2015 06:46:22",
      "content": "<p>Hi Hoyt,</p>\n<p>Thanks for the code. During prediction, I am getting different prediction levels for the same image in different runs. Please help me to sort out this. The results are as follows for consecutive runs of same set of images.</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.751127591425 |<br>| 11730_right | 0 | 0.358918072124 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.423049250692 |<br>| 11730_right | 0 | 0.326992818052 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.70960257402 |<br>| 11730_right | 0 | 0.357797767757 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.408557624363 |<br>| 11730_right | 0 | 0.355133798881 |<br>+-------------+-------+----------------+</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 81044,
      "author_name": "fisherman",
      "author_url": "",
      "post_date": "06/05/2015 19:35:40",
      "content": "<p>Hi Hoyt,</p>\n<p>Thanks for posting this started code. I am impressed by what Graphlab can do with just a few hundred lines of Python code! I tried to run your Github code but looks like the Graphlab API crashed in sframe sorting - see below and the log file attached. If I comment out the sorting code then it will run fine. Any idea?</p>\n<p>I'm running Ubuntu 14.04 on a desktop Intel Core i7-5820K CPU @ 3.30GHz &#215; 12, with 8GB Memory plus GeForce GTX 970 GPUs.</p>\n<p>=========================</p>\n<p>Traceback (most recent call last):<br> File &quot;create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 5077, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 31, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.</p>\n<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 81046,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "06/05/2015 19:52:09",
      "content": "<p>Hi Rajkumar,</p>\n<p>It turns out that this is a known issue that the NN people here are working on. &nbsp;I'm not sure there's a way around it now. &nbsp;I think if you use the boosted tree regression as a post-processing step on the activation layers, it's a bit more robust.&nbsp;</p>\n<p>-- Hoyt</p>\n\n<p>[quote=Rajkumar;80553]</p>\n<p>Hi Hoyt,</p>\n<p>Thanks for the code. During prediction, I am getting different prediction levels for the same image in different runs. Please help me to sort out this. The results are as follows for consecutive runs of same set of images.</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.751127591425 |<br>| 11730_right | 0 | 0.358918072124 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.423049250692 |<br>| 11730_right | 0 | 0.326992818052 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 1 | 0.70960257402 |<br>| 11730_right | 0 | 0.357797767757 |<br>+-------------+-------+----------------+</p>\n<p>+-------------+-------+----------------+<br>| name | level | actual_level |<br>+-------------+-------+----------------+<br>| 10046_right | 0 | 0.408557624363 |<br>| 11730_right | 0 | 0.355133798881 |<br>+-------------+-------+----------------+</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 81048,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "06/05/2015 20:01:48",
      "content": "<p>Hello fisherman,</p>\n<p>In this case, I kept hitting issues on sort running out of memory, etc. as well. &nbsp;I'm not sure of a workaround except to use a machine with more memory. &nbsp;I did talk with the people here at Dato responsible for the sort stuff, and it's getting better. &nbsp;</p>\n<p>The solution I ended up using was to&nbsp;modify the code so that it just used the file paths up until the final part when it generated the sframes, at which point I load them using a lambda function. &nbsp;This makes it significantly faster and the sort is much easier.</p>\n<p>Something like:&nbsp;</p>\n<p><code>X[&quot;path&quot;] = list(chain(*[[abspath(join(root, f)) for f in files<br></code><code><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if f.endswith('jpeg')]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for root, dir_list, files in os.walk(image_path)]))<br></code><code><br><br></code></code></p>\n<p>Then use something like</p>\n<p><code>X[&quot;image&quot;] = X[&quot;path&quot;].apply(lambda p: gl.Image(p)) </code></p>\n<p>to load it at the end.&nbsp;</p>\n<p>Hope something like that helps!&nbsp;</p>\n<p>-- Hoyt</p>\n<p>[quote=fisherman;81044]</p>\n<p>Hi Hoyt,</p>\n<p>Thanks for posting this started code. I am impressed by what Graphlab can do with just a few hundred lines of Python code! I tried to run your Github code but looks like the Graphlab API crashed in sframe sorting - see below and the log file attached. If I comment out the sorting code then it will run fine. Any idea?</p>\n<p>I'm running Ubuntu 14.04 on a desktop Intel Core i7-5820K CPU @ 3.30GHz &#215; 12, with 8GB Memory plus GeForce GTX 970 GPUs.</p>\n<p>=========================</p>\n<p>Traceback (most recent call last):<br> File &quot;create_image_sframes.py&quot;, line 47, in &lt;module&gt;<br> X_train = X_train.sort(&quot;_random_&quot;)<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/data_structures/sframe.py&quot;, line 5077, in sort<br> return SFrame(_proxy=self.__proxy__.sort(sort_column_names, sort_column_orders))<br> File &quot;/home/ekang/anaconda/lib/python2.7/dist-packages/graphlab/cython/context.py&quot;, line 31, in __exit__<br> raise exc_type(exc_value)<br>RuntimeError: Communication Failure: 113.</p>\n<p>[quote=Hoyt Koepke;66564]</p>\n<p>Hello,&nbsp;</p>\n<p>I just put up code that I used to get a leaderboard&nbsp;score of ~0.46412, and I would like to hear feedback on it. &nbsp;Feel free to use it, and I am open to feedback. &nbsp;I am sure many parts of it can be improved.&nbsp;</p>\n<p><a href=\"https://github.com/hoytak/diabetic-retinopathy-code\">https://github.com/hoytak/diabetic-retinopathy-code</a></p>\n<p>This code uses the ImageMagick convert tool to preprocess the images, then uses the neural net toolkit (basically cxxnet) and the boosted regression trees in Dato's Graphlab create toolkits&nbsp;to build the classifier. &nbsp;</p>\n<p>Main Ideas:</p>\n<ol>\n<li><span style=\"line-height: 1.4\">Use ImageMagick's convert tool to trim off the blank space to the sides of the images, then pad them so that they are all 256x256. Thus the eye is always centered with edges against the edges of the image. I avoided scaling the images to improve the neural net performance.</span></li>\n<li><span style=\"line-height: 1.4\">Create multiple versions of each image varying by hue and contrast and white balance. &nbsp;This step can be simplified to get to a quick though less accurate solution.</span></li>\n<li><span style=\"line-height: 1.4\">Duplicate each class so each class is represented equally, then shuffle the data.</span></li>\n<li><span style=\"line-height: 1.4\">Train several neural nets, one trained to predict the level membership, and another 4 to distinguish 0 vs 1-4, 0-1 vs. 2-4, 0-2 vs. 3-4, and 0-3 vs. 4.</span></li>\n<li><span style=\"line-height: 1.4\">For each image, extract both class predictions and the values of the final neural net layer. Pool all of these features across models and variations of the same images.</span></li>\n<li><span style=\"line-height: 1.4\">Train boosted regression trees on these pooled features to predict the level.</span></li>\n<li><span style=\"line-height: 1.4\">Round to the nearest integer as the class prediction.</span></li>\n</ol>\n<p>I know that each of these steps could definitely be tuned for greater accuracy, and I would love to hear your feedback. In particular, I don't think the neural net architecture is optimal for these images.&nbsp;</p>\n<p>[/quote]</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 82722,
      "author_name": "tommoran",
      "author_url": "",
      "post_date": "06/25/2015 11:12:43",
      "content": "<p>Hi Hoyt, thanks for the startercode to begin with.</p>\n<p>I'm having an issue when I run&nbsp;<span style=\"line-height: 1.4\">create_image_sframes.py</span></p>\n<p>I get the message below repeatedly</p>\n<p>PROGRESS: Unexpected JPEG decode failure file: /mnt/my-data/kaggle-data/processed/run-normal/test/21697_left.jpeg</p>\n<p>Any ideas off-hand how to fix this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 82746,
      "author_name": "hoytkoepke1",
      "author_url": "",
      "post_date": "06/25/2015 17:50:51",
      "content": "<p>Hey Octave1,</p>\n<p>I think it's happening because somehow the shrinking part of the imagemagick script ended up shrinking one of those images to zero. &nbsp;It should be fine if you just delete that image, or rerun the convert command for just that image.&nbsp;</p>\n<p>-- Hoyt</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84211,
      "author_name": "killniu",
      "author_url": "",
      "post_date": "07/12/2015 15:02:54",
      "content": "<p>thank you Hoyt, I can run following your guides.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86129,
      "author_name": "sangramkapre",
      "author_url": "",
      "post_date": "07/20/2015 06:58:23",
      "content": "<p>Hi Hoyt,</p>\n\n<p>I tried running your code. I did it only to train <code>which_model = 4</code> i.e. predicting 4-vs-all retinopathy grade.\nMy training and validation accuracy figures were quite low so I checked the features that have been stored in <code>Xty</code> and <code>Xtst</code> in <code>create_nn_model.py</code> and it shows <code>nan</code> values in both <code>scores_4</code> and <code>features_4</code> columns. Any idea about what is it that I did wrong?</p>\n\n<p>Following are the details of the model that was stored:</p>\n\n<pre><code>Class               : NeuralNetClassifier\n\nSchema\n------\nExamples            : 848387\nFeatures            : 1\nTarget column       : class\n\nTraining Summary\n----------------\nTraining accuracy   : 0.4992\nValidation accuracy : 0.5416\nTraining time (sec) : 85850.3034\n</code></pre>\n\n<p>Following are the details of <code>Xf_train</code> from <code>create_submission.py</code> code.</p>\n\n<pre><code>Columns:\n    name    str\n    scores_4    dict\n    level   int\n    features_4  dict\n\nRows: 34424\n\nData:\n+-------------+-------------------------------+-------+\n|     name    |            scores_4           | level |\n+-------------+-------------------------------+-------+\n|  2157_right | {'contrast-2.1': nan, 'con... |   0   |\n|  8649_right | {'contrast-2.1': nan, 'con... |   0   |\n|  497_right  | {'contrast-2.1': nan, 'con... |   0   |\n| 34024_right | {'contrast-2.1': nan, 'con... |   0   |\n|  2613_left  | {'contrast-2.1': nan, 'con... |   2   |\n|  21769_left | {'contrast-2.1': nan, 'con... |   0   |\n| 25197_right | {'contrast-2.1': nan, 'con... |   2   |\n|  21833_left | {'contrast-2.1': nan, 'con... |   1   |\n|  11417_left | {'contrast-2.1': nan, 'con... |   2   |\n|  5096_right | {'contrast-2.1': nan, 'con... |   0   |\n+-------------+-------------------------------+-------+\n+-------------------------------+\n|           features_4          |\n+-------------------------------+\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n| {'sat-2.23': nan, 'sat-2.2... |\n+-------------------------------+\n[34424 rows x 4 columns]\nNote: Only the head of the SFrame is printed.\nYou can use print_rows(num_rows=m, num_columns=n) to print more rows and columns.\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "66564": "",
    "66589": "",
    "66601": "",
    "66607": "",
    "66626": "",
    "66654": "",
    "67948": "",
    "67987": "",
    "67989": "",
    "68033": "",
    "68057": "",
    "68061": "",
    "68570": "",
    "68602": "",
    "68603": "",
    "68903": "",
    "68904": "",
    "68914": "",
    "68915": "",
    "68947": "",
    "68963": "",
    "68980": "",
    "68986": "",
    "69072": "",
    "69074": "",
    "69703": "",
    "69772": "",
    "69782": "",
    "70063": "",
    "70065": "",
    "70066": "",
    "70071": "",
    "70072": "",
    "70078": "",
    "70079": "",
    "70083": "",
    "70084": "",
    "70085": "",
    "70086": "",
    "70088": "",
    "70177": "",
    "70186": "",
    "71221": "",
    "71224": "",
    "71227": "",
    "71229": "",
    "71248": "",
    "71249": "",
    "71382": "",
    "71397": "",
    "71399": "",
    "71415": "",
    "71417": "",
    "71418": "",
    "72504": "",
    "72719": "",
    "72986": "",
    "73007": "",
    "73184": "",
    "74244": "",
    "74358": "",
    "74360": "",
    "74361": "",
    "74364": "",
    "74417": "",
    "74445": "",
    "75145": "",
    "75242": "",
    "77297": "",
    "80553": "",
    "81044": "",
    "81046": "",
    "81048": "",
    "82722": "",
    "82746": "",
    "84211": "thank you Hoyt, I can run following your guides.",
    "86129": "Hi Hoyt,\r\n\r\nI tried running your code. I did it only to train `which_model = 4` i.e. predicting 4-vs-all retinopathy grade.\r\nMy training and validation accuracy figures were quite low so I checked the features that have been stored in `Xty` and `Xtst` in `create_nn_model.py` and it shows `nan` values in both `scores_4` and `features_4` columns. Any idea about what is it that I did wrong?\r\n\r\nFollowing are the details of the model that was stored:\r\n\r\n       \r\n    Class               : NeuralNetClassifier\r\n    \r\n    Schema\r\n    ------\r\n    Examples            : 848387\r\n    Features            : 1\r\n    Target column       : class\r\n    \r\n    Training Summary\r\n    ----------------\r\n    Training accuracy   : 0.4992\r\n    Validation accuracy : 0.5416\r\n    Training time (sec) : 85850.3034\r\n\r\n\r\nFollowing are the details of `Xf_train` from `create_submission.py` code.\r\n\r\n    Columns:\r\n    \tname\tstr\r\n    \tscores_4\tdict\r\n    \tlevel\tint\r\n    \tfeatures_4\tdict\r\n    \r\n    Rows: 34424\r\n    \r\n    Data:\r\n    +-------------+-------------------------------+-------+\r\n    |     name    |            scores_4           | level |\r\n    +-------------+-------------------------------+-------+\r\n    |  2157_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    |  8649_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    |  497_right  | {'contrast-2.1': nan, 'con... |   0   |\r\n    | 34024_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    |  2613_left  | {'contrast-2.1': nan, 'con... |   2   |\r\n    |  21769_left | {'contrast-2.1': nan, 'con... |   0   |\r\n    | 25197_right | {'contrast-2.1': nan, 'con... |   2   |\r\n    |  21833_left | {'contrast-2.1': nan, 'con... |   1   |\r\n    |  11417_left | {'contrast-2.1': nan, 'con... |   2   |\r\n    |  5096_right | {'contrast-2.1': nan, 'con... |   0   |\r\n    +-------------+-------------------------------+-------+\r\n    +-------------------------------+\r\n    |           features_4          |\r\n    +-------------------------------+\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    | {'sat-2.23': nan, 'sat-2.2... |\r\n    +-------------------------------+\r\n    [34424 rows x 4 columns]\r\n    Note: Only the head of the SFrame is printed.\r\n    You can use print_rows(num_rows=m, num_columns=n) to print more rows and columns."
  },
  "source": "meta"
}