{
  "id": 19336,
  "title": "Using caffe for this competition",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19336",
  "author_name": "",
  "post_date": "2016-03-05T21:58:37.317Z",
  "votes": null,
  "comment_count": 4,
  "views": 1079,
  "content": "<p>So in the  very popular End to End Deep Learning tutorial with MXNet post, the input data seems to be 30x64x64 (which is basically 30 frames of a slice of image)</p>\n\n<p>The output is a vector with 600 entries because the required prediction result should be in CDF. </p>\n\n<p>Now, I'm trying to apply the same architecture in caffe. In caffe however, the input data dimension is 4D (batch size* channel * height * width)</p>\n\n<p>FIRST QUESTION: \nI know that 64 should be the height and width. But should the 30 be the number of channels or the batch size?</p>\n\n<p>I am guessing that 30 should be the channel number because if it were the batch size then each iteration would generate 30 outputs? (Correct me if I'm wrong please!)</p>\n\n<p>SECOND QUESTION:\nNow since the output is a 1D array but Caffe does everything in 4D, how should I prepare the label data to feed into Caffe as a 4D array? Should it be (600*1*1*1) or (1*600*1*1) or  (1*1*600*1) or (1*1*1*600)?   </p>",
  "messages": [
    {
      "id": "110481",
      "postDate": "03/05/2016 21:58:37",
      "content": "<p>So in the  very popular End to End Deep Learning tutorial with MXNet post, the input data seems to be 30x64x64 (which is basically 30 frames of a slice of image)</p>\n\n<p>The output is a vector with 600 entries because the required prediction result should be in CDF. </p>\n\n<p>Now, I'm trying to apply the same architecture in caffe. In caffe however, the input data dimension is 4D (batch size* channel * height * width)</p>\n\n<p>FIRST QUESTION: \nI know that 64 should be the height and width. But should the 30 be the number of channels or the batch size?</p>\n\n<p>I am guessing that 30 should be the channel number because if it were the batch size then each iteration would generate 30 outputs? (Correct me if I'm wrong please!)</p>\n\n<p>SECOND QUESTION:\nNow since the output is a 1D array but Caffe does everything in 4D, how should I prepare the label data to feed into Caffe as a 4D array? Should it be (600*1*1*1) or (1*600*1*1) or  (1*1*600*1) or (1*1*1*600)?   </p>",
      "rawMarkdown": "So in the  very popular End to End Deep Learning tutorial with MXNet post, the input data seems to be 30x64x64 (which is basically 30 frames of a slice of image)\r\n\r\nThe output is a vector with 600 entries because the required prediction result should be in CDF. \r\n\r\nNow, I'm trying to apply the same architecture in caffe. In caffe however, the input data dimension is 4D (batch size* channel * height * width)\r\n\r\nFIRST QUESTION: \r\nI know that 64 should be the height and width. But should the 30 be the number of channels or the batch size?\r\n\r\nI am guessing that 30 should be the channel number because if it were the batch size then each iteration would generate 30 outputs? (Correct me if I'm wrong please!)\r\n\r\nSECOND QUESTION:\r\nNow since the output is a 1D array but Caffe does everything in 4D, how should I prepare the label data to feed into Caffe as a 4D array? Should it be (600*1*1*1) or (1*600*1*1) or  (1*1*600*1) or (1*1*1*600)?",
      "votes": null
    },
    {
      "id": "110903",
      "postDate": "03/09/2016 14:42:56",
      "content": "<p>You can input to caffe a 30 channel by 64 height by 64 width image, basically a stack of pixel arrays from one sax. A batch of size N of images like that gives blob size Nx30x64x64. </p>\n\n<p>You should be able to load label data in Nx600 batches - batch size x number of labels. Details here: <a href=\"https://github.com/BVLC/caffe/pull/2049\">https://github.com/BVLC/caffe/pull/2049</a> and here: <a href=\"https://github.com/BVLC/caffe/issues/1341\">https://github.com/BVLC/caffe/issues/1341</a></p>",
      "rawMarkdown": "You can input to caffe a 30 channel by 64 height by 64 width image, basically a stack of pixel arrays from one sax. A batch of size N of images like that gives blob size Nx30x64x64. \r\n\r\nYou should be able to load label data in Nx600 batches - batch size x number of labels. Details here: https://github.com/BVLC/caffe/pull/2049 and here: https://github.com/BVLC/caffe/issues/1341",
      "votes": null
    },
    {
      "id": "110935",
      "postDate": "03/09/2016 19:10:08",
      "content": "<p>About labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width]) are necessary for convolutional layers.</p>",
      "rawMarkdown": "About labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width]) are necessary for convolutional layers.",
      "votes": null
    },
    {
      "id": "110937",
      "postDate": "03/09/2016 19:18:08",
      "content": "<p>Hello Dimtry, </p>\n\n<p>Yes I was able to train a small network with labels formated as [number of training examples x 600]. However I'm not sure which loss layer I should use. </p>\n\n<p>What loss layer would you say that is the most appropriate for this type of labels? I'm using Euclidean distance right now. I know Bing Xu's post created his own loss function. But I'm wondering if there are off the shelf loss layer from Caffe that work well also.</p>\n\n<p>I also read somewhere that fully connected layer + euclidean distance is basically a linear regression problem? But it's only least square for the last two layers. How can I train the previous layers using SGD from Caffe and somehow use traditional least square methods to solve the weight for the last layer?</p>\n\n<p>[quote=Dmitry Efimov;110935]</p>\n\n<p>About labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width] are necessary for convolutional layers.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Hello Dimtry, \r\n\r\nYes I was able to train a small network with labels formated as [number of training examples x 600]. However I'm not sure which loss layer I should use. \r\n\r\nWhat loss layer would you say that is the most appropriate for this type of labels? I'm using Euclidean distance right now. I know Bing Xu's post created his own loss function. But I'm wondering if there are off the shelf loss layer from Caffe that work well also.\r\n\r\nI also read somewhere that fully connected layer + euclidean distance is basically a linear regression problem? But it's only least square for the last two layers. How can I train the previous layers using SGD from Caffe and somehow use traditional least square methods to solve the weight for the last layer?\r\n\r\n[quote=Dmitry Efimov;110935]\r\n\r\nAbout labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width] are necessary for convolutional layers.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "111001",
      "postDate": "03/10/2016 10:46:51",
      "content": "<p>You can try these last two layers:</p>\n\n<pre><code>layer {\n  name: &quot;full&quot;\n  type: &quot;InnerProduct&quot;\n  bottom: &quot;name_of_previous_layer&quot;\n  top: &quot;full&quot;\n  inner_product_param {\n    num_output: 600\n    weight_filler {\n      type: &quot;xavier&quot;\n    }\n  }\n}\nlayer {\n  name: &quot;loss&quot;\n  type: &quot;EuclideanLoss&quot;\n  bottom: &quot;full&quot;\n  bottom: &quot;label&quot;\n  top: &quot;loss&quot;\n}\n</code></pre>\n\n<p>About the SGD there are two different things implemented in caffe: loss and accuracy. It uses loss to propagate along the neural network and update weights, accuracy is necessary if you want to see the accuracy only. When you specify the type of layer EuclideanLoss if automatically uses SGD to update weights. If you want you can add one more layer for Accuracy which will be not propagated.</p>",
      "rawMarkdown": "You can try these last two layers:\r\n\r\n    layer {\r\n      name: \"full\"\r\n      type: \"InnerProduct\"\r\n      bottom: \"name_of_previous_layer\"\r\n      top: \"full\"\r\n      inner_product_param {\r\n        num_output: 600\r\n        weight_filler {\r\n          type: \"xavier\"\r\n        }\r\n      }\r\n    }\r\n    layer {\r\n      name: \"loss\"\r\n      type: \"EuclideanLoss\"\r\n      bottom: \"full\"\r\n      bottom: \"label\"\r\n      top: \"loss\"\r\n    }\r\n\r\nAbout the SGD there are two different things implemented in caffe: loss and accuracy. It uses loss to propagate along the neural network and update weights, accuracy is necessary if you want to see the accuracy only. When you specify the type of layer EuclideanLoss if automatically uses SGD to update weights. If you want you can add one more layer for Accuracy which will be not propagated.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 110903,
      "author_name": "jwjohnson314",
      "author_url": "",
      "post_date": "03/09/2016 14:42:56",
      "content": "<p>You can input to caffe a 30 channel by 64 height by 64 width image, basically a stack of pixel arrays from one sax. A batch of size N of images like that gives blob size Nx30x64x64. </p>\n\n<p>You should be able to load label data in Nx600 batches - batch size x number of labels. Details here: <a href=\"https://github.com/BVLC/caffe/pull/2049\">https://github.com/BVLC/caffe/pull/2049</a> and here: <a href=\"https://github.com/BVLC/caffe/issues/1341\">https://github.com/BVLC/caffe/issues/1341</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110935,
      "author_name": "efimov",
      "author_url": "",
      "post_date": "03/09/2016 19:10:08",
      "content": "<p>About labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width]) are necessary for convolutional layers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110937,
      "author_name": "ilovesolder",
      "author_url": "",
      "post_date": "03/09/2016 19:18:08",
      "content": "<p>Hello Dimtry, </p>\n\n<p>Yes I was able to train a small network with labels formated as [number of training examples x 600]. However I'm not sure which loss layer I should use. </p>\n\n<p>What loss layer would you say that is the most appropriate for this type of labels? I'm using Euclidean distance right now. I know Bing Xu's post created his own loss function. But I'm wondering if there are off the shelf loss layer from Caffe that work well also.</p>\n\n<p>I also read somewhere that fully connected layer + euclidean distance is basically a linear regression problem? But it's only least square for the last two layers. How can I train the previous layers using SGD from Caffe and somehow use traditional least square methods to solve the weight for the last layer?</p>\n\n<p>[quote=Dmitry Efimov;110935]</p>\n\n<p>About labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width] are necessary for convolutional layers.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111001,
      "author_name": "efimov",
      "author_url": "",
      "post_date": "03/10/2016 10:46:51",
      "content": "<p>You can try these last two layers:</p>\n\n<pre><code>layer {\n  name: &quot;full&quot;\n  type: &quot;InnerProduct&quot;\n  bottom: &quot;name_of_previous_layer&quot;\n  top: &quot;full&quot;\n  inner_product_param {\n    num_output: 600\n    weight_filler {\n      type: &quot;xavier&quot;\n    }\n  }\n}\nlayer {\n  name: &quot;loss&quot;\n  type: &quot;EuclideanLoss&quot;\n  bottom: &quot;full&quot;\n  bottom: &quot;label&quot;\n  top: &quot;loss&quot;\n}\n</code></pre>\n\n<p>About the SGD there are two different things implemented in caffe: loss and accuracy. It uses loss to propagate along the neural network and update weights, accuracy is necessary if you want to see the accuracy only. When you specify the type of layer EuclideanLoss if automatically uses SGD to update weights. If you want you can add one more layer for Accuracy which will be not propagated.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "110481": "So in the  very popular End to End Deep Learning tutorial with MXNet post, the input data seems to be 30x64x64 (which is basically 30 frames of a slice of image)\r\n\r\nThe output is a vector with 600 entries because the required prediction result should be in CDF. \r\n\r\nNow, I'm trying to apply the same architecture in caffe. In caffe however, the input data dimension is 4D (batch size* channel * height * width)\r\n\r\nFIRST QUESTION: \r\nI know that 64 should be the height and width. But should the 30 be the number of channels or the batch size?\r\n\r\nI am guessing that 30 should be the channel number because if it were the batch size then each iteration would generate 30 outputs? (Correct me if I'm wrong please!)\r\n\r\nSECOND QUESTION:\r\nNow since the output is a 1D array but Caffe does everything in 4D, how should I prepare the label data to feed into Caffe as a 4D array? Should it be (600*1*1*1) or (1*600*1*1) or  (1*1*600*1) or (1*1*1*600)?",
    "110903": "You can input to caffe a 30 channel by 64 height by 64 width image, basically a stack of pixel arrays from one sax. A batch of size N of images like that gives blob size Nx30x64x64. \r\n\r\nYou should be able to load label data in Nx600 batches - batch size x number of labels. Details here: https://github.com/BVLC/caffe/pull/2049 and here: https://github.com/BVLC/caffe/issues/1341",
    "110935": "About labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width]) are necessary for convolutional layers.",
    "110937": "Hello Dimtry, \r\n\r\nYes I was able to train a small network with labels formated as [number of training examples x 600]. However I'm not sure which loss layer I should use. \r\n\r\nWhat loss layer would you say that is the most appropriate for this type of labels? I'm using Euclidean distance right now. I know Bing Xu's post created his own loss function. But I'm wondering if there are off the shelf loss layer from Caffe that work well also.\r\n\r\nI also read somewhere that fully connected layer + euclidean distance is basically a linear regression problem? But it's only least square for the last two layers. How can I train the previous layers using SGD from Caffe and somehow use traditional least square methods to solve the weight for the last layer?\r\n\r\n[quote=Dmitry Efimov;110935]\r\n\r\nAbout labels, you can use array [number of training examples x 600], because your last layer should be the Fully Connected layer and it accepts any type of data. If I understood correctly, blobs (array [batch size x channels x height x width] are necessary for convolutional layers.\r\n\r\n[/quote]",
    "111001": "You can try these last two layers:\r\n\r\n    layer {\r\n      name: \"full\"\r\n      type: \"InnerProduct\"\r\n      bottom: \"name_of_previous_layer\"\r\n      top: \"full\"\r\n      inner_product_param {\r\n        num_output: 600\r\n        weight_filler {\r\n          type: \"xavier\"\r\n        }\r\n      }\r\n    }\r\n    layer {\r\n      name: \"loss\"\r\n      type: \"EuclideanLoss\"\r\n      bottom: \"full\"\r\n      bottom: \"label\"\r\n      top: \"loss\"\r\n    }\r\n\r\nAbout the SGD there are two different things implemented in caffe: loss and accuracy. It uses loss to propagate along the neural network and update weights, accuracy is necessary if you want to see the accuracy only. When you specify the type of layer EuclideanLoss if automatically uses SGD to update weights. If you want you can add one more layer for Accuracy which will be not propagated."
  },
  "source": "meta"
}