{
  "id": 398179,
  "title": "Help with CNN architecture",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/398179",
  "author_name": "Nick Potter",
  "post_date": "2023-03-28T23:14:34.294000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone, I'm trying to get a simple (actually any) tensorflow/keras CNN model up and running, but without any luck. Training doesn't converge to anything. Output is either random noise or a constant value regardless of input.</p>\n<p>Francois Chollet's <a href=\"https://www.kaggle.com/code/fchollet/a-simple-high-performance-tf-data-pipeline\" target=\"_blank\">excellent data pipeline notebook</a> essentially gives us batches of data with the dimensions:</p>\n<p><code>(BATCH_SIZE, X_DIM, Y_DIM, Z_DIM)</code> where X_DIM and Y_DIM are the size of the window, and Z_DIM gives us the height of the stack of images. </p>\n<p>I have tried Conv3d layers, e.g. with kernel size <code>(Z_DIM, Z_DIM, Z_DIM)</code>, which should return a tensor of size similar to <code>(BATCH_SIZE, X_DIM, Y_DIM, N_FILTERS)</code> and then dense layers with and without pooling beforehand.</p>\n<p>Also I tried Conv2d layers (not sure what the resulting dimension ends up being here).</p>\n<p>I use sigmoid activation function for the final layer.</p>\n<p>A couple of questions about this architecture compared to the <a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">example pytorch CNN model</a>:</p>\n<ol>\n<li><p>The pytorch model looks like it has a single pixel as output, whereas I thought the model would/should be returning an image window sized equal to the input window?</p></li>\n<li><p>The loss function is binary cross entropy, I have used this as well as MSE, which is better?</p></li>\n<li><p>What is the difference between using conv2d and conv3d layers? Can we use conv2d layers? What about if Z_DIM=1 (i.e. we are predicting just based on a single tiff image from the stack)?</p></li>\n</ol>",
  "messages": [
    {
      "id": 2200953,
      "postDate": "2023-03-28T23:35:28.053Z",
      "content": "<p>Without your code being shared it's difficult to be of real value for your issues.  So not likely you will get the 'fix' unless you share your code.</p>\n<ol>\n<li>I think the pytorch model is a single pixel - which is a reason it's so slow - I gave up trying to get a LB score with it.</li>\n<li>Using the pipeline you referenced I was successful in getting the other shared <a href=\"https://www.kaggle.com/code/fchollet/keras-starter-kit-unet-train-on-full-dataset\" target=\"_blank\">keras</a> model to run.  Did you try that model?</li>\n<li>I did have 2 out of 12 versions of model not train.  Both had 512x512 windows but not sure of actual root cause.  There seems to be a sweet spot in window size vs local cv.</li>\n<li>When Z_DIM = 1 you are using just a single tiff.  Not sure if the tiff's in our supplied data are 4 or 8 micron thick slices, but pretty sure a single tiff is thinner than the ink layer.  Since the fragments not perfectly flat a single tif will not likely work.  For a model that I used there is a relationship between local cv accuracy and number of layers used.</li>\n<li>For ink on/off sigmoid is right and I also used 'binary_crossentropy'.</li>\n</ol>",
      "rawMarkdown": "Without your code being shared it's difficult to be of real value for your issues.  So not likely you will get the 'fix' unless you share your code.\n\n1.  I think the pytorch model is a single pixel - which is a reason it's so slow - I gave up trying to get a LB score with it.\n2.  Using the pipeline you referenced I was successful in getting the other shared [keras](https://www.kaggle.com/code/fchollet/keras-starter-kit-unet-train-on-full-dataset) model to run.  Did you try that model?\n3.  I did have 2 out of 12 versions of model not train.  Both had 512x512 windows but not sure of actual root cause.  There seems to be a sweet spot in window size vs local cv.\n4. When Z_DIM = 1 you are using just a single tiff.  Not sure if the tiff's in our supplied data are 4 or 8 micron thick slices, but pretty sure a single tiff is thinner than the ink layer.  Since the fragments not perfectly flat a single tif will not likely work.  For a model that I used there is a relationship between local cv accuracy and number of layers used.\n5.  For ink on/off sigmoid is right and I also used 'binary_crossentropy'.",
      "votes": 1,
      "replies": [
        {
          "id": 2200955,
          "postDate": "2023-03-28T23:37:21.473Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcjimmy\" target=\"_blank\">@pcjimmy</a>. This is really useful.</p>",
          "rawMarkdown": "Thanks @pcjimmy. This is really useful."
        }
      ]
    },
    {
      "id": 2200930,
      "postDate": "2023-03-28T23:14:34.293Z",
      "content": "<p>Hi everyone, I'm trying to get a simple (actually any) tensorflow/keras CNN model up and running, but without any luck. Training doesn't converge to anything. Output is either random noise or a constant value regardless of input.</p>\n<p>Francois Chollet's <a href=\"https://www.kaggle.com/code/fchollet/a-simple-high-performance-tf-data-pipeline\" target=\"_blank\">excellent data pipeline notebook</a> essentially gives us batches of data with the dimensions:</p>\n<p><code>(BATCH_SIZE, X_DIM, Y_DIM, Z_DIM)</code> where X_DIM and Y_DIM are the size of the window, and Z_DIM gives us the height of the stack of images. </p>\n<p>I have tried Conv3d layers, e.g. with kernel size <code>(Z_DIM, Z_DIM, Z_DIM)</code>, which should return a tensor of size similar to <code>(BATCH_SIZE, X_DIM, Y_DIM, N_FILTERS)</code> and then dense layers with and without pooling beforehand.</p>\n<p>Also I tried Conv2d layers (not sure what the resulting dimension ends up being here).</p>\n<p>I use sigmoid activation function for the final layer.</p>\n<p>A couple of questions about this architecture compared to the <a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">example pytorch CNN model</a>:</p>\n<ol>\n<li><p>The pytorch model looks like it has a single pixel as output, whereas I thought the model would/should be returning an image window sized equal to the input window?</p></li>\n<li><p>The loss function is binary cross entropy, I have used this as well as MSE, which is better?</p></li>\n<li><p>What is the difference between using conv2d and conv3d layers? Can we use conv2d layers? What about if Z_DIM=1 (i.e. we are predicting just based on a single tiff image from the stack)?</p></li>\n</ol>",
      "rawMarkdown": "Hi everyone, I'm trying to get a simple (actually any) tensorflow/keras CNN model up and running, but without any luck. Training doesn't converge to anything. Output is either random noise or a constant value regardless of input.\n\nFrancois Chollet's [excellent data pipeline notebook](https://www.kaggle.com/code/fchollet/a-simple-high-performance-tf-data-pipeline) essentially gives us batches of data with the dimensions:\n\n`(BATCH_SIZE, X_DIM, Y_DIM, Z_DIM)` where X_DIM and Y_DIM are the size of the window, and Z_DIM gives us the height of the stack of images. \n\nI have tried Conv3d layers, e.g. with kernel size `(Z_DIM, Z_DIM, Z_DIM)`, which should return a tensor of size similar to `(BATCH_SIZE, X_DIM, Y_DIM, N_FILTERS)` and then dense layers with and without pooling beforehand.\n\nAlso I tried Conv2d layers (not sure what the resulting dimension ends up being here).\n\nI use sigmoid activation function for the final layer.\n\nA couple of questions about this architecture compared to the [example pytorch CNN model](https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial):\n\n1. The pytorch model looks like it has a single pixel as output, whereas I thought the model would/should be returning an image window sized equal to the input window?\n\n2. The loss function is binary cross entropy, I have used this as well as MSE, which is better?\n\n3. What is the difference between using conv2d and conv3d layers? Can we use conv2d layers? What about if Z_DIM=1 (i.e. we are predicting just based on a single tiff image from the stack)?",
      "votes": 2
    },
    {
      "id": 2204895,
      "postDate": "2023-04-01T04:26:37.613Z",
      "content": "<p>I tell my students to do everything in PyTorch. Google PyTorch/Keras/TF trend, PyTorch is dominating and increasing. The trend is your friend :).</p>",
      "rawMarkdown": "I tell my students to do everything in PyTorch. Google PyTorch/Keras/TF trend, PyTorch is dominating and increasing. The trend is your friend :).",
      "replies": [
        {
          "id": 2205879,
          "postDate": "2023-04-02T03:47:27.313Z",
          "content": "<p>Nice, but from some software engineers in the field, I have heard that TF is very popular in the industry. One (who taught us a course) said that he prefers PyTorch, but since many ML experts in the industry have expertise in it, it is often more used, so it's better to learn that.</p>",
          "rawMarkdown": "Nice, but from some software engineers in the field, I have heard that TF is very popular in the industry. One (who taught us a course) said that he prefers PyTorch, but since many ML experts in the industry have expertise in it, it is often more used, so it's better to learn that."
        },
        {
          "id": 2205990,
          "postDate": "2023-04-02T06:32:28.960Z",
          "content": "<p>I have been telling my friends that bicycles and horses are dominating.  Was wrong for nearly 75 years but I see things turning a corner.</p>\n<p>The trend seldom stays the trend over a lifetime.</p>",
          "rawMarkdown": "I have been telling my friends that bicycles and horses are dominating.  Was wrong for nearly 75 years but I see things turning a corner.\n\nThe trend seldom stays the trend over a lifetime."
        }
      ]
    },
    {
      "id": 2217833,
      "postDate": "2023-04-11T07:45:36.880Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2218083,
          "postDate": "2023-04-11T12:13:56.377Z",
          "content": "<p>Be careful with such a nice account<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Fc1a11c7aa205d8b20aacd95015fd0b03%2FScreenshot_2.png?generation=1681215192581582&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Be careful with such a nice account\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Fc1a11c7aa205d8b20aacd95015fd0b03%2FScreenshot_2.png?generation=1681215192581582&alt=media)",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2200953,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-03-28T23:35:28.053000",
      "content": "<p>Without your code being shared it's difficult to be of real value for your issues.  So not likely you will get the 'fix' unless you share your code.</p>\n<ol>\n<li>I think the pytorch model is a single pixel - which is a reason it's so slow - I gave up trying to get a LB score with it.</li>\n<li>Using the pipeline you referenced I was successful in getting the other shared <a href=\"https://www.kaggle.com/code/fchollet/keras-starter-kit-unet-train-on-full-dataset\" target=\"_blank\">keras</a> model to run.  Did you try that model?</li>\n<li>I did have 2 out of 12 versions of model not train.  Both had 512x512 windows but not sure of actual root cause.  There seems to be a sweet spot in window size vs local cv.</li>\n<li>When Z_DIM = 1 you are using just a single tiff.  Not sure if the tiff's in our supplied data are 4 or 8 micron thick slices, but pretty sure a single tiff is thinner than the ink layer.  Since the fragments not perfectly flat a single tif will not likely work.  For a model that I used there is a relationship between local cv accuracy and number of layers used.</li>\n<li>For ink on/off sigmoid is right and I also used 'binary_crossentropy'.</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 2200955,
          "author_name": "Nick Potter",
          "author_url": "",
          "post_date": "2023-03-28T23:37:21.473000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcjimmy\" target=\"_blank\">@pcjimmy</a>. This is really useful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2204895,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "2023-04-01T04:26:37.613000",
      "content": "<p>I tell my students to do everything in PyTorch. Google PyTorch/Keras/TF trend, PyTorch is dominating and increasing. The trend is your friend :).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2205879,
          "author_name": "Grim_Reaper",
          "author_url": "",
          "post_date": "2023-04-02T03:47:27.313000",
          "content": "<p>Nice, but from some software engineers in the field, I have heard that TF is very popular in the industry. One (who taught us a course) said that he prefers PyTorch, but since many ML experts in the industry have expertise in it, it is often more used, so it's better to learn that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2205990,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2023-04-02T06:32:28.960000",
          "content": "<p>I have been telling my friends that bicycles and horses are dominating.  Was wrong for nearly 75 years but I see things turning a corner.</p>\n<p>The trend seldom stays the trend over a lifetime.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2217833,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-11T07:45:36.880000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2218083,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-04-11T12:13:56.377000",
          "content": "<p>Be careful with such a nice account<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Fc1a11c7aa205d8b20aacd95015fd0b03%2FScreenshot_2.png?generation=1681215192581582&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2200953": "Without your code being shared it's difficult to be of real value for your issues.  So not likely you will get the 'fix' unless you share your code.\n\n1.  I think the pytorch model is a single pixel - which is a reason it's so slow - I gave up trying to get a LB score with it.\n2.  Using the pipeline you referenced I was successful in getting the other shared [keras](https://www.kaggle.com/code/fchollet/keras-starter-kit-unet-train-on-full-dataset) model to run.  Did you try that model?\n3.  I did have 2 out of 12 versions of model not train.  Both had 512x512 windows but not sure of actual root cause.  There seems to be a sweet spot in window size vs local cv.\n4. When Z_DIM = 1 you are using just a single tiff.  Not sure if the tiff's in our supplied data are 4 or 8 micron thick slices, but pretty sure a single tiff is thinner than the ink layer.  Since the fragments not perfectly flat a single tif will not likely work.  For a model that I used there is a relationship between local cv accuracy and number of layers used.\n5.  For ink on/off sigmoid is right and I also used 'binary_crossentropy'.",
    "2200930": "Hi everyone, I'm trying to get a simple (actually any) tensorflow/keras CNN model up and running, but without any luck. Training doesn't converge to anything. Output is either random noise or a constant value regardless of input.\n\nFrancois Chollet's [excellent data pipeline notebook](https://www.kaggle.com/code/fchollet/a-simple-high-performance-tf-data-pipeline) essentially gives us batches of data with the dimensions:\n\n`(BATCH_SIZE, X_DIM, Y_DIM, Z_DIM)` where X_DIM and Y_DIM are the size of the window, and Z_DIM gives us the height of the stack of images. \n\nI have tried Conv3d layers, e.g. with kernel size `(Z_DIM, Z_DIM, Z_DIM)`, which should return a tensor of size similar to `(BATCH_SIZE, X_DIM, Y_DIM, N_FILTERS)` and then dense layers with and without pooling beforehand.\n\nAlso I tried Conv2d layers (not sure what the resulting dimension ends up being here).\n\nI use sigmoid activation function for the final layer.\n\nA couple of questions about this architecture compared to the [example pytorch CNN model](https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial):\n\n1. The pytorch model looks like it has a single pixel as output, whereas I thought the model would/should be returning an image window sized equal to the input window?\n\n2. The loss function is binary cross entropy, I have used this as well as MSE, which is better?\n\n3. What is the difference between using conv2d and conv3d layers? Can we use conv2d layers? What about if Z_DIM=1 (i.e. we are predicting just based on a single tiff image from the stack)?",
    "2204895": "I tell my students to do everything in PyTorch. Google PyTorch/Keras/TF trend, PyTorch is dominating and increasing. The trend is your friend :).",
    "2217833": ""
  }
}