{
  "id": 74527,
  "title": "New Notebook docker image is broken",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74527",
  "author_name": "",
  "post_date": "2018-12-13T06:10:30.931471300Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>With the Notebook docker image, my existing notebook stops working. I got a exception when running CuDNNLSTM (OpKernel is not available) even the GPU option was On. K.batch_dot starts complaining about mismatch shape or tensor conversion issue,  which worked well before (I have not changed anything).</p>\n\n<p>Does anyone experience the same issues?</p>\n\n<p>Regards,\nDuong nhu</p>",
  "messages": [
    {
      "id": "438120",
      "postDate": "12/13/2018 06:10:30",
      "content": "<p>Hi,</p>\n\n<p>With the Notebook docker image, my existing notebook stops working. I got a exception when running CuDNNLSTM (OpKernel is not available) even the GPU option was On. K.batch_dot starts complaining about mismatch shape or tensor conversion issue,  which worked well before (I have not changed anything).</p>\n\n<p>Does anyone experience the same issues?</p>\n\n<p>Regards,\nDuong nhu</p>",
      "rawMarkdown": "Hi,\n\nWith the Notebook docker image, my existing notebook stops working. I got a exception when running CuDNNLSTM (OpKernel is not available) even the GPU option was On. K.batch_dot starts complaining about mismatch shape or tensor conversion issue,  which worked well before (I have not changed anything).\n\nDoes anyone experience the same issues?\n\nRegards,\nDuong nhu",
      "votes": null
    },
    {
      "id": "438123",
      "postDate": "12/13/2018 06:16:14",
      "content": "<p>@Duong - I have also experienced the same thing except with capsule layers. I have not changed anything from kernel implementations.</p>",
      "rawMarkdown": "Duong - I have also experienced the same thing except with capsule layers. I have not changed anything from kernel implementations.",
      "votes": null
    },
    {
      "id": "438139",
      "postDate": "12/13/2018 06:55:40",
      "content": "<p>Yeah!there are errors in the capsule layers now!</p>",
      "rawMarkdown": "Yeah!there are errors in the capsule layers now!",
      "votes": null
    },
    {
      "id": "438154",
      "postDate": "12/13/2018 07:34:00",
      "content": "<p>I cant even run any cuda layer now</p>",
      "rawMarkdown": "I cant even run any cuda layer now",
      "votes": null
    },
    {
      "id": "438235",
      "postDate": "12/13/2018 10:12:36",
      "content": "<p>Cuda layers work for me, capsule doesn't.</p>",
      "rawMarkdown": "Cuda layers work for me, capsule doesn't.",
      "votes": null
    },
    {
      "id": "438267",
      "postDate": "12/13/2018 11:47:18",
      "content": "<p>still unable to run  Net with capsule layer using kaggle  notebook, locally everythings run perfectly </p>",
      "rawMarkdown": "still unable to run  Net with capsule layer using kaggle  notebook, locally everythings run perfectly",
      "votes": null
    },
    {
      "id": "438287",
      "postDate": "12/13/2018 12:20:05",
      "content": "<p>(Temporary) solution is to replace <code>K.batch_dot</code> with <code>tf.keras.backend.batch_dot</code></p>",
      "rawMarkdown": "(Temporary) solution is to replace `K.batch_dot` with `tf.keras.backend.batch_dot`",
      "votes": null
    },
    {
      "id": "438300",
      "postDate": "12/13/2018 12:51:05",
      "content": "<p>Tested and run smoothly, thanks for your help</p>",
      "rawMarkdown": "Tested and run smoothly, thanks for your help",
      "votes": null
    },
    {
      "id": "438493",
      "postDate": "12/13/2018 19:43:47",
      "content": "<p>I am also getting the shape mismatch from CuDNN (using PyTorch)</p>\n\n<p>Looking at the docker image it looks like PyTorch was updated, but that wouldn't explain the issues people are seeing in Keras.  I suspect CUDA has been updated also, but I don't see it in the latest commit to the docker image repository at <a href=\"https://github.com/Kaggle/docker-python\">https://github.com/Kaggle/docker-python</a></p>\n\n<p>[Update] - latest notebook image has CUDA 9.0.176 and CuDNN version 7.4.01, which is later than the versions a few weeks ago (but I forget exactly what they were)</p>",
      "rawMarkdown": "I am also getting the shape mismatch from CuDNN (using PyTorch)\n\nLooking at the docker image it looks like PyTorch was updated, but that wouldn't explain the issues people are seeing in Keras.  I suspect CUDA has been updated also, but I don't see it in the latest commit to the docker image repository at https://github.com/Kaggle/docker-python\n\n[Update] - latest notebook image has CUDA 9.0.176 and CuDNN version 7.4.01, which is later than the versions a few weeks ago (but I forget exactly what they were)",
      "votes": null
    },
    {
      "id": "438544",
      "postDate": "12/13/2018 21:35:27",
      "content": "<p>Old images seem not to work with Cuda. Latest image works. K.batch_dot still complains about mismatched shape. </p>\n\n<p>Temp solution: </p>\n\n<blockquote>\n  <p><strong>Philipp wrote</strong></p>\n  \n  <blockquote>\n    <p>(Temporary) solution is to replace <code>K.batch_dot</code> with <code>tf.keras.backend.batch_dot</code></p>\n  </blockquote>\n</blockquote>",
      "rawMarkdown": "Old images seem not to work with Cuda. Latest image works. K.batch_dot still complains about mismatched shape. \n\nTemp solution: \n&gt; **Philipp wrote**\n&gt; \n&gt; &gt; (Temporary) solution is to replace `K.batch_dot` with `tf.keras.backend.batch_dot`",
      "votes": null
    },
    {
      "id": "438991",
      "postDate": "12/14/2018 14:25:31",
      "content": "<p>Further update - it's more specific than LSTMs generally.  It's actually an incompatibility between PyTorch 1.0 and some code for Drop-connect that comes from originally from a Merity paper (Salesforce) at <a href=\"https://github.com/salesforce/awd-lstm-lm\">https://github.com/salesforce/awd-lstm-lm</a>\nI'm trying to find a work-around currently and will report back if I am able to find one.</p>",
      "rawMarkdown": "Further update - it's more specific than LSTMs generally.  It's actually an incompatibility between PyTorch 1.0 and some code for Drop-connect that comes from originally from a Merity paper (Salesforce) at https://github.com/salesforce/awd-lstm-lm\nI'm trying to find a work-around currently and will report back if I am able to find one.",
      "votes": null
    },
    {
      "id": "440514",
      "postDate": "12/17/2018 16:41:56",
      "content": "<p>I hacked WeightDrop around in such a way as to make it work (I <strong>think</strong> correctly) with PyTorch 1.0.  See my last post on <a href=\"https://github.com/salesforce/awd-lstm-lm/issues/86\">https://github.com/salesforce/awd-lstm-lm/issues/86</a></p>",
      "rawMarkdown": "I hacked WeightDrop around in such a way as to make it work (I **think** correctly) with PyTorch 1.0.  See my last post on https://github.com/salesforce/awd-lstm-lm/issues/86",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 438123,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "12/13/2018 06:16:14",
      "content": "<p>@Duong - I have also experienced the same thing except with capsule layers. I have not changed anything from kernel implementations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 438139,
          "author_name": "gmhost",
          "author_url": "",
          "post_date": "12/13/2018 06:55:40",
          "content": "<p>Yeah!there are errors in the capsule layers now!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438154,
          "author_name": "",
          "author_url": "",
          "post_date": "12/13/2018 07:34:00",
          "content": "<p>I cant even run any cuda layer now</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438235,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/13/2018 10:12:36",
          "content": "<p>Cuda layers work for me, capsule doesn't.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 438267,
      "author_name": "malekbadreddine",
      "author_url": "",
      "post_date": "12/13/2018 11:47:18",
      "content": "<p>still unable to run  Net with capsule layer using kaggle  notebook, locally everythings run perfectly </p>",
      "votes": null,
      "replies": [
        {
          "id": 438287,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/13/2018 12:20:05",
          "content": "<p>(Temporary) solution is to replace <code>K.batch_dot</code> with <code>tf.keras.backend.batch_dot</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438300,
          "author_name": "malekbadreddine",
          "author_url": "",
          "post_date": "12/13/2018 12:51:05",
          "content": "<p>Tested and run smoothly, thanks for your help</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 438493,
      "author_name": "stevedraper",
      "author_url": "",
      "post_date": "12/13/2018 19:43:47",
      "content": "<p>I am also getting the shape mismatch from CuDNN (using PyTorch)</p>\n\n<p>Looking at the docker image it looks like PyTorch was updated, but that wouldn't explain the issues people are seeing in Keras.  I suspect CUDA has been updated also, but I don't see it in the latest commit to the docker image repository at <a href=\"https://github.com/Kaggle/docker-python\">https://github.com/Kaggle/docker-python</a></p>\n\n<p>[Update] - latest notebook image has CUDA 9.0.176 and CuDNN version 7.4.01, which is later than the versions a few weeks ago (but I forget exactly what they were)</p>",
      "votes": null,
      "replies": [
        {
          "id": 438991,
          "author_name": "stevedraper",
          "author_url": "",
          "post_date": "12/14/2018 14:25:31",
          "content": "<p>Further update - it's more specific than LSTMs generally.  It's actually an incompatibility between PyTorch 1.0 and some code for Drop-connect that comes from originally from a Merity paper (Salesforce) at <a href=\"https://github.com/salesforce/awd-lstm-lm\">https://github.com/salesforce/awd-lstm-lm</a>\nI'm trying to find a work-around currently and will report back if I am able to find one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440514,
          "author_name": "stevedraper",
          "author_url": "",
          "post_date": "12/17/2018 16:41:56",
          "content": "<p>I hacked WeightDrop around in such a way as to make it work (I <strong>think</strong> correctly) with PyTorch 1.0.  See my last post on <a href=\"https://github.com/salesforce/awd-lstm-lm/issues/86\">https://github.com/salesforce/awd-lstm-lm/issues/86</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 438544,
      "author_name": "",
      "author_url": "",
      "post_date": "12/13/2018 21:35:27",
      "content": "<p>Old images seem not to work with Cuda. Latest image works. K.batch_dot still complains about mismatched shape. </p>\n\n<p>Temp solution: </p>\n\n<blockquote>\n  <p><strong>Philipp wrote</strong></p>\n  \n  <blockquote>\n    <p>(Temporary) solution is to replace <code>K.batch_dot</code> with <code>tf.keras.backend.batch_dot</code></p>\n  </blockquote>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "438120": "Hi,\n\nWith the Notebook docker image, my existing notebook stops working. I got a exception when running CuDNNLSTM (OpKernel is not available) even the GPU option was On. K.batch_dot starts complaining about mismatch shape or tensor conversion issue,  which worked well before (I have not changed anything).\n\nDoes anyone experience the same issues?\n\nRegards,\nDuong nhu",
    "438123": "Duong - I have also experienced the same thing except with capsule layers. I have not changed anything from kernel implementations.",
    "438139": "Yeah!there are errors in the capsule layers now!",
    "438154": "I cant even run any cuda layer now",
    "438235": "Cuda layers work for me, capsule doesn't.",
    "438267": "still unable to run  Net with capsule layer using kaggle  notebook, locally everythings run perfectly",
    "438287": "(Temporary) solution is to replace `K.batch_dot` with `tf.keras.backend.batch_dot`",
    "438300": "Tested and run smoothly, thanks for your help",
    "438493": "I am also getting the shape mismatch from CuDNN (using PyTorch)\n\nLooking at the docker image it looks like PyTorch was updated, but that wouldn't explain the issues people are seeing in Keras.  I suspect CUDA has been updated also, but I don't see it in the latest commit to the docker image repository at https://github.com/Kaggle/docker-python\n\n[Update] - latest notebook image has CUDA 9.0.176 and CuDNN version 7.4.01, which is later than the versions a few weeks ago (but I forget exactly what they were)",
    "438544": "Old images seem not to work with Cuda. Latest image works. K.batch_dot still complains about mismatched shape. \n\nTemp solution: \n&gt; **Philipp wrote**\n&gt; \n&gt; &gt; (Temporary) solution is to replace `K.batch_dot` with `tf.keras.backend.batch_dot`",
    "438991": "Further update - it's more specific than LSTMs generally.  It's actually an incompatibility between PyTorch 1.0 and some code for Drop-connect that comes from originally from a Merity paper (Salesforce) at https://github.com/salesforce/awd-lstm-lm\nI'm trying to find a work-around currently and will report back if I am able to find one.",
    "440514": "I hacked WeightDrop around in such a way as to make it work (I **think** correctly) with PyTorch 1.0.  See my last post on https://github.com/salesforce/awd-lstm-lm/issues/86"
  },
  "source": "meta"
}