{
  "id": 107876,
  "title": "Kernel Threw Exception",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107876",
  "author_name": "",
  "post_date": "2019-09-07T15:18:07.882568600Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>We have only been able to submit one version successfully as we keep getting an error </p>\n\n<p><code>Kernel Threw Exception</code></p>\n\n<p>Looking at the log, it is not clear at all why the version is failing, i.e. </p>\n\n<p>```\nPredicting</p>\n\n<p>2019-09-07 07:41:52.231132: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0</p>\n\n<p>2019-09-07 07:41:52.940301: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7</p>\n\n<p>y_pred Shape:  (1928, 5)</p>\n\n<p>[[0.009 0.296 0.398 0.111 0.186]\n [0.015 0.156 0.273 0.276 0.281]\n [0.009 0.293 0.042 0.356 0.3  ]\n ...\n [0.021 0.148 0.283 0.34  0.208]\n [0.007 0.197 0.516 0.136 0.144]\n [0.995 0.002 0.001 0.    0.001]]</p>\n\n<p>(1928, 2)</p>\n\n<p>```</p>\n\n<p>It seems to complete the test data without error.  Any ideas why we are getting the failure?</p>",
  "messages": [
    {
      "id": "620499",
      "postDate": "09/07/2019 15:18:07",
      "content": "<p>We have only been able to submit one version successfully as we keep getting an error </p>\n\n<p><code>Kernel Threw Exception</code></p>\n\n<p>Looking at the log, it is not clear at all why the version is failing, i.e. </p>\n\n<p>```\nPredicting</p>\n\n<p>2019-09-07 07:41:52.231132: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0</p>\n\n<p>2019-09-07 07:41:52.940301: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7</p>\n\n<p>y_pred Shape:  (1928, 5)</p>\n\n<p>[[0.009 0.296 0.398 0.111 0.186]\n [0.015 0.156 0.273 0.276 0.281]\n [0.009 0.293 0.042 0.356 0.3  ]\n ...\n [0.021 0.148 0.283 0.34  0.208]\n [0.007 0.197 0.516 0.136 0.144]\n [0.995 0.002 0.001 0.    0.001]]</p>\n\n<p>(1928, 2)</p>\n\n<p>```</p>\n\n<p>It seems to complete the test data without error.  Any ideas why we are getting the failure?</p>",
      "rawMarkdown": "We have only been able to submit one version successfully as we keep getting an error \n\n`Kernel Threw Exception`\n\nLooking at the log, it is not clear at all why the version is failing, i.e. \n\n```\nPredicting\n\n2019-09-07 07:41:52.231132: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0\n\n2019-09-07 07:41:52.940301: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7\n\ny_pred Shape:  (1928, 5)\n\n[[0.009 0.296 0.398 0.111 0.186]\n [0.015 0.156 0.273 0.276 0.281]\n [0.009 0.293 0.042 0.356 0.3  ]\n ...\n [0.021 0.148 0.283 0.34  0.208]\n [0.007 0.197 0.516 0.136 0.144]\n [0.995 0.002 0.001 0.    0.001]]\n\n(1928, 2)\n\n\n```\n\nIt seems to complete the test data without error.  Any ideas why we are getting the failure?",
      "votes": null
    },
    {
      "id": "620533",
      "postDate": "09/07/2019 16:14:20",
      "content": "<p>Looks like when the kernel is running with whole 13000 images, the kernel is throwing up this error. I had faced similar issue. It was able to predict for entire 1928 images in less than 1 hour but when the kernel was submitted for the entire test and scoring data, the kernel was timed out. </p>",
      "rawMarkdown": "Looks like when the kernel is running with whole 13000 images, the kernel is throwing up this error. I had faced similar issue. It was able to predict for entire 1928 images in less than 1 hour but when the kernel was submitted for the entire test and scoring data, the kernel was timed out.",
      "votes": null
    },
    {
      "id": "620548",
      "postDate": "09/07/2019 16:52:33",
      "content": "<p>So basically don't do them all together.  Is it best to work with chunks e.g. 1000 at a time or singly? </p>",
      "rawMarkdown": "So basically don't do them all together.  Is it best to work with chunks e.g. 1000 at a time or singly?",
      "votes": null
    },
    {
      "id": "620572",
      "postDate": "09/07/2019 17:25:46",
      "content": "<p>No <a href=\"/seanjmcm\">@seanjmcm</a> . These are Kernels only competition. When you submit the kernel, it automatically runs for the all images. Only 15% of data results are shared to us. Rest will be shared once the competition gets over. Only advice is to keep the kernel running time to as much as low possible. </p>",
      "rawMarkdown": "No @seanjmcm . These are Kernels only competition. When you submit the kernel, it automatically runs for the all images. Only 15% of data results are shared to us. Rest will be shared once the competition gets over. Only advice is to keep the kernel running time to as much as low possible.",
      "votes": null
    },
    {
      "id": "620640",
      "postDate": "09/07/2019 19:06:02",
      "content": "<p>'Only advice is to keep the kernel running time to as much as low possible.' </p>\n\n<p>Ok, I'm assuming less training epochs = faster kernel? </p>",
      "rawMarkdown": "'Only advice is to keep the kernel running time to as much as low possible.' \n\nOk, I'm assuming less training epochs = faster kernel?",
      "votes": null
    },
    {
      "id": "620669",
      "postDate": "09/07/2019 20:09:28",
      "content": "<p>I think that error could be received in several ways 😁. There must be an error while your kernel runs on private test set.\nFor example, once I received that because I was renaming image ID including '.png' while you need to maintain the original name.\nAnother case was because of a bug in one of the first cropping functions shared.</p>",
      "rawMarkdown": "I think that error could be received in several ways 😁. There must be an error while your kernel runs on private test set.\nFor example, once I received that because I was renaming image ID including '.png' while you need to maintain the original name.\nAnother case was because of a bug in one of the first cropping functions shared.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 620533,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "09/07/2019 16:14:20",
      "content": "<p>Looks like when the kernel is running with whole 13000 images, the kernel is throwing up this error. I had faced similar issue. It was able to predict for entire 1928 images in less than 1 hour but when the kernel was submitted for the entire test and scoring data, the kernel was timed out. </p>",
      "votes": null,
      "replies": [
        {
          "id": 620548,
          "author_name": "seanjmcm",
          "author_url": "",
          "post_date": "09/07/2019 16:52:33",
          "content": "<p>So basically don't do them all together.  Is it best to work with chunks e.g. 1000 at a time or singly? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620572,
          "author_name": "manojprabhaakr",
          "author_url": "",
          "post_date": "09/07/2019 17:25:46",
          "content": "<p>No <a href=\"/seanjmcm\">@seanjmcm</a> . These are Kernels only competition. When you submit the kernel, it automatically runs for the all images. Only 15% of data results are shared to us. Rest will be shared once the competition gets over. Only advice is to keep the kernel running time to as much as low possible. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620640,
          "author_name": "seanjmcm",
          "author_url": "",
          "post_date": "09/07/2019 19:06:02",
          "content": "<p>'Only advice is to keep the kernel running time to as much as low possible.' </p>\n\n<p>Ok, I'm assuming less training epochs = faster kernel? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 620669,
      "author_name": "raimonds1993",
      "author_url": "",
      "post_date": "09/07/2019 20:09:28",
      "content": "<p>I think that error could be received in several ways 😁. There must be an error while your kernel runs on private test set.\nFor example, once I received that because I was renaming image ID including '.png' while you need to maintain the original name.\nAnother case was because of a bug in one of the first cropping functions shared.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "620499": "We have only been able to submit one version successfully as we keep getting an error \n\n`Kernel Threw Exception`\n\nLooking at the log, it is not clear at all why the version is failing, i.e. \n\n```\nPredicting\n\n2019-09-07 07:41:52.231132: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0\n\n2019-09-07 07:41:52.940301: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7\n\ny_pred Shape:  (1928, 5)\n\n[[0.009 0.296 0.398 0.111 0.186]\n [0.015 0.156 0.273 0.276 0.281]\n [0.009 0.293 0.042 0.356 0.3  ]\n ...\n [0.021 0.148 0.283 0.34  0.208]\n [0.007 0.197 0.516 0.136 0.144]\n [0.995 0.002 0.001 0.    0.001]]\n\n(1928, 2)\n\n\n```\n\nIt seems to complete the test data without error.  Any ideas why we are getting the failure?",
    "620533": "Looks like when the kernel is running with whole 13000 images, the kernel is throwing up this error. I had faced similar issue. It was able to predict for entire 1928 images in less than 1 hour but when the kernel was submitted for the entire test and scoring data, the kernel was timed out.",
    "620548": "So basically don't do them all together.  Is it best to work with chunks e.g. 1000 at a time or singly?",
    "620572": "No @seanjmcm . These are Kernels only competition. When you submit the kernel, it automatically runs for the all images. Only 15% of data results are shared to us. Rest will be shared once the competition gets over. Only advice is to keep the kernel running time to as much as low possible.",
    "620640": "'Only advice is to keep the kernel running time to as much as low possible.' \n\nOk, I'm assuming less training epochs = faster kernel?",
    "620669": "I think that error could be received in several ways 😁. There must be an error while your kernel runs on private test set.\nFor example, once I received that because I was renaming image ID including '.png' while you need to maintain the original name.\nAnother case was because of a bug in one of the first cropping functions shared."
  },
  "source": "meta"
}