{
  "id": 396211,
  "title": "Different scoring time",
  "url": "/competitions/asl-signs/discussion/396211",
  "author_name": "",
  "post_date": "2023-03-20T18:29:24.418081500Z",
  "votes": 15,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Big part of our submissions looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F9adee9bd1973a697dd3e26368489bbb4%2F2023-03-20%20%2021.25.19.png?generation=1679336777439578&amp;alt=media\" alt=\"\"></p>\n<p>We submit the same code with different weights (and with the same weights). Check inference time in local and after commit, usually it's under 100ms. But usually submissions falls with error and sometimes everything is ok.</p>\n<p>Is there any way to make submissions stable? Or make sure that submissions will have success before submitting?</p>",
  "messages": [
    {
      "id": "2189764",
      "postDate": "03/20/2023 18:29:24",
      "content": "<p>Big part of our submissions looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F9adee9bd1973a697dd3e26368489bbb4%2F2023-03-20%20%2021.25.19.png?generation=1679336777439578&amp;alt=media\" alt=\"\"></p>\n<p>We submit the same code with different weights (and with the same weights). Check inference time in local and after commit, usually it's under 100ms. But usually submissions falls with error and sometimes everything is ok.</p>\n<p>Is there any way to make submissions stable? Or make sure that submissions will have success before submitting?</p>",
      "rawMarkdown": "Big part of our submissions looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F9adee9bd1973a697dd3e26368489bbb4%2F2023-03-20%20%2021.25.19.png?generation=1679336777439578&alt=media)\n\nWe submit the same code with different weights (and with the same weights). Check inference time in local and after commit, usually it's under 100ms. But usually submissions falls with error and sometimes everything is ok.\n\nIs there any way to make submissions stable? Or make sure that submissions will have success before submitting?",
      "votes": null
    },
    {
      "id": "2194103",
      "postDate": "03/23/2023 18:20:28",
      "content": "<p>Is that what's happening? An inference timeout is throwing a scoring error? I was wondering why some of my submissions were ending up like this. </p>",
      "rawMarkdown": "Is that what's happening? An inference timeout is throwing a scoring error? I was wondering why some of my submissions were ending up like this.",
      "votes": null
    },
    {
      "id": "2194468",
      "postDate": "03/24/2023 00:27:02",
      "content": "<p>This is the one of the reasons. I might be different size of model, different name, format. Or wrong call function in model etc. <br>\nBut in my case I tell about out of time error because same code works at some submits</p>",
      "rawMarkdown": "This is the one of the reasons. I might be different size of model, different name, format. Or wrong call function in model etc. \nBut in my case I tell about out of time error because same code works at some submits",
      "votes": null
    },
    {
      "id": "2195540",
      "postDate": "03/24/2023 17:55:26",
      "content": "<p>It only happens when I add more parameters. I wish the error messages would be more clear, but it seems like there's a built in timeout within the scoring method that's different from the 9 hour notebook timeout.</p>",
      "rawMarkdown": "It only happens when I add more parameters. I wish the error messages would be more clear, but it seems like there's a built in timeout within the scoring method that's different from the 9 hour notebook timeout.",
      "votes": null
    },
    {
      "id": "2197306",
      "postDate": "03/26/2023 04:23:06",
      "content": "<p>Till now, I have faced only storage related issues. Model with size around 39 MB starts throwing the error (assuming there should be some overheads making it to cross the given threshold).</p>\n<p>Interesting to note that same code is working at some submits. Could you please check if your size is close to 40 MB? May be the overheads are changing somehow. </p>",
      "rawMarkdown": "Till now, I have faced only storage related issues. Model with size around 39 MB starts throwing the error (assuming there should be some overheads making it to cross the given threshold).\n\nInteresting to note that same code is working at some submits. Could you please check if your size is close to 40 MB? May be the overheads are changing somehow.",
      "votes": null
    },
    {
      "id": "2198645",
      "postDate": "03/27/2023 07:48:09",
      "content": "<p>Hi, I'm facing the same issue. Didn't you manage to find out a reason so far?</p>",
      "rawMarkdown": "Hi, I'm facing the same issue. Didn't you manage to find out a reason so far?",
      "votes": null
    },
    {
      "id": "2198654",
      "postDate": "03/27/2023 07:56:03",
      "content": "<p>No, it's still 1-2 succeeded from 10 submissions<br>\nTried fusing two models to one, different onnx conversions, raw keras models. And it's the same</p>\n<p>btw ONNX 5-7 times faster than TFLite for me, but we unable to use it </p>",
      "rawMarkdown": "No, it's still 1-2 succeeded from 10 submissions\nTried fusing two models to one, different onnx conversions, raw keras models. And it's the same\n\nbtw ONNX 5-7 times faster than TFLite for me, but we unable to use it",
      "votes": null
    },
    {
      "id": "2198709",
      "postDate": "03/27/2023 09:02:49",
      "content": "<p>Yeah, I just pressed on a submission button of the same notebook twice and got different results</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2F55fa57a566db3c8bda4f7edf2f3d8645%2Fphoto_2023-03-27_12-01-19.jpg?generation=1679907702543463&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Yeah, I just pressed on a submission button of the same notebook twice and got different results\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2F55fa57a566db3c8bda4f7edf2f3d8645%2Fphoto_2023-03-27_12-01-19.jpg?generation=1679907702543463&alt=media)",
      "votes": null
    },
    {
      "id": "2198842",
      "postDate": "03/27/2023 10:26:51",
      "content": "<p>Unfortunately, Kaggle servers have been known to be inconsistent regarding runtimes. It has nothing to do with code, it's because of the hardware (afaik)</p>\n<p>In \"classical\" code competitions the same code could take 8 hours +/- 1 hour to run. This is usually fine because 9 hours is enough to fit a whole ensemble. </p>\n<p>However in our case here that's an issue because we're trying to respect the 100ms per iteration limit, which makes the overall submission process really unpleasant.</p>",
      "rawMarkdown": "Unfortunately, Kaggle servers have been known to be inconsistent regarding runtimes. It has nothing to do with code, it's because of the hardware (afaik)\n\nIn \"classical\" code competitions the same code could take 8 hours +/- 1 hour to run. This is usually fine because 9 hours is enough to fit a whole ensemble. \n\nHowever in our case here that's an issue because we're trying to respect the 100ms per iteration limit, which makes the overall submission process really unpleasant.",
      "votes": null
    },
    {
      "id": "2199082",
      "postDate": "03/27/2023 13:46:46",
      "content": "<p>Unfortunately.</p>\n<p>Maybe in this type of competition would be honest to run each submission few times. Ordinal 9 hours ~ 540 // 75minutes. And take mean or min inference time.</p>\n<p>For resources economy: break if one of runs succeed</p>",
      "rawMarkdown": "Unfortunately.\n\nMaybe in this type of competition would be honest to run each submission few times. Ordinal 9 hours ~ 540 // 75minutes. And take mean or min inference time.\n\nFor resources economy: break if one of runs succeed",
      "votes": null
    },
    {
      "id": "2199373",
      "postDate": "03/27/2023 17:12:39",
      "content": "<p>Even running the submitted model several times would reduce the number of failures on the edge cases (like 99ms vs 101ms) but would not fix it at all. After all, the constraint of competition is to evaluate our models under 100ms but there is no hardware named to be used, no standard deviation mentioned, and no background payload is specified. In the end, we do not know the exact distribution of the length of sequences in the hidden set. We do not know any distribution of the hidden layer set (in case you are using dynamic NN). Having all that put al us under unknown constraints. If your model is running out of time, you probably should to reduce that size in order to get stable submissions. Otherwise, you may try your luck submitting big enough models dozens times and hope your best model will be accepted. </p>\n<p>To sum up my thoughts - there are no strict hardware/software requirements specified, thus there is no good reason o expect a stable result under unstable constrains. </p>",
      "rawMarkdown": "Even running the submitted model several times would reduce the number of failures on the edge cases (like 99ms vs 101ms) but would not fix it at all. After all, the constraint of competition is to evaluate our models under 100ms but there is no hardware named to be used, no standard deviation mentioned, and no background payload is specified. In the end, we do not know the exact distribution of the length of sequences in the hidden set. We do not know any distribution of the hidden layer set (in case you are using dynamic NN). Having all that put al us under unknown constraints. If your model is running out of time, you probably should to reduce that size in order to get stable submissions. Otherwise, you may try your luck submitting big enough models dozens times and hope your best model will be accepted. \n\nTo sum up my thoughts - there are no strict hardware/software requirements specified, thus there is no good reason o expect a stable result under unstable constrains.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2194103,
      "author_name": "abdoue",
      "author_url": "",
      "post_date": "03/23/2023 18:20:28",
      "content": "<p>Is that what's happening? An inference timeout is throwing a scoring error? I was wondering why some of my submissions were ending up like this. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2194468,
          "author_name": "kolyaforrat",
          "author_url": "",
          "post_date": "03/24/2023 00:27:02",
          "content": "<p>This is the one of the reasons. I might be different size of model, different name, format. Or wrong call function in model etc. <br>\nBut in my case I tell about out of time error because same code works at some submits</p>",
          "votes": null,
          "replies": [
            {
              "id": 2195540,
              "author_name": "abdoue",
              "author_url": "",
              "post_date": "03/24/2023 17:55:26",
              "content": "<p>It only happens when I add more parameters. I wish the error messages would be more clear, but it seems like there's a built in timeout within the scoring method that's different from the 9 hour notebook timeout.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2197306,
      "author_name": "mannyjain",
      "author_url": "",
      "post_date": "03/26/2023 04:23:06",
      "content": "<p>Till now, I have faced only storage related issues. Model with size around 39 MB starts throwing the error (assuming there should be some overheads making it to cross the given threshold).</p>\n<p>Interesting to note that same code is working at some submits. Could you please check if your size is close to 40 MB? May be the overheads are changing somehow. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2198645,
      "author_name": "vadimtimakin",
      "author_url": "",
      "post_date": "03/27/2023 07:48:09",
      "content": "<p>Hi, I'm facing the same issue. Didn't you manage to find out a reason so far?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2198654,
          "author_name": "kolyaforrat",
          "author_url": "",
          "post_date": "03/27/2023 07:56:03",
          "content": "<p>No, it's still 1-2 succeeded from 10 submissions<br>\nTried fusing two models to one, different onnx conversions, raw keras models. And it's the same</p>\n<p>btw ONNX 5-7 times faster than TFLite for me, but we unable to use it </p>",
          "votes": null,
          "replies": [
            {
              "id": 2198709,
              "author_name": "vadimtimakin",
              "author_url": "",
              "post_date": "03/27/2023 09:02:49",
              "content": "<p>Yeah, I just pressed on a submission button of the same notebook twice and got different results</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2F55fa57a566db3c8bda4f7edf2f3d8645%2Fphoto_2023-03-27_12-01-19.jpg?generation=1679907702543463&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2198842,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "03/27/2023 10:26:51",
      "content": "<p>Unfortunately, Kaggle servers have been known to be inconsistent regarding runtimes. It has nothing to do with code, it's because of the hardware (afaik)</p>\n<p>In \"classical\" code competitions the same code could take 8 hours +/- 1 hour to run. This is usually fine because 9 hours is enough to fit a whole ensemble. </p>\n<p>However in our case here that's an issue because we're trying to respect the 100ms per iteration limit, which makes the overall submission process really unpleasant.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2199082,
          "author_name": "kolyaforrat",
          "author_url": "",
          "post_date": "03/27/2023 13:46:46",
          "content": "<p>Unfortunately.</p>\n<p>Maybe in this type of competition would be honest to run each submission few times. Ordinal 9 hours ~ 540 // 75minutes. And take mean or min inference time.</p>\n<p>For resources economy: break if one of runs succeed</p>",
          "votes": null,
          "replies": [
            {
              "id": 2199373,
              "author_name": "meowmeowmeowmeowmeow",
              "author_url": "",
              "post_date": "03/27/2023 17:12:39",
              "content": "<p>Even running the submitted model several times would reduce the number of failures on the edge cases (like 99ms vs 101ms) but would not fix it at all. After all, the constraint of competition is to evaluate our models under 100ms but there is no hardware named to be used, no standard deviation mentioned, and no background payload is specified. In the end, we do not know the exact distribution of the length of sequences in the hidden set. We do not know any distribution of the hidden layer set (in case you are using dynamic NN). Having all that put al us under unknown constraints. If your model is running out of time, you probably should to reduce that size in order to get stable submissions. Otherwise, you may try your luck submitting big enough models dozens times and hope your best model will be accepted. </p>\n<p>To sum up my thoughts - there are no strict hardware/software requirements specified, thus there is no good reason o expect a stable result under unstable constrains. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2189764": "Big part of our submissions looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4212496%2F9adee9bd1973a697dd3e26368489bbb4%2F2023-03-20%20%2021.25.19.png?generation=1679336777439578&alt=media)\n\nWe submit the same code with different weights (and with the same weights). Check inference time in local and after commit, usually it's under 100ms. But usually submissions falls with error and sometimes everything is ok.\n\nIs there any way to make submissions stable? Or make sure that submissions will have success before submitting?",
    "2194103": "Is that what's happening? An inference timeout is throwing a scoring error? I was wondering why some of my submissions were ending up like this.",
    "2194468": "This is the one of the reasons. I might be different size of model, different name, format. Or wrong call function in model etc. \nBut in my case I tell about out of time error because same code works at some submits",
    "2195540": "It only happens when I add more parameters. I wish the error messages would be more clear, but it seems like there's a built in timeout within the scoring method that's different from the 9 hour notebook timeout.",
    "2197306": "Till now, I have faced only storage related issues. Model with size around 39 MB starts throwing the error (assuming there should be some overheads making it to cross the given threshold).\n\nInteresting to note that same code is working at some submits. Could you please check if your size is close to 40 MB? May be the overheads are changing somehow.",
    "2198645": "Hi, I'm facing the same issue. Didn't you manage to find out a reason so far?",
    "2198654": "No, it's still 1-2 succeeded from 10 submissions\nTried fusing two models to one, different onnx conversions, raw keras models. And it's the same\n\nbtw ONNX 5-7 times faster than TFLite for me, but we unable to use it",
    "2198709": "Yeah, I just pressed on a submission button of the same notebook twice and got different results\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5007869%2F55fa57a566db3c8bda4f7edf2f3d8645%2Fphoto_2023-03-27_12-01-19.jpg?generation=1679907702543463&alt=media)",
    "2198842": "Unfortunately, Kaggle servers have been known to be inconsistent regarding runtimes. It has nothing to do with code, it's because of the hardware (afaik)\n\nIn \"classical\" code competitions the same code could take 8 hours +/- 1 hour to run. This is usually fine because 9 hours is enough to fit a whole ensemble. \n\nHowever in our case here that's an issue because we're trying to respect the 100ms per iteration limit, which makes the overall submission process really unpleasant.",
    "2199082": "Unfortunately.\n\nMaybe in this type of competition would be honest to run each submission few times. Ordinal 9 hours ~ 540 // 75minutes. And take mean or min inference time.\n\nFor resources economy: break if one of runs succeed",
    "2199373": "Even running the submitted model several times would reduce the number of failures on the edge cases (like 99ms vs 101ms) but would not fix it at all. After all, the constraint of competition is to evaluate our models under 100ms but there is no hardware named to be used, no standard deviation mentioned, and no background payload is specified. In the end, we do not know the exact distribution of the length of sequences in the hidden set. We do not know any distribution of the hidden layer set (in case you are using dynamic NN). Having all that put al us under unknown constraints. If your model is running out of time, you probably should to reduce that size in order to get stable submissions. Otherwise, you may try your luck submitting big enough models dozens times and hope your best model will be accepted. \n\nTo sum up my thoughts - there are no strict hardware/software requirements specified, thus there is no good reason o expect a stable result under unstable constrains."
  },
  "source": "meta"
}