{
  "id": 170773,
  "title": "Exceeded Allowed Compute Error？model parameter size limit?GPU memory limit?",
  "url": "/competitions/landmark-retrieval-2020/discussion/170773",
  "author_name": "smileimg",
  "post_date": "2020-07-29T03:17:03.537000",
  "votes": 5,
  "comment_count": 17,
  "views": 0,
  "content": "<p>As we all know, fusing multi-model output in feature retrieval competition is effective for improving map in Google Landmark Retrieval.So in order to test whether it is feasible, we simply copied multiple resnet101 networks and combine them into a single model output in tensorflow, which can forward normally locally.But we get Notebook Exceeded Allowed Compute Error in submissions. And we combine 2 resnet101 model,the submission result is normal. So where did we exceed the limit?Memory ? model parameter's size? or others ?</p>\n\n<p>Detail:\n1.copied and combine 6 resnet101, and embedding the result to 4096 global descriptor;\n2.model size: 887.79MB\n3.After running for about 3-4 hours, an error was reported. Less than 9 hours timeout\n4. combine 2 resnet101, embedding result to 1024, submission result is normal </p>\n\n<p>So we would like to consult, are there any restrictions on the parameter size, GPU memory, embedding size, etc. of the submitted model?</p>",
  "messages": [
    {
      "id": 950012,
      "postDate": "2020-07-29T05:56:16.210Z",
      "content": "<p>Maybe try submitting a single model with 4096 global descriptor to make sure it is not a memory error?</p>\n\n<p>Btw you say \"we\" but you seem to be a solo competitor? Who is \"we\"?</p>",
      "rawMarkdown": "Maybe try submitting a single model with 4096 global descriptor to make sure it is not a memory error?\n\nBtw you say \"we\" but you seem to be a solo competitor? Who is \"we\"?",
      "votes": 14,
      "replies": [
        {
          "id": 950109,
          "postDate": "2020-07-29T07:30:48.020Z",
          "rawMarkdown": "",
          "votes": -4,
          "isDeleted": true
        },
        {
          "id": 950121,
          "postDate": "2020-07-29T07:46:08.843Z",
          "content": "<p>As long as you don't share any information with each other until you merge teams, it must be according to Kaggle rules. But I would advise you to be careful because you guys have great performance, it would be very bad if you get disqualified because of private sharing.</p>",
          "rawMarkdown": "As long as you don't share any information with each other until you merge teams, it must be according to Kaggle rules. But I would advise you to be careful because you guys have great performance, it would be very bad if you get disqualified because of private sharing.",
          "votes": 1
        },
        {
          "id": 950161,
          "postDate": "2020-07-29T08:09:27.400Z",
          "content": "<p>Thanks for your advice,Let us refocus on this issue instead of information about us.</p>",
          "rawMarkdown": "Thanks for your advice,Let us refocus on this issue instead of information about us.",
          "votes": -4
        },
        {
          "id": 950256,
          "postDate": "2020-07-29T09:52:07.850Z",
          "content": "<p><a href=\"/renhui111111\">@renhui111111</a> You should have a read at this discussion : <a href=\"https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294\">https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294</a> </p>\n\n<p>This \"team\" was removed after the end of the competition, for they shared information before merging ...</p>",
          "rawMarkdown": "@renhui111111 You should have a read at this discussion : https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294 \n\nThis \"team\" was removed after the end of the competition, for they shared information before merging ...",
          "votes": 4
        },
        {
          "id": 950395,
          "postDate": "2020-07-29T11:21:13.577Z",
          "content": "<p>Maybe try submitting a single model with 4096 global descriptor to make sure it is not a memory error?We try submmit it, and the result is normal</p>",
          "rawMarkdown": "Maybe try submitting a single model with 4096 global descriptor to make sure it is not a memory error?We try submmit it, and the result is normal"
        },
        {
          "id": 950423,
          "postDate": "2020-07-29T11:40:42.427Z",
          "content": "<p>Okay, then at least it narrows it down. So the issue is not the matching part of the pipeline.</p>",
          "rawMarkdown": "Okay, then at least it narrows it down. So the issue is not the matching part of the pipeline."
        }
      ]
    },
    {
      "id": 949880,
      "postDate": "2020-07-29T03:17:03.537Z",
      "content": "<p>As we all know, fusing multi-model output in feature retrieval competition is effective for improving map in Google Landmark Retrieval.So in order to test whether it is feasible, we simply copied multiple resnet101 networks and combine them into a single model output in tensorflow, which can forward normally locally.But we get Notebook Exceeded Allowed Compute Error in submissions. And we combine 2 resnet101 model,the submission result is normal. So where did we exceed the limit?Memory ? model parameter's size? or others ?</p>\n\n<p>Detail:\n1.copied and combine 6 resnet101, and embedding the result to 4096 global descriptor;\n2.model size: 887.79MB\n3.After running for about 3-4 hours, an error was reported. Less than 9 hours timeout\n4. combine 2 resnet101, embedding result to 1024, submission result is normal </p>\n\n<p>So we would like to consult, are there any restrictions on the parameter size, GPU memory, embedding size, etc. of the submitted model?</p>",
      "rawMarkdown": "As we all know, fusing multi-model output in feature retrieval competition is effective for improving map in Google Landmark Retrieval.So in order to test whether it is feasible, we simply copied multiple resnet101 networks and combine them into a single model output in tensorflow, which can forward normally locally.But we get Notebook Exceeded Allowed Compute Error in submissions. And we combine 2 resnet101 model,the submission result is normal. So where did we exceed the limit?Memory ? model parameter's size? or others ?\n\nDetail:\n1.copied and combine 6 resnet101, and embedding the result to 4096 global descriptor;\n2.model size: 887.79MB\n3.After running for about 3-4 hours, an error was reported. Less than 9 hours timeout\n4. combine 2 resnet101, embedding result to 1024, submission result is normal \n\nSo we would like to consult, are there any restrictions on the parameter size, GPU memory, embedding size, etc. of the submitted model?",
      "votes": 5
    },
    {
      "id": 970985,
      "postDate": "2020-08-15T04:24:30.453Z",
      "content": "<p>The same problem. No matter the final dimension is 512 or 1024. Do you have any new progress? </p>",
      "rawMarkdown": "The same problem. No matter the final dimension is 512 or 1024. Do you have any new progress? ",
      "votes": 1
    },
    {
      "id": 967076,
      "postDate": "2020-08-11T22:31:03.830Z",
      "content": "<p>[Off Topic] May I know why all of your team members have the same profile pic? At first look, it appears to be a single user with multiple accounts! 😅</p>",
      "rawMarkdown": "[Off Topic] May I know why all of your team members have the same profile pic? At first look, it appears to be a single user with multiple accounts! 😅",
      "votes": 2
    },
    {
      "id": 951794,
      "postDate": "2020-07-30T12:18:34.440Z",
      "content": "<p>Refer to this discussion  <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163508\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163508</a> . I think that might be the answer. It perhaps is not about the embedding length as the host clearly says. But the overall computation time for NN calculation+embedding generation being long for a single query, could be a possible reason. Since, you said a single model with an embedding of  4096 length is working , but an ensemble generating 4096 is not. </p>\n\n<p>If that is so, it may have a good reason, as well. It will make overall retrieval time longer.  Possibly they have set a threshold.</p>",
      "rawMarkdown": "Refer to this discussion  https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163508 . I think that might be the answer. It perhaps is not about the embedding length as the host clearly says. But the overall computation time for NN calculation+embedding generation being long for a single query, could be a possible reason. Since, you said a single model with an embedding of  4096 length is working , but an ensemble generating 4096 is not. \n\nIf that is so, it may have a good reason, as well. It will make overall retrieval time longer.  Possibly they have set a threshold."
    },
    {
      "id": 950864,
      "postDate": "2020-07-29T17:17:42.040Z",
      "content": "<p>I got this error when I submited a model with embedding size of 8192:  <code>Notebook Exceeded Allowed Compute</code></p>",
      "rawMarkdown": "I got this error when I submited a model with embedding size of 8192:  `Notebook Exceeded Allowed Compute`",
      "replies": [
        {
          "id": 951274,
          "postDate": "2020-07-30T03:08:13.323Z",
          "content": "<p>try embedding this model to 512, and the result is normal?</p>",
          "rawMarkdown": "try embedding this model to 512, and the result is normal?",
          "votes": 1
        },
        {
          "id": 952500,
          "postDate": "2020-07-31T01:46:28.023Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes"
        },
        {
          "id": 952726,
          "postDate": "2020-07-31T07:06:29.543Z",
          "content": "<p>so, embedding size Is one of the factors, but not all of it</p>",
          "rawMarkdown": "so, embedding size Is one of the factors, but not all of it"
        },
        {
          "id": 956464,
          "postDate": "2020-08-03T14:25:42.907Z",
          "content": "<p>How is the limit of the dimension ? Dimension=8192 makes memory error, but Dimension=4096 is OK ?</p>",
          "rawMarkdown": "How is the limit of the dimension ? Dimension=8192 makes memory error, but Dimension=4096 is OK ?"
        }
      ]
    },
    {
      "id": 949924,
      "postDate": "2020-07-29T04:45:30.850Z",
      "content": "<p>I think there is limit on submission run time of three hours. its not the commit runtime of 9 hours. lets see what the host says..</p>",
      "rawMarkdown": "I think there is limit on submission run time of three hours. its not the commit runtime of 9 hours. lets see what the host says..",
      "replies": [
        {
          "id": 951275,
          "postDate": "2020-07-30T03:08:42.353Z",
          "content": "<p>also waiting</p>",
          "rawMarkdown": "also waiting"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 950012,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2020-07-29T05:56:16.210000",
      "content": "<p>Maybe try submitting a single model with 4096 global descriptor to make sure it is not a memory error?</p>\n\n<p>Btw you say \"we\" but you seem to be a solo competitor? Who is \"we\"?</p>",
      "votes": 14,
      "replies": [
        {
          "id": 950109,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-29T07:30:48.020000",
          "content": "",
          "votes": -4,
          "replies": []
        },
        {
          "id": 950121,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-07-29T07:46:08.843000",
          "content": "<p>As long as you don't share any information with each other until you merge teams, it must be according to Kaggle rules. But I would advise you to be careful because you guys have great performance, it would be very bad if you get disqualified because of private sharing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 950161,
          "author_name": "smileimg",
          "author_url": "",
          "post_date": "2020-07-29T08:09:27.400000",
          "content": "<p>Thanks for your advice,Let us refocus on this issue instead of information about us.</p>",
          "votes": -4,
          "replies": []
        },
        {
          "id": 950256,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2020-07-29T09:52:07.850000",
          "content": "<p><a href=\"/renhui111111\">@renhui111111</a> You should have a read at this discussion : <a href=\"https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294\">https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294</a> </p>\n\n<p>This \"team\" was removed after the end of the competition, for they shared information before merging ...</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 950395,
          "author_name": "smileimg",
          "author_url": "",
          "post_date": "2020-07-29T11:21:13.577000",
          "content": "<p>Maybe try submitting a single model with 4096 global descriptor to make sure it is not a memory error?We try submmit it, and the result is normal</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 950423,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-07-29T11:40:42.427000",
          "content": "<p>Okay, then at least it narrows it down. So the issue is not the matching part of the pipeline.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 970985,
      "author_name": "Lei",
      "author_url": "",
      "post_date": "2020-08-15T04:24:30.453000",
      "content": "<p>The same problem. No matter the final dimension is 512 or 1024. Do you have any new progress? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 967076,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-08-11T22:31:03.830000",
      "content": "<p>[Off Topic] May I know why all of your team members have the same profile pic? At first look, it appears to be a single user with multiple accounts! 😅</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 951794,
      "author_name": "Suvronil",
      "author_url": "",
      "post_date": "2020-07-30T12:18:34.440000",
      "content": "<p>Refer to this discussion  <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163508\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163508</a> . I think that might be the answer. It perhaps is not about the embedding length as the host clearly says. But the overall computation time for NN calculation+embedding generation being long for a single query, could be a possible reason. Since, you said a single model with an embedding of  4096 length is working , but an ensemble generating 4096 is not. </p>\n\n<p>If that is so, it may have a good reason, as well. It will make overall retrieval time longer.  Possibly they have set a threshold.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 950864,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-07-29T17:17:42.040000",
      "content": "<p>I got this error when I submited a model with embedding size of 8192:  <code>Notebook Exceeded Allowed Compute</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 951274,
          "author_name": "smileimg",
          "author_url": "",
          "post_date": "2020-07-30T03:08:13.323000",
          "content": "<p>try embedding this model to 512, and the result is normal?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952500,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2020-07-31T01:46:28.023000",
          "content": "<p>yes</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 952726,
          "author_name": "smileimg",
          "author_url": "",
          "post_date": "2020-07-31T07:06:29.543000",
          "content": "<p>so, embedding size Is one of the factors, but not all of it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 956464,
          "author_name": "toshi_k",
          "author_url": "",
          "post_date": "2020-08-03T14:25:42.907000",
          "content": "<p>How is the limit of the dimension ? Dimension=8192 makes memory error, but Dimension=4096 is OK ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 949924,
      "author_name": "Uday Kumar Gurugubelli",
      "author_url": "",
      "post_date": "2020-07-29T04:45:30.850000",
      "content": "<p>I think there is limit on submission run time of three hours. its not the commit runtime of 9 hours. lets see what the host says..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 951275,
          "author_name": "smileimg",
          "author_url": "",
          "post_date": "2020-07-30T03:08:42.353000",
          "content": "<p>also waiting</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "950012": "Maybe try submitting a single model with 4096 global descriptor to make sure it is not a memory error?\n\nBtw you say \"we\" but you seem to be a solo competitor? Who is \"we\"?",
    "949880": "As we all know, fusing multi-model output in feature retrieval competition is effective for improving map in Google Landmark Retrieval.So in order to test whether it is feasible, we simply copied multiple resnet101 networks and combine them into a single model output in tensorflow, which can forward normally locally.But we get Notebook Exceeded Allowed Compute Error in submissions. And we combine 2 resnet101 model,the submission result is normal. So where did we exceed the limit?Memory ? model parameter's size? or others ?\n\nDetail:\n1.copied and combine 6 resnet101, and embedding the result to 4096 global descriptor;\n2.model size: 887.79MB\n3.After running for about 3-4 hours, an error was reported. Less than 9 hours timeout\n4. combine 2 resnet101, embedding result to 1024, submission result is normal \n\nSo we would like to consult, are there any restrictions on the parameter size, GPU memory, embedding size, etc. of the submitted model?",
    "970985": "The same problem. No matter the final dimension is 512 or 1024. Do you have any new progress? ",
    "967076": "[Off Topic] May I know why all of your team members have the same profile pic? At first look, it appears to be a single user with multiple accounts! 😅",
    "951794": "Refer to this discussion  https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163508 . I think that might be the answer. It perhaps is not about the embedding length as the host clearly says. But the overall computation time for NN calculation+embedding generation being long for a single query, could be a possible reason. Since, you said a single model with an embedding of  4096 length is working , but an ensemble generating 4096 is not. \n\nIf that is so, it may have a good reason, as well. It will make overall retrieval time longer.  Possibly they have set a threshold.",
    "950864": "I got this error when I submited a model with embedding size of 8192:  `Notebook Exceeded Allowed Compute`",
    "949924": "I think there is limit on submission run time of three hours. its not the commit runtime of 9 hours. lets see what the host says.."
  }
}