{
  "id": 179964,
  "title": "How to make a simple CNN model for this dataset?",
  "url": "/competitions/landmark-recognition-2020/discussion/179964",
  "author_name": "",
  "post_date": "2020-09-03T12:46:19.353421600Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I trying to train a simple CNN architecture and when I am trying to have linear layer with 80000 out channels I am running out of memory. How are you dealing with so many classes?<br>\nNote: I have done all of my work on CPU till now. Is GPU allowing to have 80000 out channels in the last layer?</p>",
  "messages": [
    {
      "id": "996630",
      "postDate": "09/03/2020 12:46:19",
      "content": "<p>I trying to train a simple CNN architecture and when I am trying to have linear layer with 80000 out channels I am running out of memory. How are you dealing with so many classes?<br>\nNote: I have done all of my work on CPU till now. Is GPU allowing to have 80000 out channels in the last layer?</p>",
      "rawMarkdown": "I trying to train a simple CNN architecture and when I am trying to have linear layer with 80000 out channels I am running out of memory. How are you dealing with so many classes?\nNote: I have done all of my work on CPU till now. Is GPU allowing to have 80000 out channels in the last layer?",
      "votes": null
    },
    {
      "id": "996703",
      "postDate": "09/03/2020 13:54:59",
      "content": "<p><a href=\"https://github.com/mayukh18/Google-Landmark-Recognition-Retrieval-2019\" target=\"_blank\">https://github.com/mayukh18/Google-Landmark-Recognition-Retrieval-2019</a><br>\nYou can look this code written by Mayukh Bhattacharyya.</p>",
      "rawMarkdown": "https://github.com/mayukh18/Google-Landmark-Recognition-Retrieval-2019\nYou can look this code written by Mayukh Bhattacharyya.",
      "votes": null
    },
    {
      "id": "996948",
      "postDate": "09/03/2020 16:47:12",
      "content": "<p>Memory is definitively a problem when dealing with such large number of classes. What you could do to reduce the memory requirements and number of parameters of your model is to reduce the number of features of your ensemble. </p>\n<p>For example, ResNet101 has 2048 convolutional features so your last FC layer will have 2048x80313 parameters. However, if you add an extra FC layer with bias after pooling to reduce the 2048 features to 512 your model will have 2048x512 + 512x80313, which is roughly 3.9 times less parameters and operations. Additionally, this extra FC + bias layer will also work towards whitening your features, which reduce co-occurrences and helps the score for retrieval tasks. </p>\n<p>Another option is to train your model without using classification as a proxy. In this case, you would need to train using contrastive or triplet loss.</p>\n<p>But bear in mind that this competition has a lot of data, which means you won't be able to do much using CPU. You'll likely need some GPUs or a TPU to do some work.</p>",
      "rawMarkdown": "Memory is definitively a problem when dealing with such large number of classes. What you could do to reduce the memory requirements and number of parameters of your model is to reduce the number of features of your ensemble. \n\nFor example, ResNet101 has 2048 convolutional features so your last FC layer will have 2048x80313 parameters. However, if you add an extra FC layer with bias after pooling to reduce the 2048 features to 512 your model will have 2048x512 + 512x80313, which is roughly 3.9 times less parameters and operations. Additionally, this extra FC + bias layer will also work towards whitening your features, which reduce co-occurrences and helps the score for retrieval tasks. \n\nAnother option is to train your model without using classification as a proxy. In this case, you would need to train using contrastive or triplet loss.\n\nBut bear in mind that this competition has a lot of data, which means you won't be able to do much using CPU. You'll likely need some GPUs or a TPU to do some work.",
      "votes": null
    },
    {
      "id": "1003561",
      "postDate": "09/09/2020 05:03:10",
      "content": "<p>Hi, Is it mandatory to have training loops(i.e to use private train set) in submission code ? I have pre-trained model which is trained with public training set and in submission code i have only inference part, i always get 0.0000 as public score? Can you please help me   </p>",
      "rawMarkdown": "Hi, Is it mandatory to have training loops(i.e to use private train set) in submission code ? I have pre-trained model which is trained with public training set and in submission code i have only inference part, i always get 0.0000 as public score? Can you please help me",
      "votes": null
    },
    {
      "id": "1004206",
      "postDate": "09/09/2020 14:33:52",
      "content": "<p><a href=\"https://www.kaggle.com/shanmugam212\" target=\"_blank\">@shanmugam212</a> have you got an answer to this? Even I am stuck. Somewhere they are saying that we can train outside. But yeah training loops are still there for private dataset.</p>\n<p>People are saying even for baseline model, train embeddings can be generated outside and just loaded from a file and used. But what if my embeddings are having some files which are not in their private training dataset?</p>",
      "rawMarkdown": "shanmugam212 have you got an answer to this? Even I am stuck. Somewhere they are saying that we can train outside. But yeah training loops are still there for private dataset.\n\nPeople are saying even for baseline model, train embeddings can be generated outside and just loaded from a file and used. But what if my embeddings are having some files which are not in their private training dataset?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 996703,
      "author_name": "shivam17818",
      "author_url": "",
      "post_date": "09/03/2020 13:54:59",
      "content": "<p><a href=\"https://github.com/mayukh18/Google-Landmark-Recognition-Retrieval-2019\" target=\"_blank\">https://github.com/mayukh18/Google-Landmark-Recognition-Retrieval-2019</a><br>\nYou can look this code written by Mayukh Bhattacharyya.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 996948,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "09/03/2020 16:47:12",
      "content": "<p>Memory is definitively a problem when dealing with such large number of classes. What you could do to reduce the memory requirements and number of parameters of your model is to reduce the number of features of your ensemble. </p>\n<p>For example, ResNet101 has 2048 convolutional features so your last FC layer will have 2048x80313 parameters. However, if you add an extra FC layer with bias after pooling to reduce the 2048 features to 512 your model will have 2048x512 + 512x80313, which is roughly 3.9 times less parameters and operations. Additionally, this extra FC + bias layer will also work towards whitening your features, which reduce co-occurrences and helps the score for retrieval tasks. </p>\n<p>Another option is to train your model without using classification as a proxy. In this case, you would need to train using contrastive or triplet loss.</p>\n<p>But bear in mind that this competition has a lot of data, which means you won't be able to do much using CPU. You'll likely need some GPUs or a TPU to do some work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1003561,
          "author_name": "shanmugam212",
          "author_url": "",
          "post_date": "09/09/2020 05:03:10",
          "content": "<p>Hi, Is it mandatory to have training loops(i.e to use private train set) in submission code ? I have pre-trained model which is trained with public training set and in submission code i have only inference part, i always get 0.0000 as public score? Can you please help me   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1004206,
          "author_name": "rohitdeepu17",
          "author_url": "",
          "post_date": "09/09/2020 14:33:52",
          "content": "<p><a href=\"https://www.kaggle.com/shanmugam212\" target=\"_blank\">@shanmugam212</a> have you got an answer to this? Even I am stuck. Somewhere they are saying that we can train outside. But yeah training loops are still there for private dataset.</p>\n<p>People are saying even for baseline model, train embeddings can be generated outside and just loaded from a file and used. But what if my embeddings are having some files which are not in their private training dataset?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "996630": "I trying to train a simple CNN architecture and when I am trying to have linear layer with 80000 out channels I am running out of memory. How are you dealing with so many classes?\nNote: I have done all of my work on CPU till now. Is GPU allowing to have 80000 out channels in the last layer?",
    "996703": "https://github.com/mayukh18/Google-Landmark-Recognition-Retrieval-2019\nYou can look this code written by Mayukh Bhattacharyya.",
    "996948": "Memory is definitively a problem when dealing with such large number of classes. What you could do to reduce the memory requirements and number of parameters of your model is to reduce the number of features of your ensemble. \n\nFor example, ResNet101 has 2048 convolutional features so your last FC layer will have 2048x80313 parameters. However, if you add an extra FC layer with bias after pooling to reduce the 2048 features to 512 your model will have 2048x512 + 512x80313, which is roughly 3.9 times less parameters and operations. Additionally, this extra FC + bias layer will also work towards whitening your features, which reduce co-occurrences and helps the score for retrieval tasks. \n\nAnother option is to train your model without using classification as a proxy. In this case, you would need to train using contrastive or triplet loss.\n\nBut bear in mind that this competition has a lot of data, which means you won't be able to do much using CPU. You'll likely need some GPUs or a TPU to do some work.",
    "1003561": "Hi, Is it mandatory to have training loops(i.e to use private train set) in submission code ? I have pre-trained model which is trained with public training set and in submission code i have only inference part, i always get 0.0000 as public score? Can you please help me",
    "1004206": "shanmugam212 have you got an answer to this? Even I am stuck. Somewhere they are saying that we can train outside. But yeah training loops are still there for private dataset.\n\nPeople are saying even for baseline model, train embeddings can be generated outside and just loaded from a file and used. But what if my embeddings are having some files which are not in their private training dataset?"
  },
  "source": "meta"
}