{
  "id": 88334,
  "title": "Is deeper better?",
  "url": "/competitions/imet-2019-fgvc6/discussion/88334",
  "author_name": "Peiyuan Liao",
  "post_date": "2019-04-08T00:01:17.206000",
  "votes": 4,
  "comment_count": 20,
  "views": 0,
  "content": "<p>[deleted due to new findings] See below.</p>",
  "messages": [
    {
      "id": 510807,
      "postDate": "2019-04-09T13:42:02.760Z",
      "content": "<p>After my trying, deeper is better, my LB score ResNet152(0.610) &gt; ResNet101 &gt; ResNet50</p>",
      "rawMarkdown": "After my trying, deeper is better, my LB score ResNet152(0.610) &gt; ResNet101 &gt; ResNet50",
      "votes": 4,
      "replies": [
        {
          "id": 510918,
          "postDate": "2019-04-09T15:49:08.557Z",
          "content": "<p>Hi <a href=\"/yanglehan\">@yanglehan</a>, could you please share which image size and batch size do you use? </p>",
          "rawMarkdown": "Hi @yanglehan, could you please share which image size and batch size do you use? "
        },
        {
          "id": 511480,
          "postDate": "2019-04-10T04:17:59.787Z",
          "content": "<p>288</p>",
          "rawMarkdown": "288",
          "votes": 1
        },
        {
          "id": 516357,
          "postDate": "2019-04-14T03:48:30.170Z",
          "content": "<p>Hi <a href=\"/yanglehan\">@yanglehan</a> </p>\n\n<blockquote>\n  <p>ResNet152(0.610) </p>\n</blockquote>\n\n<p>Is it KFOLD results or just single fold?</p>",
          "rawMarkdown": "Hi @yanglehan \n&gt; ResNet152(0.610) \n\nIs it KFOLD results or just single fold?"
        },
        {
          "id": 516385,
          "postDate": "2019-04-14T04:28:09.513Z",
          "content": "<p>5 folds</p>",
          "rawMarkdown": "5 folds\n"
        },
        {
          "id": 519763,
          "postDate": "2019-04-19T15:50:48.870Z",
          "content": "<p>I have a questions：the kernel have the time to run 5-CV? Or you train the model in your GPU, and inference in Kernel by your trained weights.</p>",
          "rawMarkdown": "I have a questions：the kernel have the time to run 5-CV? Or you train the model in your GPU, and inference in Kernel by your trained weights."
        }
      ]
    },
    {
      "id": 510237,
      "postDate": "2019-04-08T21:19:28.107Z",
      "content": "<p>It's not a number of layers itself that makes a particular architecture good, it's more about module design, module connectivity and training process. The rule of thumb is \"if a particular CNN gets a better score on ImageNet, it's typically better on downstream tasks\".</p>",
      "rawMarkdown": "It's not a number of layers itself that makes a particular architecture good, it's more about module design, module connectivity and training process. The rule of thumb is \"if a particular CNN gets a better score on ImageNet, it's typically better on downstream tasks\".",
      "votes": 4,
      "replies": [
        {
          "id": 516368,
          "postDate": "2019-04-14T04:07:28.663Z",
          "content": "<p>excuse me ,what do you mean about 'downstream tasks'.</p>",
          "rawMarkdown": "excuse me ,what do you mean about 'downstream tasks'."
        }
      ]
    },
    {
      "id": 509540,
      "postDate": "2019-04-08T00:01:17.207Z",
      "content": "<p>[deleted due to new findings] See below.</p>",
      "rawMarkdown": "[deleted due to new findings] See below.",
      "votes": 4
    },
    {
      "id": 516335,
      "postDate": "2019-04-14T02:47:27.783Z",
      "content": "<p>Been trying to get a working fractalnet, dunno how much deeper I can go. </p>",
      "rawMarkdown": "Been trying to get a working fractalnet, dunno how much deeper I can go. ",
      "votes": 1,
      "replies": [
        {
          "id": 519985,
          "postDate": "2019-04-19T22:43:40.670Z",
          "content": "<p>Any luck in fractalnet? Besides deeper, wider helps too, it seems.</p>",
          "rawMarkdown": "Any luck in fractalnet? Besides deeper, wider helps too, it seems."
        },
        {
          "id": 519999,
          "postDate": "2019-04-19T23:41:29.753Z",
          "content": "<p>No luck with fractalnet yet, global columns seem to be the culprits in constantly giving me size mismatches. It should be possible to get working, though.</p>\n\n<p>I gave up on it for the time being as something else has grabbed my interest. My kernel with it on the imet data should be ready soon.</p>\n\n<p>*edit: kernel on random wired network is up. </p>",
          "rawMarkdown": "No luck with fractalnet yet, global columns seem to be the culprits in constantly giving me size mismatches. It should be possible to get working, though.\n\nI gave up on it for the time being as something else has grabbed my interest. My kernel with it on the imet data should be ready soon.\n\n*edit: kernel on random wired network is up. "
        },
        {
          "id": 520233,
          "postDate": "2019-04-20T13:36:52.280Z",
          "content": "<p><a href=\"https://github.com/khanrc/pt.fractalnet\">https://github.com/khanrc/pt.fractalnet</a> People are having a hard time reproducing fractalnet's experiment.</p>",
          "rawMarkdown": "https://github.com/khanrc/pt.fractalnet People are having a hard time reproducing fractalnet's experiment."
        },
        {
          "id": 520272,
          "postDate": "2019-04-20T15:57:16.213Z",
          "content": "<p>Myself included.</p>\n\n<p>I think this one is better, <a href=\"https://github.com/osmr/imgclsmob/blob/master/pytorch/pytorchcv/models/fractalnet_cifar.py\">https://github.com/osmr/imgclsmob/blob/master/pytorch/pytorchcv/models/fractalnet_cifar.py</a></p>\n\n<p>Though it still is quite troublesome. If you get it to work, please let me know.</p>",
          "rawMarkdown": "Myself included.\n\nI think this one is better, [https://github.com/osmr/imgclsmob/blob/master/pytorch/pytorchcv/models/fractalnet_cifar.py](https://github.com/osmr/imgclsmob/blob/master/pytorch/pytorchcv/models/fractalnet_cifar.py)\n\nThough it still is quite troublesome. If you get it to work, please let me know."
        }
      ]
    },
    {
      "id": 509633,
      "postDate": "2019-04-08T04:29:18.870Z",
      "content": "<p>I am training ResNet152, I will compare it with 101and 50</p>",
      "rawMarkdown": "I am training ResNet152, I will compare it with 101and 50",
      "votes": 1
    },
    {
      "id": 509585,
      "postDate": "2019-04-08T02:06:03.203Z",
      "content": "<p>Densenet is so big.</p>",
      "rawMarkdown": "Densenet is so big.",
      "votes": 1,
      "replies": [
        {
          "id": 509720,
          "postDate": "2019-04-08T07:30:17.607Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 516268,
      "postDate": "2019-04-13T22:35:09.857Z",
      "content": "<p>The point i'm trying to make here is that deeper nets suffered from being harder to train and having longer inference to time, e.g., I've failed to train a proper NASnet in the doodle competition due to the limitation of my GPU resources, and my top-ranking model in TGS competition only used a resnet34 backbone. And considering the multi-faced nature of \"deeper\" - increase of GFLOPs does not always align with the increase of memory usage-, I'm opening this question to those who would like to share their insights on this dataset particularly.</p>",
      "rawMarkdown": "The point i'm trying to make here is that deeper nets suffered from being harder to train and having longer inference to time, e.g., I've failed to train a proper NASnet in the doodle competition due to the limitation of my GPU resources, and my top-ranking model in TGS competition only used a resnet34 backbone. And considering the multi-faced nature of \"deeper\" - increase of GFLOPs does not always align with the increase of memory usage-, I'm opening this question to those who would like to share their insights on this dataset particularly."
    },
    {
      "id": 513190,
      "postDate": "2019-04-11T03:16:04.907Z",
      "content": "<p>Shall we go deeper ? May be that's a problem. In many cases, deeper means better, but no all. Just like the number of training times, with  increasingly of  frequency of training, the problem of overfitting emerge. So, in my viewpoint, the suitabel deepth and training is vital impotance of improving  accuracy.</p>",
      "rawMarkdown": "Shall we go deeper ? May be that's a problem. In many cases, deeper means better, but no all. Just like the number of training times, with  increasingly of  frequency of training, the problem of overfitting emerge. So, in my viewpoint, the suitabel deepth and training is vital impotance of improving  accuracy."
    },
    {
      "id": 511533,
      "postDate": "2019-04-10T05:11:04.040Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 509547,
      "postDate": "2019-04-08T00:32:39.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 510807,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-09T13:42:02.760000",
      "content": "<p>After my trying, deeper is better, my LB score ResNet152(0.610) &gt; ResNet101 &gt; ResNet50</p>",
      "votes": 4,
      "replies": [
        {
          "id": 510918,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-04-09T15:49:08.557000",
          "content": "<p>Hi <a href=\"/yanglehan\">@yanglehan</a>, could you please share which image size and batch size do you use? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 511480,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-10T04:17:59.787000",
          "content": "<p>288</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 516357,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-04-14T03:48:30.170000",
          "content": "<p>Hi <a href=\"/yanglehan\">@yanglehan</a> </p>\n\n<blockquote>\n  <p>ResNet152(0.610) </p>\n</blockquote>\n\n<p>Is it KFOLD results or just single fold?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 516385,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-14T04:28:09.513000",
          "content": "<p>5 folds</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 519763,
          "author_name": "Harleys Zhang",
          "author_url": "",
          "post_date": "2019-04-19T15:50:48.870000",
          "content": "<p>I have a questions：the kernel have the time to run 5-CV? Or you train the model in your GPU, and inference in Kernel by your trained weights.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 510237,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-04-08T21:19:28.107000",
      "content": "<p>It's not a number of layers itself that makes a particular architecture good, it's more about module design, module connectivity and training process. The rule of thumb is \"if a particular CNN gets a better score on ImageNet, it's typically better on downstream tasks\".</p>",
      "votes": 4,
      "replies": [
        {
          "id": 516368,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-14T04:07:28.663000",
          "content": "<p>excuse me ,what do you mean about 'downstream tasks'.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 516335,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-04-14T02:47:27.783000",
      "content": "<p>Been trying to get a working fractalnet, dunno how much deeper I can go. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 519985,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2019-04-19T22:43:40.670000",
          "content": "<p>Any luck in fractalnet? Besides deeper, wider helps too, it seems.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 519999,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-04-19T23:41:29.753000",
          "content": "<p>No luck with fractalnet yet, global columns seem to be the culprits in constantly giving me size mismatches. It should be possible to get working, though.</p>\n\n<p>I gave up on it for the time being as something else has grabbed my interest. My kernel with it on the imet data should be ready soon.</p>\n\n<p>*edit: kernel on random wired network is up. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 520233,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2019-04-20T13:36:52.280000",
          "content": "<p><a href=\"https://github.com/khanrc/pt.fractalnet\">https://github.com/khanrc/pt.fractalnet</a> People are having a hard time reproducing fractalnet's experiment.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 520272,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-04-20T15:57:16.213000",
          "content": "<p>Myself included.</p>\n\n<p>I think this one is better, <a href=\"https://github.com/osmr/imgclsmob/blob/master/pytorch/pytorchcv/models/fractalnet_cifar.py\">https://github.com/osmr/imgclsmob/blob/master/pytorch/pytorchcv/models/fractalnet_cifar.py</a></p>\n\n<p>Though it still is quite troublesome. If you get it to work, please let me know.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 509633,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-08T04:29:18.870000",
      "content": "<p>I am training ResNet152, I will compare it with 101and 50</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 509585,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-04-08T02:06:03.203000",
      "content": "<p>Densenet is so big.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 509720,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-08T07:30:17.607000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 516268,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2019-04-13T22:35:09.857000",
      "content": "<p>The point i'm trying to make here is that deeper nets suffered from being harder to train and having longer inference to time, e.g., I've failed to train a proper NASnet in the doodle competition due to the limitation of my GPU resources, and my top-ranking model in TGS competition only used a resnet34 backbone. And considering the multi-faced nature of \"deeper\" - increase of GFLOPs does not always align with the increase of memory usage-, I'm opening this question to those who would like to share their insights on this dataset particularly.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 513190,
      "author_name": "xutao",
      "author_url": "",
      "post_date": "2019-04-11T03:16:04.907000",
      "content": "<p>Shall we go deeper ? May be that's a problem. In many cases, deeper means better, but no all. Just like the number of training times, with  increasingly of  frequency of training, the problem of overfitting emerge. So, in my viewpoint, the suitabel deepth and training is vital impotance of improving  accuracy.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 511533,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-10T05:11:04.040000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 509547,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-08T00:32:39.550000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "510807": "After my trying, deeper is better, my LB score ResNet152(0.610) &gt; ResNet101 &gt; ResNet50",
    "510237": "It's not a number of layers itself that makes a particular architecture good, it's more about module design, module connectivity and training process. The rule of thumb is \"if a particular CNN gets a better score on ImageNet, it's typically better on downstream tasks\".",
    "509540": "[deleted due to new findings] See below.",
    "516335": "Been trying to get a working fractalnet, dunno how much deeper I can go. ",
    "509633": "I am training ResNet152, I will compare it with 101and 50",
    "509585": "Densenet is so big.",
    "516268": "The point i'm trying to make here is that deeper nets suffered from being harder to train and having longer inference to time, e.g., I've failed to train a proper NASnet in the doodle competition due to the limitation of my GPU resources, and my top-ranking model in TGS competition only used a resnet34 backbone. And considering the multi-faced nature of \"deeper\" - increase of GFLOPs does not always align with the increase of memory usage-, I'm opening this question to those who would like to share their insights on this dataset particularly.",
    "513190": "Shall we go deeper ? May be that's a problem. In many cases, deeper means better, but no all. Just like the number of training times, with  increasingly of  frequency of training, the problem of overfitting emerge. So, in my viewpoint, the suitabel deepth and training is vital impotance of improving  accuracy.",
    "511533": "",
    "509547": ""
  }
}