{
  "id": 139348,
  "title": "CV vs. public LB",
  "url": "/competitions/imet-2020-fgvc7/discussion/139348",
  "author_name": "",
  "post_date": "2020-03-28T13:47:34.985696800Z",
  "votes": 8,
  "comment_count": 24,
  "views": 0,
  "content": "<p>As shown in <a href=\"https://www.kaggle.com/yasufuminakama/imet-2020-pytorch-resnet18-inference?scriptVersionId=30995441\">my notebook</a></p>\n\n<p>model: Resnet18 \nsingle fold CV: 0.5585\nLB: 0.613</p>\n\n<p>Much difference🤔 </p>",
  "messages": [
    {
      "id": "789243",
      "postDate": "03/28/2020 13:47:34",
      "content": "<p>As shown in <a href=\"https://www.kaggle.com/yasufuminakama/imet-2020-pytorch-resnet18-inference?scriptVersionId=30995441\">my notebook</a></p>\n\n<p>model: Resnet18 \nsingle fold CV: 0.5585\nLB: 0.613</p>\n\n<p>Much difference🤔 </p>",
      "rawMarkdown": "As shown in [my notebook](https://www.kaggle.com/yasufuminakama/imet-2020-pytorch-resnet18-inference?scriptVersionId=30995441)\n\nmodel: Resnet18 \nsingle fold CV: 0.5585\nLB: 0.613\n\nMuch difference🤔",
      "votes": null
    },
    {
      "id": "790934",
      "postDate": "03/30/2020 00:54:51",
      "content": "<p>model: seresnext50 \nsingle fold CV: 0.666\nLB: 0.581</p>",
      "rawMarkdown": "model: seresnext50 \nsingle fold CV: 0.666\nLB: 0.581",
      "votes": null
    },
    {
      "id": "791092",
      "postDate": "03/30/2020 05:20:21",
      "content": "<p>Model: ResNet50\nCV: 0.557\nLB: 0.623</p>",
      "rawMarkdown": "Model: ResNet50\nCV: 0.557\nLB: 0.623",
      "votes": null
    },
    {
      "id": "791675",
      "postDate": "03/30/2020 15:33:56",
      "content": "<p>The public lb contains validation data. Also we have more fine-grained attributes this time. This could be the reason for this difference.</p>",
      "rawMarkdown": "The public lb contains validation data. Also we have more fine-grained attributes this time. This could be the reason for this difference.",
      "votes": null
    },
    {
      "id": "791718",
      "postDate": "03/30/2020 16:01:03",
      "content": "<p>\"validation data\" is part of the training data??</p>",
      "rawMarkdown": "\"validation data\" is part of the training data??",
      "votes": null
    },
    {
      "id": "791734",
      "postDate": "03/30/2020 16:09:36",
      "content": "<p>No, \"validation data\" is equivalent to \"public test\".</p>",
      "rawMarkdown": "No, \"validation data\" is equivalent to \"public test\".",
      "votes": null
    },
    {
      "id": "792173",
      "postDate": "03/31/2020 00:29:11",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!",
      "votes": null
    },
    {
      "id": "792968",
      "postDate": "03/31/2020 17:28:55",
      "content": "<p>had a bug. 0.715 is the correct score</p>",
      "rawMarkdown": "had a bug. 0.715 is the correct score",
      "votes": null
    },
    {
      "id": "793180",
      "postDate": "03/31/2020 21:28:05",
      "content": "<p>[update]\nmodel: seresnext50 (training 3 hours on kaggle TPU)\nsingle fold CV: 0.618\nLB: 0.659\n(I need augmentations &amp; more epochs &amp; larger image size)</p>",
      "rawMarkdown": "[update]\nmodel: seresnext50 (training 3 hours on kaggle TPU)\nsingle fold CV: 0.618\nLB: 0.659\n(I need augmentations &amp; more epochs &amp; larger image size)",
      "votes": null
    },
    {
      "id": "793217",
      "postDate": "03/31/2020 22:18:00",
      "content": "<p>thanks <a href=\"/igorkrashenyi\">@igorkrashenyi</a> that's great! I'm wondering which is going to be the best score this year, because you already improved last year best score 0.672 :)</p>",
      "rawMarkdown": "thanks @igorkrashenyi that's great! I'm wondering which is going to be the best score this year, because you already improved last year best score 0.672 :)",
      "votes": null
    },
    {
      "id": "793219",
      "postDate": "03/31/2020 22:19:34",
      "content": "<p>Great start! <a href=\"/yasufuminakama\">@yasufuminakama</a> did you create the tfrecords? is the dataset with tfrecords available?</p>",
      "rawMarkdown": "Great start! @yasufuminakama did you create the tfrecords? is the dataset with tfrecords available?",
      "votes": null
    },
    {
      "id": "793623",
      "postDate": "04/01/2020 06:50:50",
      "content": "<p>I used TPU with PyTorch-xla :)\nSo I haven't created tfrecords.</p>",
      "rawMarkdown": "I used TPU with PyTorch-xla :)\nSo I haven't created tfrecords.",
      "votes": null
    },
    {
      "id": "797846",
      "postDate": "04/05/2020 00:24:35",
      "content": "<p>[update]\nsingle fold CV: 0.656\nLB: 0.707</p>",
      "rawMarkdown": "[update]\nsingle fold CV: 0.656\nLB: 0.707",
      "votes": null
    },
    {
      "id": "799476",
      "postDate": "04/06/2020 13:17:06",
      "content": "<p>I guess not more than 0.8 </p>",
      "rawMarkdown": "I guess not more than 0.8",
      "votes": null
    },
    {
      "id": "800333",
      "postDate": "04/07/2020 10:03:34",
      "content": "<p>you are still using pytorch-xla and kaggle TPU? nice :)</p>",
      "rawMarkdown": "you are still using pytorch-xla and kaggle TPU? nice :)",
      "votes": null
    },
    {
      "id": "800335",
      "postDate": "04/07/2020 10:06:54",
      "content": "<p>awesome <a href=\"/igorkrashenyi\">@igorkrashenyi</a>. Unless someone finds the magic feature :)\nHow long it takes to train. Is it possible with kernels?</p>",
      "rawMarkdown": "awesome @igorkrashenyi. Unless someone finds the magic feature :)\nHow long it takes to train. Is it possible with kernels?",
      "votes": null
    },
    {
      "id": "800396",
      "postDate": "04/07/2020 11:27:46",
      "content": "<p>of course, it is possible :) \nIt takes around 3-6 hours to get 0.666 on my local CV depending on the model architecture.  </p>",
      "rawMarkdown": "of course, it is possible :) \nIt takes around 3-6 hours to get 0.666 on my local CV depending on the model architecture.",
      "votes": null
    },
    {
      "id": "800417",
      "postDate": "04/07/2020 11:47:50",
      "content": "<p>I mean, to train in kernel for &lt;9h  and score .75 :) single model. if that is possible I could afford 3 runs weekly :D ensemble and voila you get 0.8 :D</p>",
      "rawMarkdown": "I mean, to train in kernel for &lt;9h  and score .75 :) single model. if that is possible I could afford 3 runs weekly :D ensemble and voila you get 0.8 :D",
      "votes": null
    },
    {
      "id": "801334",
      "postDate": "04/08/2020 11:27:34",
      "content": "<p>I fine-tuned with more augmentations  &amp; larger image size on GPU using TPU-trained model :)</p>",
      "rawMarkdown": "I fine-tuned with more augmentations  &amp; larger image size on GPU using TPU-trained model :)",
      "votes": null
    },
    {
      "id": "809959",
      "postDate": "04/16/2020 16:09:47",
      "content": "<p>model: resnet50\nsingle fold CV: 0.649\nLB: 0.699</p>",
      "rawMarkdown": "model: resnet50\nsingle fold CV: 0.649\nLB: 0.699",
      "votes": null
    },
    {
      "id": "833971",
      "postDate": "05/05/2020 07:39:58",
      "content": "<p>How did you set up pytorch xla without internet?</p>",
      "rawMarkdown": "How did you set up pytorch xla without internet?",
      "votes": null
    },
    {
      "id": "833977",
      "postDate": "05/05/2020 07:47:55",
      "content": "<p>I set up pytorch xla with internet for training, but I don't use pytorch xla for inference(No internet access enabled for inference)</p>",
      "rawMarkdown": "I set up pytorch xla with internet for training, but I don't use pytorch xla for inference(No internet access enabled for inference)",
      "votes": null
    },
    {
      "id": "834281",
      "postDate": "05/05/2020 13:03:39",
      "content": "<p>Is there any way how can pre-trained weights on TPU execute without internet on kaggle notebook for inference?</p>",
      "rawMarkdown": "Is there any way how can pre-trained weights on TPU execute without internet on kaggle notebook for inference?",
      "votes": null
    },
    {
      "id": "834307",
      "postDate": "05/05/2020 13:29:48",
      "content": "<p>You need to save model as CPU after you trained model on TPU.\nFor example, \n```</p>\n\n<h1>Accelerator: TPU</h1>\n\n<p>def _run():\n    device = xm.xla_device()\n    model = my_model.to(device)\n    # train model\n    # save model on TPU\n    torch.save({'model': model, 'threshold': best_thresh}, 'model_tpu.pth')</p>\n\n<h1>Start training processes</h1>\n\n<p>def _mp_fn(rank, flags):\n    torch.set_default_tensor_type('torch.FloatTensor')\n    a = _run()</p>\n\n<p>FLAGS = {}\nxmp.spawn(_mp_fn, args=(FLAGS,), nprocs=1, start_method='fork')</p>\n\n<h1>save model as CPU</h1>\n\n<p>state = torch.load('model_tpu.pth')\ntorch.save({'model': state['model'].to('cpu').state_dict(), 'threshold': state['threshold']}, 'model_cpu.pth')\n```\nHope this helps :)</p>",
      "rawMarkdown": "You need to save model as CPU after you trained model on TPU.\nFor example, \n```\n# Accelerator: TPU\ndef _run():\n    device = xm.xla_device()\n    model = my_model.to(device)\n    # train model\n    # save model on TPU\n    torch.save({'model': model, 'threshold': best_thresh}, 'model_tpu.pth')\n\n# Start training processes\ndef _mp_fn(rank, flags):\n    torch.set_default_tensor_type('torch.FloatTensor')\n    a = _run()\n\nFLAGS = {}\nxmp.spawn(_mp_fn, args=(FLAGS,), nprocs=1, start_method='fork')\n\n# save model as CPU\nstate = torch.load('model_tpu.pth')\ntorch.save({'model': state['model'].to('cpu').state_dict(), 'threshold': state['threshold']}, 'model_cpu.pth')\n```\nHope this helps :)",
      "votes": null
    },
    {
      "id": "834704",
      "postDate": "05/05/2020 17:58:26",
      "content": "<p>thank you very much.</p>",
      "rawMarkdown": "thank you very much.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 790934,
      "author_name": "igorkrashenyi",
      "author_url": "",
      "post_date": "03/30/2020 00:54:51",
      "content": "<p>model: seresnext50 \nsingle fold CV: 0.666\nLB: 0.581</p>",
      "votes": null,
      "replies": [
        {
          "id": 792968,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "03/31/2020 17:28:55",
          "content": "<p>had a bug. 0.715 is the correct score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793217,
          "author_name": "jesucristo",
          "author_url": "",
          "post_date": "03/31/2020 22:18:00",
          "content": "<p>thanks <a href=\"/igorkrashenyi\">@igorkrashenyi</a> that's great! I'm wondering which is going to be the best score this year, because you already improved last year best score 0.672 :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799476,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "04/06/2020 13:17:06",
          "content": "<p>I guess not more than 0.8 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800335,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "04/07/2020 10:06:54",
          "content": "<p>awesome <a href=\"/igorkrashenyi\">@igorkrashenyi</a>. Unless someone finds the magic feature :)\nHow long it takes to train. Is it possible with kernels?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800396,
          "author_name": "igorkrashenyi",
          "author_url": "",
          "post_date": "04/07/2020 11:27:46",
          "content": "<p>of course, it is possible :) \nIt takes around 3-6 hours to get 0.666 on my local CV depending on the model architecture.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800417,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "04/07/2020 11:47:50",
          "content": "<p>I mean, to train in kernel for &lt;9h  and score .75 :) single model. if that is possible I could afford 3 runs weekly :D ensemble and voila you get 0.8 :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 791092,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "03/30/2020 05:20:21",
      "content": "<p>Model: ResNet50\nCV: 0.557\nLB: 0.623</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 791675,
      "author_name": "codingcliff",
      "author_url": "",
      "post_date": "03/30/2020 15:33:56",
      "content": "<p>The public lb contains validation data. Also we have more fine-grained attributes this time. This could be the reason for this difference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 791718,
          "author_name": "inoueu1",
          "author_url": "",
          "post_date": "03/30/2020 16:01:03",
          "content": "<p>\"validation data\" is part of the training data??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 791734,
          "author_name": "codingcliff",
          "author_url": "",
          "post_date": "03/30/2020 16:09:36",
          "content": "<p>No, \"validation data\" is equivalent to \"public test\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 792173,
          "author_name": "inoueu1",
          "author_url": "",
          "post_date": "03/31/2020 00:29:11",
          "content": "<p>thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 793180,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "03/31/2020 21:28:05",
      "content": "<p>[update]\nmodel: seresnext50 (training 3 hours on kaggle TPU)\nsingle fold CV: 0.618\nLB: 0.659\n(I need augmentations &amp; more epochs &amp; larger image size)</p>",
      "votes": null,
      "replies": [
        {
          "id": 793219,
          "author_name": "jesucristo",
          "author_url": "",
          "post_date": "03/31/2020 22:19:34",
          "content": "<p>Great start! <a href=\"/yasufuminakama\">@yasufuminakama</a> did you create the tfrecords? is the dataset with tfrecords available?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 793623,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "04/01/2020 06:50:50",
          "content": "<p>I used TPU with PyTorch-xla :)\nSo I haven't created tfrecords.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 797846,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "04/05/2020 00:24:35",
          "content": "<p>[update]\nsingle fold CV: 0.656\nLB: 0.707</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800333,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "04/07/2020 10:03:34",
          "content": "<p>you are still using pytorch-xla and kaggle TPU? nice :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 801334,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "04/08/2020 11:27:34",
          "content": "<p>I fine-tuned with more augmentations  &amp; larger image size on GPU using TPU-trained model :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 833971,
          "author_name": "sanantoha",
          "author_url": "",
          "post_date": "05/05/2020 07:39:58",
          "content": "<p>How did you set up pytorch xla without internet?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 833977,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "05/05/2020 07:47:55",
          "content": "<p>I set up pytorch xla with internet for training, but I don't use pytorch xla for inference(No internet access enabled for inference)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 834281,
          "author_name": "sanantoha",
          "author_url": "",
          "post_date": "05/05/2020 13:03:39",
          "content": "<p>Is there any way how can pre-trained weights on TPU execute without internet on kaggle notebook for inference?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 834307,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "05/05/2020 13:29:48",
          "content": "<p>You need to save model as CPU after you trained model on TPU.\nFor example, \n```</p>\n\n<h1>Accelerator: TPU</h1>\n\n<p>def _run():\n    device = xm.xla_device()\n    model = my_model.to(device)\n    # train model\n    # save model on TPU\n    torch.save({'model': model, 'threshold': best_thresh}, 'model_tpu.pth')</p>\n\n<h1>Start training processes</h1>\n\n<p>def _mp_fn(rank, flags):\n    torch.set_default_tensor_type('torch.FloatTensor')\n    a = _run()</p>\n\n<p>FLAGS = {}\nxmp.spawn(_mp_fn, args=(FLAGS,), nprocs=1, start_method='fork')</p>\n\n<h1>save model as CPU</h1>\n\n<p>state = torch.load('model_tpu.pth')\ntorch.save({'model': state['model'].to('cpu').state_dict(), 'threshold': state['threshold']}, 'model_cpu.pth')\n```\nHope this helps :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 834704,
          "author_name": "sanantoha",
          "author_url": "",
          "post_date": "05/05/2020 17:58:26",
          "content": "<p>thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 809959,
      "author_name": "alimbekovkz",
      "author_url": "",
      "post_date": "04/16/2020 16:09:47",
      "content": "<p>model: resnet50\nsingle fold CV: 0.649\nLB: 0.699</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "789243": "As shown in [my notebook](https://www.kaggle.com/yasufuminakama/imet-2020-pytorch-resnet18-inference?scriptVersionId=30995441)\n\nmodel: Resnet18 \nsingle fold CV: 0.5585\nLB: 0.613\n\nMuch difference🤔",
    "790934": "model: seresnext50 \nsingle fold CV: 0.666\nLB: 0.581",
    "791092": "Model: ResNet50\nCV: 0.557\nLB: 0.623",
    "791675": "The public lb contains validation data. Also we have more fine-grained attributes this time. This could be the reason for this difference.",
    "791718": "\"validation data\" is part of the training data??",
    "791734": "No, \"validation data\" is equivalent to \"public test\".",
    "792173": "thanks!",
    "792968": "had a bug. 0.715 is the correct score",
    "793180": "[update]\nmodel: seresnext50 (training 3 hours on kaggle TPU)\nsingle fold CV: 0.618\nLB: 0.659\n(I need augmentations &amp; more epochs &amp; larger image size)",
    "793217": "thanks @igorkrashenyi that's great! I'm wondering which is going to be the best score this year, because you already improved last year best score 0.672 :)",
    "793219": "Great start! @yasufuminakama did you create the tfrecords? is the dataset with tfrecords available?",
    "793623": "I used TPU with PyTorch-xla :)\nSo I haven't created tfrecords.",
    "797846": "[update]\nsingle fold CV: 0.656\nLB: 0.707",
    "799476": "I guess not more than 0.8",
    "800333": "you are still using pytorch-xla and kaggle TPU? nice :)",
    "800335": "awesome @igorkrashenyi. Unless someone finds the magic feature :)\nHow long it takes to train. Is it possible with kernels?",
    "800396": "of course, it is possible :) \nIt takes around 3-6 hours to get 0.666 on my local CV depending on the model architecture.",
    "800417": "I mean, to train in kernel for &lt;9h  and score .75 :) single model. if that is possible I could afford 3 runs weekly :D ensemble and voila you get 0.8 :D",
    "801334": "I fine-tuned with more augmentations  &amp; larger image size on GPU using TPU-trained model :)",
    "809959": "model: resnet50\nsingle fold CV: 0.649\nLB: 0.699",
    "833971": "How did you set up pytorch xla without internet?",
    "833977": "I set up pytorch xla with internet for training, but I don't use pytorch xla for inference(No internet access enabled for inference)",
    "834281": "Is there any way how can pre-trained weights on TPU execute without internet on kaggle notebook for inference?",
    "834307": "You need to save model as CPU after you trained model on TPU.\nFor example, \n```\n# Accelerator: TPU\ndef _run():\n    device = xm.xla_device()\n    model = my_model.to(device)\n    # train model\n    # save model on TPU\n    torch.save({'model': model, 'threshold': best_thresh}, 'model_tpu.pth')\n\n# Start training processes\ndef _mp_fn(rank, flags):\n    torch.set_default_tensor_type('torch.FloatTensor')\n    a = _run()\n\nFLAGS = {}\nxmp.spawn(_mp_fn, args=(FLAGS,), nprocs=1, start_method='fork')\n\n# save model as CPU\nstate = torch.load('model_tpu.pth')\ntorch.save({'model': state['model'].to('cpu').state_dict(), 'threshold': state['threshold']}, 'model_cpu.pth')\n```\nHope this helps :)",
    "834704": "thank you very much."
  },
  "source": "meta"
}