{
  "id": 120972,
  "title": "[solved] strange prediction accuracy",
  "url": "/competitions/pku-autonomous-driving/discussion/120972",
  "author_name": "",
  "post_date": "2019-12-10T08:19:17.943542Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>As I followed the same configuration as <a href=\"https://www.kaggle.com/phoenix9032/center-resnet-starter/notebook\">center-resnet-starter</a>, and I also tried <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">centernet-baseline</a>, but both got LB around 0.01~0.02, much lower than 0.038, I trained offline, but I checked almost every part is the same as in the two notebooks...</p>\n\n<p>Some evidence:\nTraining loss\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F0ab42b9f09519238c49bf3f0acb59c5c%2F2019-12-10%2011.25.36.png?generation=1575991574987840&amp;alt=media\" alt=\"\"></p>\n\n<p>Visualizing prediction on testset\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F2119e9ade9119d02dc75fc1a041d500b%2Fpred0.png?generation=1575965589843154&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff90412af102a2ce9b8ccb82d0bb1902a%2Fpred1.png?generation=1575965608992000&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff6a7ea678dfe97232dab5ba23ced46d6%2Fpred2.png?generation=1575965634695279&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "691522",
      "postDate": "12/10/2019 08:19:17",
      "content": "<p>As I followed the same configuration as <a href=\"https://www.kaggle.com/phoenix9032/center-resnet-starter/notebook\">center-resnet-starter</a>, and I also tried <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">centernet-baseline</a>, but both got LB around 0.01~0.02, much lower than 0.038, I trained offline, but I checked almost every part is the same as in the two notebooks...</p>\n\n<p>Some evidence:\nTraining loss\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F0ab42b9f09519238c49bf3f0acb59c5c%2F2019-12-10%2011.25.36.png?generation=1575991574987840&amp;alt=media\" alt=\"\"></p>\n\n<p>Visualizing prediction on testset\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F2119e9ade9119d02dc75fc1a041d500b%2Fpred0.png?generation=1575965589843154&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff90412af102a2ce9b8ccb82d0bb1902a%2Fpred1.png?generation=1575965608992000&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff6a7ea678dfe97232dab5ba23ced46d6%2Fpred2.png?generation=1575965634695279&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "As I followed the same configuration as [center-resnet-starter](https://www.kaggle.com/phoenix9032/center-resnet-starter/notebook), and I also tried [centernet-baseline](https://www.kaggle.com/hocop1/centernet-baseline), but both got LB around 0.01~0.02, much lower than 0.038, I trained offline, but I checked almost every part is the same as in the two notebooks...\n\nSome evidence:\nTraining loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F0ab42b9f09519238c49bf3f0acb59c5c%2F2019-12-10%2011.25.36.png?generation=1575991574987840&amp;alt=media)\n\n\n\nVisualizing prediction on testset\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F2119e9ade9119d02dc75fc1a041d500b%2Fpred0.png?generation=1575965589843154&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff90412af102a2ce9b8ccb82d0bb1902a%2Fpred1.png?generation=1575965608992000&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff6a7ea678dfe97232dab5ba23ced46d6%2Fpred2.png?generation=1575965634695279&amp;alt=media)",
      "votes": null
    },
    {
      "id": "691552",
      "postDate": "12/10/2019 09:02:30",
      "content": "<p>Hello <a href=\"/niuddd\">@niuddd</a> that is indeed very weird . I have few questions . \n1. DId you stop your training at 11th epoch and then run the predictions ? I ask this because Model tends to get very bad overfit at the end if you run longer </p>\n\n<ol>\n<li><p>Did you tweak anything in the post processing part ?</p></li>\n<li><p>How is your validation set comparison ?</p></li>\n<li><p>Most importantly can you please post your validation loss list epoch by epoch , same way that you have posted train loss ?</p></li>\n</ol>",
      "rawMarkdown": "Hello @niuddd that is indeed very weird . I have few questions . \n1. DId you stop your training at 11th epoch and then run the predictions ? I ask this because Model tends to get very bad overfit at the end if you run longer \n\n2. Did you tweak anything in the post processing part ?\n\n3. How is your validation set comparison ?\n\n4. Most importantly can you please post your validation loss list epoch by epoch , same way that you have posted train loss ?",
      "votes": null
    },
    {
      "id": "691862",
      "postDate": "12/10/2019 15:22:50",
      "content": "<p>Hi, <a href=\"/phoenix9032\">@phoenix9032</a> , thanks for your reply.\n1. set num_epochs=12, predict with checkpoint of 12th epoch\n2. yes, threshold of logits, tried both 0 and -1 for center-net (efficient-net backbone, will try resnet soon), as I said, LB are between 0.01~0.02\n3. 9% for valid-set, what does \"comparison\" mean?\n4. updated both train&amp;valid loss above</p>\n\n<p>Additional information: \n1. I managed to copy the evaluation notebook to calculate local CV on valid-set, it's &lt;0.01...\n2. I checked every part of code block by block, found 2 difference (very sure): I use batch size 8 instead of 2 because of 4 gpus, I use torch.optim.Adam instead of AdamW</p>",
      "rawMarkdown": "Hi, @phoenix9032 , thanks for your reply.\n1. set num_epochs=12, predict with checkpoint of 12th epoch\n2. yes, threshold of logits, tried both 0 and -1 for center-net (efficient-net backbone, will try resnet soon), as I said, LB are between 0.01~0.02\n3. 9% for valid-set, what does \"comparison\" mean?\n4. updated both train&amp;valid loss above\n\nAdditional information: \n1. I managed to copy the evaluation notebook to calculate local CV on valid-set, it's &lt;0.01...\n2. I checked every part of code block by block, found 2 difference (very sure): I use batch size 8 instead of 2 because of 4 gpus, I use torch.optim.Adam instead of AdamW",
      "votes": null
    },
    {
      "id": "691891",
      "postDate": "12/10/2019 16:04:38",
      "content": "<p>What happens if you run exact same code as kernel? It's better to check it's because of code or hardware at first.</p>",
      "rawMarkdown": "What happens if you run exact same code as kernel? It's better to check it's because of code or hardware at first.",
      "votes": null
    },
    {
      "id": "692238",
      "postDate": "12/11/2019 04:00:31",
      "content": "<p>I think the difference lies in the training part because using checkpoint from <a href=\"/phoenix9032\">@phoenix9032</a> notebook, I managed to reproduce the test-set prediction (99.99% similar). Local cv shows 0.08.</p>",
      "rawMarkdown": "I think the difference lies in the training part because using checkpoint from @phoenix9032 notebook, I managed to reproduce the test-set prediction (99.99% similar). Local cv shows 0.08.",
      "votes": null
    },
    {
      "id": "692594",
      "postDate": "12/11/2019 14:07:07",
      "content": "<p>The problem is subtle to me that I ignored......batch size. I used to set batch size to 8 on 4 gpus, as well as gradient accumulation 2 as default training setting......which turns to be a bad setting for this competition, I changed batch size to 4, it works then. I haven't yet figured out the reason why using a bigger (but still normal) batch size (16 vs 4) could cause drastic failure in training in this task, seems batch size is tunable.</p>",
      "rawMarkdown": "The problem is subtle to me that I ignored......batch size. I used to set batch size to 8 on 4 gpus, as well as gradient accumulation 2 as default training setting......which turns to be a bad setting for this competition, I changed batch size to 4, it works then. I haven't yet figured out the reason why using a bigger (but still normal) batch size (16 vs 4) could cause drastic failure in training in this task, seems batch size is tunable.",
      "votes": null
    },
    {
      "id": "692678",
      "postDate": "12/11/2019 16:01:29",
      "content": "<p>I have failed miserably with gradient accumulation . This is an open problem for me as well . I thought may be I am doing something wrong </p>",
      "rawMarkdown": "I have failed miserably with gradient accumulation . This is an open problem for me as well . I thought may be I am doing something wrong",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 691552,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "12/10/2019 09:02:30",
      "content": "<p>Hello <a href=\"/niuddd\">@niuddd</a> that is indeed very weird . I have few questions . \n1. DId you stop your training at 11th epoch and then run the predictions ? I ask this because Model tends to get very bad overfit at the end if you run longer </p>\n\n<ol>\n<li><p>Did you tweak anything in the post processing part ?</p></li>\n<li><p>How is your validation set comparison ?</p></li>\n<li><p>Most importantly can you please post your validation loss list epoch by epoch , same way that you have posted train loss ?</p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 691862,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "12/10/2019 15:22:50",
          "content": "<p>Hi, <a href=\"/phoenix9032\">@phoenix9032</a> , thanks for your reply.\n1. set num_epochs=12, predict with checkpoint of 12th epoch\n2. yes, threshold of logits, tried both 0 and -1 for center-net (efficient-net backbone, will try resnet soon), as I said, LB are between 0.01~0.02\n3. 9% for valid-set, what does \"comparison\" mean?\n4. updated both train&amp;valid loss above</p>\n\n<p>Additional information: \n1. I managed to copy the evaluation notebook to calculate local CV on valid-set, it's &lt;0.01...\n2. I checked every part of code block by block, found 2 difference (very sure): I use batch size 8 instead of 2 because of 4 gpus, I use torch.optim.Adam instead of AdamW</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 691891,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "12/10/2019 16:04:38",
      "content": "<p>What happens if you run exact same code as kernel? It's better to check it's because of code or hardware at first.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 692238,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "12/11/2019 04:00:31",
      "content": "<p>I think the difference lies in the training part because using checkpoint from <a href=\"/phoenix9032\">@phoenix9032</a> notebook, I managed to reproduce the test-set prediction (99.99% similar). Local cv shows 0.08.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 692594,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "12/11/2019 14:07:07",
      "content": "<p>The problem is subtle to me that I ignored......batch size. I used to set batch size to 8 on 4 gpus, as well as gradient accumulation 2 as default training setting......which turns to be a bad setting for this competition, I changed batch size to 4, it works then. I haven't yet figured out the reason why using a bigger (but still normal) batch size (16 vs 4) could cause drastic failure in training in this task, seems batch size is tunable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 692678,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "12/11/2019 16:01:29",
          "content": "<p>I have failed miserably with gradient accumulation . This is an open problem for me as well . I thought may be I am doing something wrong </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "691522": "As I followed the same configuration as [center-resnet-starter](https://www.kaggle.com/phoenix9032/center-resnet-starter/notebook), and I also tried [centernet-baseline](https://www.kaggle.com/hocop1/centernet-baseline), but both got LB around 0.01~0.02, much lower than 0.038, I trained offline, but I checked almost every part is the same as in the two notebooks...\n\nSome evidence:\nTraining loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F0ab42b9f09519238c49bf3f0acb59c5c%2F2019-12-10%2011.25.36.png?generation=1575991574987840&amp;alt=media)\n\n\n\nVisualizing prediction on testset\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2F2119e9ade9119d02dc75fc1a041d500b%2Fpred0.png?generation=1575965589843154&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff90412af102a2ce9b8ccb82d0bb1902a%2Fpred1.png?generation=1575965608992000&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F274731%2Ff6a7ea678dfe97232dab5ba23ced46d6%2Fpred2.png?generation=1575965634695279&amp;alt=media)",
    "691552": "Hello @niuddd that is indeed very weird . I have few questions . \n1. DId you stop your training at 11th epoch and then run the predictions ? I ask this because Model tends to get very bad overfit at the end if you run longer \n\n2. Did you tweak anything in the post processing part ?\n\n3. How is your validation set comparison ?\n\n4. Most importantly can you please post your validation loss list epoch by epoch , same way that you have posted train loss ?",
    "691862": "Hi, @phoenix9032 , thanks for your reply.\n1. set num_epochs=12, predict with checkpoint of 12th epoch\n2. yes, threshold of logits, tried both 0 and -1 for center-net (efficient-net backbone, will try resnet soon), as I said, LB are between 0.01~0.02\n3. 9% for valid-set, what does \"comparison\" mean?\n4. updated both train&amp;valid loss above\n\nAdditional information: \n1. I managed to copy the evaluation notebook to calculate local CV on valid-set, it's &lt;0.01...\n2. I checked every part of code block by block, found 2 difference (very sure): I use batch size 8 instead of 2 because of 4 gpus, I use torch.optim.Adam instead of AdamW",
    "691891": "What happens if you run exact same code as kernel? It's better to check it's because of code or hardware at first.",
    "692238": "I think the difference lies in the training part because using checkpoint from @phoenix9032 notebook, I managed to reproduce the test-set prediction (99.99% similar). Local cv shows 0.08.",
    "692594": "The problem is subtle to me that I ignored......batch size. I used to set batch size to 8 on 4 gpus, as well as gradient accumulation 2 as default training setting......which turns to be a bad setting for this competition, I changed batch size to 4, it works then. I haven't yet figured out the reason why using a bigger (but still normal) batch size (16 vs 4) could cause drastic failure in training in this task, seems batch size is tunable.",
    "692678": "I have failed miserably with gradient accumulation . This is an open problem for me as well . I thought may be I am doing something wrong"
  },
  "source": "meta"
}