{
  "id": 273256,
  "title": "pytorch vs tensorflow  convergence",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/273256",
  "author_name": "",
  "post_date": "2021-09-20T05:36:22.433188Z",
  "votes": 8,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I tried pytorch and tensorflow in this competition.</p>\n<p>Tensorflow implementation gets good score, but pytorch implementation does not get to tensorflow score.<br>\nPytorch score converge and looks overfitting.</p>\n<p>I think this comes from large minibatch size of tensorflow implementation, so I tried loss accumulation of pytorch. The result doesn't change.</p>\n<p>Which library do you use?</p>",
  "messages": [
    {
      "id": "1517828",
      "postDate": "09/20/2021 05:36:22",
      "content": "<p>I tried pytorch and tensorflow in this competition.</p>\n<p>Tensorflow implementation gets good score, but pytorch implementation does not get to tensorflow score.<br>\nPytorch score converge and looks overfitting.</p>\n<p>I think this comes from large minibatch size of tensorflow implementation, so I tried loss accumulation of pytorch. The result doesn't change.</p>\n<p>Which library do you use?</p>",
      "rawMarkdown": "I tried pytorch and tensorflow in this competition.\n\nTensorflow implementation gets good score, but pytorch implementation does not get to tensorflow score.\nPytorch score converge and looks overfitting.\n\nI think this comes from large minibatch size of tensorflow implementation, so I tried loss accumulation of pytorch. The result doesn't change.\n\nWhich library do you use?",
      "votes": null
    },
    {
      "id": "1517970",
      "postDate": "09/20/2021 08:56:29",
      "content": "<p>I have used both and the main difference is that TF + TPUs seems to run faster. I don't have any numbers to back this up for now, only a hunch. </p>",
      "rawMarkdown": "I have used both and the main difference is that TF + TPUs seems to run faster. I don't have any numbers to back this up for now, only a hunch.",
      "votes": null
    },
    {
      "id": "1518272",
      "postDate": "09/20/2021 14:28:46",
      "content": "<p>There are many differences between TF and Pytorch, like built-in layers initialization.</p>",
      "rawMarkdown": "There are many differences between TF and Pytorch, like built-in layers initialization.",
      "votes": null
    },
    {
      "id": "1518339",
      "postDate": "09/20/2021 15:18:16",
      "content": "<p>This is correct. I remember a whole discussion in a previous competition about why WaveNet implementations between TF and PyTorch had different results and it came down to initialization.</p>",
      "rawMarkdown": "This is correct. I remember a whole discussion in a previous competition about why WaveNet implementations between TF and PyTorch had different results and it came down to initialization.",
      "votes": null
    },
    {
      "id": "1518691",
      "postDate": "09/20/2021 23:58:01",
      "content": "<p>can you give an example value of what is good tf score and bad pytorch score?</p>\n<p>loss accumulation is not good for BN.</p>",
      "rawMarkdown": "can you give an example value of what is good tf score and bad pytorch score?\n\n loss accumulation is not good for BN.",
      "votes": null
    },
    {
      "id": "1518984",
      "postDate": "09/21/2021 08:33:14",
      "content": "<p>I didn't know that, why it's bad for bn?</p>",
      "rawMarkdown": "I didn't know that, why it's bad for bn?",
      "votes": null
    },
    {
      "id": "1519044",
      "postDate": "09/21/2021 09:25:24",
      "content": "<p><a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> - Gradient accumulation works by accumulating gradients over multiple batches. But the gradients (and therefore BN stats) are still calculated per single batch. So if your BS is too small to calculate reliable BN statistics, gradient accumulation will not help.</p>\n<p>Models without BN (e.g. transformers that use layer norm) do benefit from gradient accumulation</p>",
      "rawMarkdown": "bakeryproducts - Gradient accumulation works by accumulating gradients over multiple batches. But the gradients (and therefore BN stats) are still calculated per single batch. So if your BS is too small to calculate reliable BN statistics, gradient accumulation will not help.\n\nModels without BN (e.g. transformers that use layer norm) do benefit from gradient accumulation",
      "votes": null
    },
    {
      "id": "1519254",
      "postDate": "09/21/2021 13:08:37",
      "content": "<p>Thank you.<br>\nI sometimes hear of there are implementation differences beween tf and pytorch in other models.</p>\n<p>I see that pytorch imlementation converged faster and looks overfitting in about 4 epochs. I doubt this really is from the difference of layer initialization.<br>\nI used pretrained model when learning in both tf and pytorch.</p>",
      "rawMarkdown": "Thank you.\nI sometimes hear of there are implementation differences beween tf and pytorch in other models.\n\nI see that pytorch imlementation converged faster and looks overfitting in about 4 epochs. I doubt this really is from the difference of layer initialization.\nI used pretrained model when learning in both tf and pytorch.",
      "votes": null
    },
    {
      "id": "1519258",
      "postDate": "09/21/2021 13:12:55",
      "content": "<p>Thanks.</p>\n<blockquote>\n  <p>loss accumulation is not good for BN.</p>\n</blockquote>\n<p>Yeah , I don't want to use loss accumulation actively, but the one of implemenation difference is batch size in this time.<br>\nTF implementation uses TPU, so it realizes large batch size. I think TF large batch size using some TPU clusters  is theoretically equal to loss accumulation, but is this wrong?</p>",
      "rawMarkdown": "Thanks.\n> loss accumulation is not good for BN.\n\nYeah , I don't want to use loss accumulation actively, but the one of implemenation difference is batch size in this time.\nTF implementation uses TPU, so it realizes large batch size. I think TF large batch size using some TPU clusters  is theoretically equal to loss accumulation, but is this wrong?",
      "votes": null
    },
    {
      "id": "1519505",
      "postDate": "09/21/2021 17:11:32",
      "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> I dont know about tf (should be the same as pytorch), but bn in pytorch keeps running estimates of its computed mean and variance. Default momentum on running stats is low, but you can turn it up obviously. My point is that bn stats do not calculated per single batch</p>",
      "rawMarkdown": "anjum48 I dont know about tf (should be the same as pytorch), but bn in pytorch keeps running estimates of its computed mean and variance. Default momentum on running stats is low, but you can turn it up obviously. My point is that bn stats do not calculated per single batch",
      "votes": null
    },
    {
      "id": "1519541",
      "postDate": "09/21/2021 18:09:27",
      "content": "<blockquote>\n  <p>My point is that bn stats do not calculated per single batch</p>\n</blockquote>\n<p>If the batch size is too small (e.g. &lt; 4) then the update to the rolling variance will be noisy. My point was that gradient accumulation does not fix this</p>",
      "rawMarkdown": ">My point is that bn stats do not calculated per single batch\n\nIf the batch size is too small (e.g. < 4) then the update to the rolling variance will be noisy. My point was that gradient accumulation does not fix this",
      "votes": null
    },
    {
      "id": "1519555",
      "postDate": "09/21/2021 18:29:14",
      "content": "<blockquote>\n  <p>I doubt this really is from the difference of layer initialization.</p>\n</blockquote>\n<p>If you say so.</p>",
      "rawMarkdown": "> I doubt this really is from the difference of layer initialization.\n\nIf you say so.",
      "votes": null
    },
    {
      "id": "1519579",
      "postDate": "09/21/2021 18:54:37",
      "content": "<p>According to public notebooks tf also overfits after 4 epochs.</p>",
      "rawMarkdown": "According to public notebooks tf also overfits after 4 epochs.",
      "votes": null
    },
    {
      "id": "1519752",
      "postDate": "09/21/2021 21:34:47",
      "content": "<p>Ok, but it started with g.a. is bad for bn, not g.a can fix bn. There is nothing really to fix as if single update is noisy just add more momentum to bn. I dont really see how grad accumulation relate to bn to became \"worse\"</p>",
      "rawMarkdown": "Ok, but it started with g.a. is bad for bn, not g.a can fix bn. There is nothing really to fix as if single update is noisy just add more momentum to bn. I dont really see how grad accumulation relate to bn to became \"worse\"",
      "votes": null
    },
    {
      "id": "1519774",
      "postDate": "09/21/2021 22:15:59",
      "content": "<p>Thank you.<br>\nI will try to improve score using both tf and pytorch.</p>",
      "rawMarkdown": "Thank you.\nI will try to improve score using both tf and pytorch.",
      "votes": null
    },
    {
      "id": "1519804",
      "postDate": "09/21/2021 23:42:26",
      "content": "<p>you can refer to the papers that proposed alternative normization to batch norm and syn batch norm (synchronization across gpus).</p>\n<p>they should have experimental results showing how bn performs across the different batch sizes.<br>\ni recalled reading some papers that did experiments on gradient accumulations too.</p>",
      "rawMarkdown": "you can refer to the papers that proposed alternative normization to batch norm and syn batch norm (synchronization across gpus).\n\nthey should have experimental results showing how bn performs across the different batch sizes.\ni recalled reading some papers that did experiments on gradient accumulations too.",
      "votes": null
    },
    {
      "id": "1528238",
      "postDate": "09/29/2021 13:47:22",
      "content": "<p><a href=\"https://openaccess.thecvf.com/content_CVPR_2020/papers/Singh_Filter_Response_Normalization_Layer_Eliminating_Batch_Dependence_in_the_Training_CVPR_2020_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPR_2020/papers/Singh_Filter_Response_Normalization_Layer_Eliminating_Batch_Dependence_in_the_Training_CVPR_2020_paper.pdf</a></p>\n<p>this paper has a graph on accuracy vs BN batch size for multiple GPU (equivalent to gradient accumulation) </p>",
      "rawMarkdown": "https://openaccess.thecvf.com/content_CVPR_2020/papers/Singh_Filter_Response_Normalization_Layer_Eliminating_Batch_Dependence_in_the_Training_CVPR_2020_paper.pdf\n\nthis paper has a graph on accuracy vs BN batch size for multiple GPU (equivalent to gradient accumulation)",
      "votes": null
    },
    {
      "id": "1559743",
      "postDate": "10/27/2021 07:11:05",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1517970,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "09/20/2021 08:56:29",
      "content": "<p>I have used both and the main difference is that TF + TPUs seems to run faster. I don't have any numbers to back this up for now, only a hunch. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1518272,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/20/2021 14:28:46",
      "content": "<p>There are many differences between TF and Pytorch, like built-in layers initialization.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1518339,
          "author_name": "marktenenholtz",
          "author_url": "",
          "post_date": "09/20/2021 15:18:16",
          "content": "<p>This is correct. I remember a whole discussion in a previous competition about why WaveNet implementations between TF and PyTorch had different results and it came down to initialization.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519254,
          "author_name": "yoshito",
          "author_url": "",
          "post_date": "09/21/2021 13:08:37",
          "content": "<p>Thank you.<br>\nI sometimes hear of there are implementation differences beween tf and pytorch in other models.</p>\n<p>I see that pytorch imlementation converged faster and looks overfitting in about 4 epochs. I doubt this really is from the difference of layer initialization.<br>\nI used pretrained model when learning in both tf and pytorch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519555,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/21/2021 18:29:14",
          "content": "<blockquote>\n  <p>I doubt this really is from the difference of layer initialization.</p>\n</blockquote>\n<p>If you say so.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519579,
          "author_name": "hannes82",
          "author_url": "",
          "post_date": "09/21/2021 18:54:37",
          "content": "<p>According to public notebooks tf also overfits after 4 epochs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519774,
          "author_name": "yoshito",
          "author_url": "",
          "post_date": "09/21/2021 22:15:59",
          "content": "<p>Thank you.<br>\nI will try to improve score using both tf and pytorch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1518691,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/20/2021 23:58:01",
      "content": "<p>can you give an example value of what is good tf score and bad pytorch score?</p>\n<p>loss accumulation is not good for BN.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1518984,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "09/21/2021 08:33:14",
          "content": "<p>I didn't know that, why it's bad for bn?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519044,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "09/21/2021 09:25:24",
          "content": "<p><a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> - Gradient accumulation works by accumulating gradients over multiple batches. But the gradients (and therefore BN stats) are still calculated per single batch. So if your BS is too small to calculate reliable BN statistics, gradient accumulation will not help.</p>\n<p>Models without BN (e.g. transformers that use layer norm) do benefit from gradient accumulation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519258,
          "author_name": "yoshito",
          "author_url": "",
          "post_date": "09/21/2021 13:12:55",
          "content": "<p>Thanks.</p>\n<blockquote>\n  <p>loss accumulation is not good for BN.</p>\n</blockquote>\n<p>Yeah , I don't want to use loss accumulation actively, but the one of implemenation difference is batch size in this time.<br>\nTF implementation uses TPU, so it realizes large batch size. I think TF large batch size using some TPU clusters  is theoretically equal to loss accumulation, but is this wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519505,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "09/21/2021 17:11:32",
          "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> I dont know about tf (should be the same as pytorch), but bn in pytorch keeps running estimates of its computed mean and variance. Default momentum on running stats is low, but you can turn it up obviously. My point is that bn stats do not calculated per single batch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519541,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "09/21/2021 18:09:27",
          "content": "<blockquote>\n  <p>My point is that bn stats do not calculated per single batch</p>\n</blockquote>\n<p>If the batch size is too small (e.g. &lt; 4) then the update to the rolling variance will be noisy. My point was that gradient accumulation does not fix this</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519752,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "09/21/2021 21:34:47",
          "content": "<p>Ok, but it started with g.a. is bad for bn, not g.a can fix bn. There is nothing really to fix as if single update is noisy just add more momentum to bn. I dont really see how grad accumulation relate to bn to became \"worse\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1519804,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/21/2021 23:42:26",
          "content": "<p>you can refer to the papers that proposed alternative normization to batch norm and syn batch norm (synchronization across gpus).</p>\n<p>they should have experimental results showing how bn performs across the different batch sizes.<br>\ni recalled reading some papers that did experiments on gradient accumulations too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1528238,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/29/2021 13:47:22",
      "content": "<p><a href=\"https://openaccess.thecvf.com/content_CVPR_2020/papers/Singh_Filter_Response_Normalization_Layer_Eliminating_Batch_Dependence_in_the_Training_CVPR_2020_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPR_2020/papers/Singh_Filter_Response_Normalization_Layer_Eliminating_Batch_Dependence_in_the_Training_CVPR_2020_paper.pdf</a></p>\n<p>this paper has a graph on accuracy vs BN batch size for multiple GPU (equivalent to gradient accumulation) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1559743,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 07:11:05",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1517828": "I tried pytorch and tensorflow in this competition.\n\nTensorflow implementation gets good score, but pytorch implementation does not get to tensorflow score.\nPytorch score converge and looks overfitting.\n\nI think this comes from large minibatch size of tensorflow implementation, so I tried loss accumulation of pytorch. The result doesn't change.\n\nWhich library do you use?",
    "1517970": "I have used both and the main difference is that TF + TPUs seems to run faster. I don't have any numbers to back this up for now, only a hunch.",
    "1518272": "There are many differences between TF and Pytorch, like built-in layers initialization.",
    "1518339": "This is correct. I remember a whole discussion in a previous competition about why WaveNet implementations between TF and PyTorch had different results and it came down to initialization.",
    "1518691": "can you give an example value of what is good tf score and bad pytorch score?\n\n loss accumulation is not good for BN.",
    "1518984": "I didn't know that, why it's bad for bn?",
    "1519044": "bakeryproducts - Gradient accumulation works by accumulating gradients over multiple batches. But the gradients (and therefore BN stats) are still calculated per single batch. So if your BS is too small to calculate reliable BN statistics, gradient accumulation will not help.\n\nModels without BN (e.g. transformers that use layer norm) do benefit from gradient accumulation",
    "1519254": "Thank you.\nI sometimes hear of there are implementation differences beween tf and pytorch in other models.\n\nI see that pytorch imlementation converged faster and looks overfitting in about 4 epochs. I doubt this really is from the difference of layer initialization.\nI used pretrained model when learning in both tf and pytorch.",
    "1519258": "Thanks.\n> loss accumulation is not good for BN.\n\nYeah , I don't want to use loss accumulation actively, but the one of implemenation difference is batch size in this time.\nTF implementation uses TPU, so it realizes large batch size. I think TF large batch size using some TPU clusters  is theoretically equal to loss accumulation, but is this wrong?",
    "1519505": "anjum48 I dont know about tf (should be the same as pytorch), but bn in pytorch keeps running estimates of its computed mean and variance. Default momentum on running stats is low, but you can turn it up obviously. My point is that bn stats do not calculated per single batch",
    "1519541": ">My point is that bn stats do not calculated per single batch\n\nIf the batch size is too small (e.g. < 4) then the update to the rolling variance will be noisy. My point was that gradient accumulation does not fix this",
    "1519555": "> I doubt this really is from the difference of layer initialization.\n\nIf you say so.",
    "1519579": "According to public notebooks tf also overfits after 4 epochs.",
    "1519752": "Ok, but it started with g.a. is bad for bn, not g.a can fix bn. There is nothing really to fix as if single update is noisy just add more momentum to bn. I dont really see how grad accumulation relate to bn to became \"worse\"",
    "1519774": "Thank you.\nI will try to improve score using both tf and pytorch.",
    "1519804": "you can refer to the papers that proposed alternative normization to batch norm and syn batch norm (synchronization across gpus).\n\nthey should have experimental results showing how bn performs across the different batch sizes.\ni recalled reading some papers that did experiments on gradient accumulations too.",
    "1528238": "https://openaccess.thecvf.com/content_CVPR_2020/papers/Singh_Filter_Response_Normalization_Layer_Eliminating_Batch_Dependence_in_the_Training_CVPR_2020_paper.pdf\n\nthis paper has a graph on accuracy vs BN batch size for multiple GPU (equivalent to gradient accumulation)",
    "1559743": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}