{
  "id": 303766,
  "title": "Why Yolov5s6 gives the best result in this game",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/303766",
  "author_name": "",
  "post_date": "2022-01-29T08:52:24.291404Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, yolov5s6 is the best Public LB model in public notebook.<br>\nYou may go <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> discussion for read it:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638</a></p>\n<p>We can compare with the model structure Yolov5s6, and Yolov5s:</p>\n<h1>YOLOv5s.ymal</h1>\n<blockquote>\n  <p>head:<br>\n   [[-1, 1, Conv, [512, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 6], 1, Concat, [1]],  # cat backbone P4<br>\n    [-1, 3, C3, [512, False]],  # 13</p>\n  <p>[-1, 1, Conv, [256, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 4], 1, Concat, [1]],  # cat backbone P3<br>\n    [-1, 3, C3, [256, False]],  # 17 (P3/8-small)</p>\n  <p>[-1, 1, Conv, [256, 3, 2]],<br>\n    [[-1, 14], 1, Concat, [1]],  # cat head P4<br>\n    [-1, 3, C3, [512, False]],  # 20 (P4/16-medium)</p>\n  <p>[-1, 1, Conv, [512, 3, 2]],<br>\n    [[-1, 10], 1, Concat, [1]],  # cat head P5<br>\n    [-1, 3, C3, [1024, False]],  # 23 (P5/32-large)</p>\n  <p>[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)<br>\n   ]</p>\n</blockquote>\n<p>and</p>\n<h1>YOLOv5s6.ymal</h1>\n<blockquote>\n  <p>head:<br>\n   [[-1, 1, Conv, [768, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 8], 1, Concat, [1]],  # cat backbone P5<br>\n    [-1, 3, C3, [768, False]],  # 15</p>\n  <p>[-1, 1, Conv, [512, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 6], 1, Concat, [1]],  # cat backbone P4<br>\n    [-1, 3, C3, [512, False]],  # 19</p>\n  <p>[-1, 1, Conv, [256, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 4], 1, Concat, [1]],  # cat backbone P3<br>\n    [-1, 3, C3, [256, False]],  # 23 (P3/8-small)</p>\n  <p>[-1, 1, Conv, [256, 3, 2]],<br>\n    [[-1, 20], 1, Concat, [1]],  # cat head P4<br>\n    [-1, 3, C3, [512, False]],  # 26 (P4/16-medium)</p>\n  <p>[-1, 1, Conv, [512, 3, 2]],<br>\n    [[-1, 16], 1, Concat, [1]],  # cat head P5<br>\n    [-1, 3, C3, [768, False]],  # 29 (P5/32-large)</p>\n  <p>[-1, 1, Conv, [768, 3, 2]],<br>\n    [[-1, 12], 1, Concat, [1]],  # cat head P6<br>\n    [-1, 3, C3, [1024, False]],  # 32 (P6/64-xlarge)</p>\n  <p>[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)<br>\n   ]</p>\n</blockquote>\n<p>Then main difference is at the detection line:<br>\n<code>[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)</code><br>\n<code>[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)</code></p>\n<p><strong>The yolov5s6 has one more detection layer</strong>, and which can increase the performance on the tiny object; The COTS video hes lots of tiny starfish, so that the reason Yolov5s6 has the best performance (on the public notebook ATM).</p>",
  "messages": [
    {
      "id": "1668054",
      "postDate": "01/29/2022 08:52:24",
      "content": "<p>First of all, yolov5s6 is the best Public LB model in public notebook.<br>\nYou may go <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> discussion for read it:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638</a></p>\n<p>We can compare with the model structure Yolov5s6, and Yolov5s:</p>\n<h1>YOLOv5s.ymal</h1>\n<blockquote>\n  <p>head:<br>\n   [[-1, 1, Conv, [512, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 6], 1, Concat, [1]],  # cat backbone P4<br>\n    [-1, 3, C3, [512, False]],  # 13</p>\n  <p>[-1, 1, Conv, [256, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 4], 1, Concat, [1]],  # cat backbone P3<br>\n    [-1, 3, C3, [256, False]],  # 17 (P3/8-small)</p>\n  <p>[-1, 1, Conv, [256, 3, 2]],<br>\n    [[-1, 14], 1, Concat, [1]],  # cat head P4<br>\n    [-1, 3, C3, [512, False]],  # 20 (P4/16-medium)</p>\n  <p>[-1, 1, Conv, [512, 3, 2]],<br>\n    [[-1, 10], 1, Concat, [1]],  # cat head P5<br>\n    [-1, 3, C3, [1024, False]],  # 23 (P5/32-large)</p>\n  <p>[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)<br>\n   ]</p>\n</blockquote>\n<p>and</p>\n<h1>YOLOv5s6.ymal</h1>\n<blockquote>\n  <p>head:<br>\n   [[-1, 1, Conv, [768, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 8], 1, Concat, [1]],  # cat backbone P5<br>\n    [-1, 3, C3, [768, False]],  # 15</p>\n  <p>[-1, 1, Conv, [512, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 6], 1, Concat, [1]],  # cat backbone P4<br>\n    [-1, 3, C3, [512, False]],  # 19</p>\n  <p>[-1, 1, Conv, [256, 1, 1]],<br>\n    [-1, 1, nn.Upsample, [None, 2, 'nearest']],<br>\n    [[-1, 4], 1, Concat, [1]],  # cat backbone P3<br>\n    [-1, 3, C3, [256, False]],  # 23 (P3/8-small)</p>\n  <p>[-1, 1, Conv, [256, 3, 2]],<br>\n    [[-1, 20], 1, Concat, [1]],  # cat head P4<br>\n    [-1, 3, C3, [512, False]],  # 26 (P4/16-medium)</p>\n  <p>[-1, 1, Conv, [512, 3, 2]],<br>\n    [[-1, 16], 1, Concat, [1]],  # cat head P5<br>\n    [-1, 3, C3, [768, False]],  # 29 (P5/32-large)</p>\n  <p>[-1, 1, Conv, [768, 3, 2]],<br>\n    [[-1, 12], 1, Concat, [1]],  # cat head P6<br>\n    [-1, 3, C3, [1024, False]],  # 32 (P6/64-xlarge)</p>\n  <p>[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)<br>\n   ]</p>\n</blockquote>\n<p>Then main difference is at the detection line:<br>\n<code>[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)</code><br>\n<code>[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)</code></p>\n<p><strong>The yolov5s6 has one more detection layer</strong>, and which can increase the performance on the tiny object; The COTS video hes lots of tiny starfish, so that the reason Yolov5s6 has the best performance (on the public notebook ATM).</p>",
      "rawMarkdown": "First of all, yolov5s6 is the best Public LB model in public notebook.\nYou may go @steamedsheep discussion for read it:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638\n\n\nWe can compare with the model structure Yolov5s6, and Yolov5s:\n\n# YOLOv5s.ymal\n\n> head:\n>  [[-1, 1, Conv, [512, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 6], 1, Concat, [1]],  # cat backbone P4\n>   [-1, 3, C3, [512, False]],  # 13\n>\n>   [-1, 1, Conv, [256, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 4], 1, Concat, [1]],  # cat backbone P3\n>   [-1, 3, C3, [256, False]],  # 17 (P3/8-small)\n>\n>   [-1, 1, Conv, [256, 3, 2]],\n>   [[-1, 14], 1, Concat, [1]],  # cat head P4\n>   [-1, 3, C3, [512, False]],  # 20 (P4/16-medium)\n>\n>   [-1, 1, Conv, [512, 3, 2]],\n>   [[-1, 10], 1, Concat, [1]],  # cat head P5\n>   [-1, 3, C3, [1024, False]],  # 23 (P5/32-large)\n>\n>   [[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)\n>  ]\n\n\nand\n\n# YOLOv5s6.ymal\n\n> head:\n>  [[-1, 1, Conv, [768, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 8], 1, Concat, [1]],  # cat backbone P5\n>   [-1, 3, C3, [768, False]],  # 15\n>\n>   [-1, 1, Conv, [512, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 6], 1, Concat, [1]],  # cat backbone P4\n>   [-1, 3, C3, [512, False]],  # 19\n>\n>   [-1, 1, Conv, [256, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 4], 1, Concat, [1]],  # cat backbone P3\n>   [-1, 3, C3, [256, False]],  # 23 (P3/8-small)\n>\n>   [-1, 1, Conv, [256, 3, 2]],\n>   [[-1, 20], 1, Concat, [1]],  # cat head P4\n>   [-1, 3, C3, [512, False]],  # 26 (P4/16-medium)\n>\n>   [-1, 1, Conv, [512, 3, 2]],\n>   [[-1, 16], 1, Concat, [1]],  # cat head P5\n>   [-1, 3, C3, [768, False]],  # 29 (P5/32-large)\n>\n>   [-1, 1, Conv, [768, 3, 2]],\n>   [[-1, 12], 1, Concat, [1]],  # cat head P6\n>   [-1, 3, C3, [1024, False]],  # 32 (P6/64-xlarge)\n>\n>   [[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)\n>  ]\n\n\nThen main difference is at the detection line:\n`[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)`\n`[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)`\n\n**The yolov5s6 has one more detection layer**, and which can increase the performance on the tiny object; The COTS video hes lots of tiny starfish, so that the reason Yolov5s6 has the best performance (on the public notebook ATM).",
      "votes": null
    },
    {
      "id": "1668058",
      "postDate": "01/29/2022 09:00:26",
      "content": "<p>In my case, idk why yolov5m6 is performed better than the s6, even the same configs, even though many confirm that the s6 gives them the best performance on LB.</p>",
      "rawMarkdown": "In my case, idk why yolov5m6 is performed better than the s6, even the same configs, even though many confirm that the s6 gives them the best performance on LB.",
      "votes": null
    },
    {
      "id": "1668070",
      "postDate": "01/29/2022 09:17:25",
      "content": "<p>If we compare with yolov5m6 and yolov5m,  it will show the same structure:<br>\n<code>[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)</code><br>\n<code>[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)</code></p>\n<p>I mean with 6 behind is better than original structure.<br>\nm6 is deeper than s6, so it should have a better performance : )</p>",
      "rawMarkdown": "If we compare with yolov5m6 and yolov5m,  it will show the same structure:\n`[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)`\n`[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)`\n\nI mean with 6 behind is better than original structure.\nm6 is deeper than s6, so it should have a better performance : )",
      "votes": null
    },
    {
      "id": "1668099",
      "postDate": "01/29/2022 10:02:22",
      "content": "<p>Please read this Roboflow blog post about v6 release: <a href=\"https://blog.roboflow.com/yolov5-v6-0-is-here/\" target=\"_blank\">https://blog.roboflow.com/yolov5-v6-0-is-here/</a> <br>\nand release notes: <a href=\"https://github.com/ultralytics/yolov5/releases/tag/v6.0\" target=\"_blank\">https://github.com/ultralytics/yolov5/releases/tag/v6.0</a></p>",
      "rawMarkdown": "Please read this Roboflow blog post about v6 release: https://blog.roboflow.com/yolov5-v6-0-is-here/ \nand release notes: https://github.com/ultralytics/yolov5/releases/tag/v6.0",
      "votes": null
    },
    {
      "id": "1668365",
      "postDate": "01/29/2022 15:08:08",
      "content": "<p>'The yolov5s6 has one more detection layer, and which can increase the performance on the tiny object'</p>\n<p>I am not sure about that. I think it is the opposite. Extra detection layer comes from the later blocks which have wider receptive field. Thus, can capture bigger objects. I think boost comes from the increased input size of pretraining.(yolov5s6 trained with input size of 1280). <br>\nI tried to train with yolov5-p2(initialized with yolov5l weights) which have extra detection layer comes from earlier layers but result was worse(maybe extra detection layers couldn't be initialized with yolov5l weights since it is absent)</p>",
      "rawMarkdown": "'The yolov5s6 has one more detection layer, and which can increase the performance on the tiny object'\n\nI am not sure about that. I think it is the opposite. Extra detection layer comes from the later blocks which have wider receptive field. Thus, can capture bigger objects. I think boost comes from the increased input size of pretraining.(yolov5s6 trained with input size of 1280). \nI tried to train with yolov5-p2(initialized with yolov5l weights) which have extra detection layer comes from earlier layers but result was worse(maybe extra detection layers couldn't be initialized with yolov5l weights since it is absent)",
      "votes": null
    },
    {
      "id": "1668666",
      "postDate": "01/29/2022 21:38:04",
      "content": "<p>I think the size is relative. Increasing the receptive field of small targets can also be understood as increasing the receptive field of large targets. My understanding is that it actually increases the depth that the model can detect.<br>\nThanks for reply.</p>",
      "rawMarkdown": "I think the size is relative. Increasing the receptive field of small targets can also be understood as increasing the receptive field of large targets. My understanding is that it actually increases the depth that the model can detect.\nThanks for reply.",
      "votes": null
    },
    {
      "id": "1668667",
      "postDate": "01/29/2022 21:38:28",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1668058,
      "author_name": "locbaop",
      "author_url": "",
      "post_date": "01/29/2022 09:00:26",
      "content": "<p>In my case, idk why yolov5m6 is performed better than the s6, even the same configs, even though many confirm that the s6 gives them the best performance on LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1668070,
          "author_name": "dwchen",
          "author_url": "",
          "post_date": "01/29/2022 09:17:25",
          "content": "<p>If we compare with yolov5m6 and yolov5m,  it will show the same structure:<br>\n<code>[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)</code><br>\n<code>[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)</code></p>\n<p>I mean with 6 behind is better than original structure.<br>\nm6 is deeper than s6, so it should have a better performance : )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1668099,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "01/29/2022 10:02:22",
      "content": "<p>Please read this Roboflow blog post about v6 release: <a href=\"https://blog.roboflow.com/yolov5-v6-0-is-here/\" target=\"_blank\">https://blog.roboflow.com/yolov5-v6-0-is-here/</a> <br>\nand release notes: <a href=\"https://github.com/ultralytics/yolov5/releases/tag/v6.0\" target=\"_blank\">https://github.com/ultralytics/yolov5/releases/tag/v6.0</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1668667,
          "author_name": "dwchen",
          "author_url": "",
          "post_date": "01/29/2022 21:38:28",
          "content": "<p>Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1668365,
      "author_name": "oguzhannefsoglu",
      "author_url": "",
      "post_date": "01/29/2022 15:08:08",
      "content": "<p>'The yolov5s6 has one more detection layer, and which can increase the performance on the tiny object'</p>\n<p>I am not sure about that. I think it is the opposite. Extra detection layer comes from the later blocks which have wider receptive field. Thus, can capture bigger objects. I think boost comes from the increased input size of pretraining.(yolov5s6 trained with input size of 1280). <br>\nI tried to train with yolov5-p2(initialized with yolov5l weights) which have extra detection layer comes from earlier layers but result was worse(maybe extra detection layers couldn't be initialized with yolov5l weights since it is absent)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1668666,
          "author_name": "dwchen",
          "author_url": "",
          "post_date": "01/29/2022 21:38:04",
          "content": "<p>I think the size is relative. Increasing the receptive field of small targets can also be understood as increasing the receptive field of large targets. My understanding is that it actually increases the depth that the model can detect.<br>\nThanks for reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1668054": "First of all, yolov5s6 is the best Public LB model in public notebook.\nYou may go @steamedsheep discussion for read it:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638\n\n\nWe can compare with the model structure Yolov5s6, and Yolov5s:\n\n# YOLOv5s.ymal\n\n> head:\n>  [[-1, 1, Conv, [512, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 6], 1, Concat, [1]],  # cat backbone P4\n>   [-1, 3, C3, [512, False]],  # 13\n>\n>   [-1, 1, Conv, [256, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 4], 1, Concat, [1]],  # cat backbone P3\n>   [-1, 3, C3, [256, False]],  # 17 (P3/8-small)\n>\n>   [-1, 1, Conv, [256, 3, 2]],\n>   [[-1, 14], 1, Concat, [1]],  # cat head P4\n>   [-1, 3, C3, [512, False]],  # 20 (P4/16-medium)\n>\n>   [-1, 1, Conv, [512, 3, 2]],\n>   [[-1, 10], 1, Concat, [1]],  # cat head P5\n>   [-1, 3, C3, [1024, False]],  # 23 (P5/32-large)\n>\n>   [[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)\n>  ]\n\n\nand\n\n# YOLOv5s6.ymal\n\n> head:\n>  [[-1, 1, Conv, [768, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 8], 1, Concat, [1]],  # cat backbone P5\n>   [-1, 3, C3, [768, False]],  # 15\n>\n>   [-1, 1, Conv, [512, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 6], 1, Concat, [1]],  # cat backbone P4\n>   [-1, 3, C3, [512, False]],  # 19\n>\n>   [-1, 1, Conv, [256, 1, 1]],\n>   [-1, 1, nn.Upsample, [None, 2, 'nearest']],\n>   [[-1, 4], 1, Concat, [1]],  # cat backbone P3\n>   [-1, 3, C3, [256, False]],  # 23 (P3/8-small)\n>\n>   [-1, 1, Conv, [256, 3, 2]],\n>   [[-1, 20], 1, Concat, [1]],  # cat head P4\n>   [-1, 3, C3, [512, False]],  # 26 (P4/16-medium)\n>\n>   [-1, 1, Conv, [512, 3, 2]],\n>   [[-1, 16], 1, Concat, [1]],  # cat head P5\n>   [-1, 3, C3, [768, False]],  # 29 (P5/32-large)\n>\n>   [-1, 1, Conv, [768, 3, 2]],\n>   [[-1, 12], 1, Concat, [1]],  # cat head P6\n>   [-1, 3, C3, [1024, False]],  # 32 (P6/64-xlarge)\n>\n>   [[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)\n>  ]\n\n\nThen main difference is at the detection line:\n`[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)`\n`[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)`\n\n**The yolov5s6 has one more detection layer**, and which can increase the performance on the tiny object; The COTS video hes lots of tiny starfish, so that the reason Yolov5s6 has the best performance (on the public notebook ATM).",
    "1668058": "In my case, idk why yolov5m6 is performed better than the s6, even the same configs, even though many confirm that the s6 gives them the best performance on LB.",
    "1668070": "If we compare with yolov5m6 and yolov5m,  it will show the same structure:\n`[[23, 26, 29, 32], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5, P6)`\n`[[17, 20, 23], 1, Detect, [nc, anchors]],  # Detect(P3, P4, P5)`\n\nI mean with 6 behind is better than original structure.\nm6 is deeper than s6, so it should have a better performance : )",
    "1668099": "Please read this Roboflow blog post about v6 release: https://blog.roboflow.com/yolov5-v6-0-is-here/ \nand release notes: https://github.com/ultralytics/yolov5/releases/tag/v6.0",
    "1668365": "'The yolov5s6 has one more detection layer, and which can increase the performance on the tiny object'\n\nI am not sure about that. I think it is the opposite. Extra detection layer comes from the later blocks which have wider receptive field. Thus, can capture bigger objects. I think boost comes from the increased input size of pretraining.(yolov5s6 trained with input size of 1280). \nI tried to train with yolov5-p2(initialized with yolov5l weights) which have extra detection layer comes from earlier layers but result was worse(maybe extra detection layers couldn't be initialized with yolov5l weights since it is absent)",
    "1668666": "I think the size is relative. Increasing the receptive field of small targets can also be understood as increasing the receptive field of large targets. My understanding is that it actually increases the depth that the model can detect.\nThanks for reply.",
    "1668667": "Thanks for sharing."
  },
  "source": "meta"
}