{
  "id": 126585,
  "title": "Correct way to calc mAP",
  "url": "/competitions/pku-autonomous-driving/discussion/126585",
  "author_name": "",
  "post_date": "2020-01-18T14:40:01.587161500Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>When I calculate mAP I just randomly select 20% of the training data but that selection includes data the model was trained on. So if I train a model with 10% validation data should I calculate mAP with only that 10%? Will that reflect the model performance better?</p>",
  "messages": [
    {
      "id": "722402",
      "postDate": "01/18/2020 14:40:01",
      "content": "<p>When I calculate mAP I just randomly select 20% of the training data but that selection includes data the model was trained on. So if I train a model with 10% validation data should I calculate mAP with only that 10%? Will that reflect the model performance better?</p>",
      "rawMarkdown": "When I calculate mAP I just randomly select 20% of the training data but that selection includes data the model was trained on. So if I train a model with 10% validation data should I calculate mAP with only that 10%? Will that reflect the model performance better?",
      "votes": null
    },
    {
      "id": "722425",
      "postDate": "01/18/2020 14:57:37",
      "content": "<p>yes,only that 10%</p>",
      "rawMarkdown": "yes,only that 10%",
      "votes": null
    },
    {
      "id": "722478",
      "postDate": "01/18/2020 16:12:57",
      "content": "<p>Weird, .148 map on that 10% but only .047 on LB</p>",
      "rawMarkdown": "Weird, .148 map on that 10% but only .047 on LB",
      "votes": null
    },
    {
      "id": "722486",
      "postDate": "01/18/2020 16:19:16",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  use 20% instead,,,i see cv calculated with 20% data reflect the model performance better!</p>",
      "rawMarkdown": "greatgamedota  use 20% instead,,,i see cv calculated with 20% data reflect the model performance better!",
      "votes": null
    },
    {
      "id": "722501",
      "postDate": "01/18/2020 16:33:52",
      "content": "<p>Yes, calculating your mAP on data that you used for training has no meaning at all.\nSo only use your holdout set for calculation</p>",
      "rawMarkdown": "Yes, calculating your mAP on data that you used for training has no meaning at all.\nSo only use your holdout set for calculation",
      "votes": null
    },
    {
      "id": "722506",
      "postDate": "01/18/2020 16:36:31",
      "content": "<p>So add another 10% from the data I trained the model on?</p>",
      "rawMarkdown": "So add another 10% from the data I trained the model on?",
      "votes": null
    },
    {
      "id": "722521",
      "postDate": "01/18/2020 16:54:30",
      "content": "<p>no,i mean do 80-20 split...20% for validation and 80% for training then calculate map using that 20% validation data</p>",
      "rawMarkdown": "no,i mean do 80-20 split...20% for validation and 80% for training then calculate map using that 20% validation data",
      "votes": null
    },
    {
      "id": "722524",
      "postDate": "01/18/2020 17:06:36",
      "content": "<p>Thanks, I'll try that!</p>",
      "rawMarkdown": "Thanks, I'll try that!",
      "votes": null
    },
    {
      "id": "723811",
      "postDate": "01/20/2020 13:39:07",
      "content": "<p><a href=\"/mobassir\">@mobassir</a> Retrained models with 80-20 split and I get about the same performance and CV as 90-10.\nOLD .148 -&gt; .047\nNEW .13 -&gt; .048</p>",
      "rawMarkdown": "mobassir Retrained models with 80-20 split and I get about the same performance and CV as 90-10.\nOLD .148 -&gt; .047\nNEW .13 -&gt; .048",
      "votes": null
    },
    {
      "id": "724045",
      "postDate": "01/20/2020 18:20:21",
      "content": "<p>Arent we holding more data without the training ..?\ni used 90:10 CV 0.180+ transforms to .07+ .Highest one i get with this .210 lets see how much does it takes to</p>",
      "rawMarkdown": "Arent we holding more data without the training ..?\ni used 90:10 CV 0.180+ transforms to .07+ .Highest one i get with this .210 lets see how much does it takes to",
      "votes": null
    },
    {
      "id": "724130",
      "postDate": "01/20/2020 20:38:12",
      "content": "<p>always do your submission with a 100% train set but the same hyperparameters as you got from validation</p>",
      "rawMarkdown": "always do your submission with a 100% train set but the same hyperparameters as you got from validation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 722425,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "01/18/2020 14:57:37",
      "content": "<p>yes,only that 10%</p>",
      "votes": null,
      "replies": [
        {
          "id": 722478,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/18/2020 16:12:57",
          "content": "<p>Weird, .148 map on that 10% but only .047 on LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722486,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/18/2020 16:19:16",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  use 20% instead,,,i see cv calculated with 20% data reflect the model performance better!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722506,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/18/2020 16:36:31",
          "content": "<p>So add another 10% from the data I trained the model on?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722521,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/18/2020 16:54:30",
          "content": "<p>no,i mean do 80-20 split...20% for validation and 80% for training then calculate map using that 20% validation data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722524,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/18/2020 17:06:36",
          "content": "<p>Thanks, I'll try that!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723811,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/20/2020 13:39:07",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> Retrained models with 80-20 split and I get about the same performance and CV as 90-10.\nOLD .148 -&gt; .047\nNEW .13 -&gt; .048</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724045,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/20/2020 18:20:21",
          "content": "<p>Arent we holding more data without the training ..?\ni used 90:10 CV 0.180+ transforms to .07+ .Highest one i get with this .210 lets see how much does it takes to</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724130,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "01/20/2020 20:38:12",
          "content": "<p>always do your submission with a 100% train set but the same hyperparameters as you got from validation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 722501,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "01/18/2020 16:33:52",
      "content": "<p>Yes, calculating your mAP on data that you used for training has no meaning at all.\nSo only use your holdout set for calculation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "722402": "When I calculate mAP I just randomly select 20% of the training data but that selection includes data the model was trained on. So if I train a model with 10% validation data should I calculate mAP with only that 10%? Will that reflect the model performance better?",
    "722425": "yes,only that 10%",
    "722478": "Weird, .148 map on that 10% but only .047 on LB",
    "722486": "greatgamedota  use 20% instead,,,i see cv calculated with 20% data reflect the model performance better!",
    "722501": "Yes, calculating your mAP on data that you used for training has no meaning at all.\nSo only use your holdout set for calculation",
    "722506": "So add another 10% from the data I trained the model on?",
    "722521": "no,i mean do 80-20 split...20% for validation and 80% for training then calculate map using that 20% validation data",
    "722524": "Thanks, I'll try that!",
    "723811": "mobassir Retrained models with 80-20 split and I get about the same performance and CV as 90-10.\nOLD .148 -&gt; .047\nNEW .13 -&gt; .048",
    "724045": "Arent we holding more data without the training ..?\ni used 90:10 CV 0.180+ transforms to .07+ .Highest one i get with this .210 lets see how much does it takes to",
    "724130": "always do your submission with a 100% train set but the same hyperparameters as you got from validation"
  },
  "source": "meta"
}