{
  "id": 254574,
  "title": "CV score-vs-LB score",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/254574",
  "author_name": "",
  "post_date": "2021-07-22T14:15:53.097362500Z",
  "votes": 18,
  "comment_count": 10,
  "views": 0,
  "content": "<h2>What is the relationship between your score on the training data and your score on LB?</h2>\n<p>Our results are as follows.</p>\n<ul>\n<li>train_score 2.465</li>\n<li>LB 3.481</li>\n</ul>\n<h4>The train_score is calculated in the following way</h4>\n<pre><code>from vincenty import vincenty\n\ndef vincenty_meter(r, lat='latDeg', lng='lngDeg', tlat='ground_truth_latDeg', tlng='ground_truth_lngDeg'):\n    return vincenty((r[lat], r[lng]), (r[tlat], r[tlng])) * 1000\n\ndef calc_score(input_df: pd.DataFrame):\n    output_df = input_df.copy()\n    output_df['meter'] = input_df[['latDeg','lngDeg','ground_truth_latDeg','ground_truth_lngDeg']].apply(vincenty_meter, axis=1)\n    meter_score = output_df['meter'].mean()\n\n    scores = []\n    for phone in output_df['phone'].unique():\n        p_50 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 50)\n        p_95 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 95)\n        scores.append(p_50)\n        scores.append(p_95)\n\n    score = sum(scores) / len(scores)\n    return score\n</code></pre>",
  "messages": [
    {
      "id": "1396825",
      "postDate": "07/22/2021 14:15:53",
      "content": "<h2>What is the relationship between your score on the training data and your score on LB?</h2>\n<p>Our results are as follows.</p>\n<ul>\n<li>train_score 2.465</li>\n<li>LB 3.481</li>\n</ul>\n<h4>The train_score is calculated in the following way</h4>\n<pre><code>from vincenty import vincenty\n\ndef vincenty_meter(r, lat='latDeg', lng='lngDeg', tlat='ground_truth_latDeg', tlng='ground_truth_lngDeg'):\n    return vincenty((r[lat], r[lng]), (r[tlat], r[tlng])) * 1000\n\ndef calc_score(input_df: pd.DataFrame):\n    output_df = input_df.copy()\n    output_df['meter'] = input_df[['latDeg','lngDeg','ground_truth_latDeg','ground_truth_lngDeg']].apply(vincenty_meter, axis=1)\n    meter_score = output_df['meter'].mean()\n\n    scores = []\n    for phone in output_df['phone'].unique():\n        p_50 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 50)\n        p_95 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 95)\n        scores.append(p_50)\n        scores.append(p_95)\n\n    score = sum(scores) / len(scores)\n    return score\n</code></pre>",
      "rawMarkdown": "## What is the relationship between your score on the training data and your score on LB?\nOur results are as follows.\n\n- train_score 2.465\n- LB 3.481\n\n\n\n#### The train_score is calculated in the following way\n\n```\nfrom vincenty import vincenty\n\ndef vincenty_meter(r, lat='latDeg', lng='lngDeg', tlat='ground_truth_latDeg', tlng='ground_truth_lngDeg'):\n    return vincenty((r[lat], r[lng]), (r[tlat], r[tlng])) * 1000\n\ndef calc_score(input_df: pd.DataFrame):\n    output_df = input_df.copy()\n    output_df['meter'] = input_df[['latDeg','lngDeg','ground_truth_latDeg','ground_truth_lngDeg']].apply(vincenty_meter, axis=1)\n    meter_score = output_df['meter'].mean()\n\n    scores = []\n    for phone in output_df['phone'].unique():\n        p_50 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 50)\n        p_95 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 95)\n        scores.append(p_50)\n        scores.append(p_95)\n\n    score = sum(scores) / len(scores)\n    return score\n```",
      "votes": null
    },
    {
      "id": "1398216",
      "postDate": "07/23/2021 21:21:47",
      "content": "<p>train : 3.080 / LB : 4.085</p>",
      "rawMarkdown": "train : 3.080 / LB : 4.085",
      "votes": null
    },
    {
      "id": "1398254",
      "postDate": "07/23/2021 23:13:50",
      "content": "<p>From what I have observed, the difference after crossing the 5.00 is always close to 1</p>\n<p>upd. LB is ~4 and CV is ~3</p>",
      "rawMarkdown": "From what I have observed, the difference after crossing the 5.00 is always close to 1\n\nupd. LB is ~4 and CV is ~3",
      "votes": null
    },
    {
      "id": "1398648",
      "postDate": "07/24/2021 11:08:57",
      "content": "<p>train_score(CV): 3.262<br>\nLB: 4.821</p>",
      "rawMarkdown": "train_score(CV): 3.262\nLB: 4.821",
      "votes": null
    },
    {
      "id": "1398649",
      "postDate": "07/24/2021 11:10:34",
      "content": "<pre><code>CV : 3.770 / LB : 5.386\nCV : 3.295 / LB : 5.147\nCV : 3.186 / LB : 4.919\nCV : 2.928 / LB : 4.570\n</code></pre>\n<p>Unlike everyone else, the CV/LB gap increases as the score improves… 😣<br>\nDoes anyone have any idea what's going on?</p>",
      "rawMarkdown": "```\nCV : 3.770 / LB : 5.386\nCV : 3.295 / LB : 5.147\nCV : 3.186 / LB : 4.919\nCV : 2.928 / LB : 4.570\n```\n\nUnlike everyone else, the CV/LB gap increases as the score improves... 😣\nDoes anyone have any idea what's going on?",
      "votes": null
    },
    {
      "id": "1398717",
      "postDate": "07/24/2021 12:31:37",
      "content": "<p>CV: 2.999 / LB: 4.169</p>",
      "rawMarkdown": "CV: 2.999 / LB: 4.169",
      "votes": null
    },
    {
      "id": "1398761",
      "postDate": "07/24/2021 13:19:54",
      "content": "<p>These values are calculated by haversine distance, but the difference should be negligible.</p>\n<p>train : 2.341<br>\nLB    : 3.772</p>\n<p>My model shouldn't overfit to the public LB, but it is highly overfitted to the training data?</p>",
      "rawMarkdown": "These values are calculated by haversine distance, but the difference should be negligible.\n\ntrain : 2.341\nLB    : 3.772\n\nMy model shouldn't overfit to the public LB, but it is highly overfitted to the training data?",
      "votes": null
    },
    {
      "id": "1399649",
      "postDate": "07/25/2021 14:42:39",
      "content": "<p>CV: 3.049 / LB: 4.519</p>",
      "rawMarkdown": "CV: 3.049 / LB: 4.519",
      "votes": null
    },
    {
      "id": "1399740",
      "postDate": "07/25/2021 16:23:07",
      "content": "<p>CV: 3.467 / LB: 4.735</p>",
      "rawMarkdown": "CV: 3.467 / LB: 4.735",
      "votes": null
    },
    {
      "id": "1400157",
      "postDate": "07/26/2021 05:35:52",
      "content": "<p>CV: 3.410 / LB: 4.972</p>",
      "rawMarkdown": "CV: 3.410 / LB: 4.972",
      "votes": null
    },
    {
      "id": "1449647",
      "postDate": "08/04/2021 23:40:22",
      "content": "<p>CV 5.11 vs LB 6.530</p>",
      "rawMarkdown": "CV 5.11 vs LB 6.530",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1398216,
      "author_name": "t88take",
      "author_url": "",
      "post_date": "07/23/2021 21:21:47",
      "content": "<p>train : 3.080 / LB : 4.085</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1398254,
      "author_name": "avtobusbratiev",
      "author_url": "",
      "post_date": "07/23/2021 23:13:50",
      "content": "<p>From what I have observed, the difference after crossing the 5.00 is always close to 1</p>\n<p>upd. LB is ~4 and CV is ~3</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1398648,
      "author_name": "takaito",
      "author_url": "",
      "post_date": "07/24/2021 11:08:57",
      "content": "<p>train_score(CV): 3.262<br>\nLB: 4.821</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1398649,
      "author_name": "sakami",
      "author_url": "",
      "post_date": "07/24/2021 11:10:34",
      "content": "<pre><code>CV : 3.770 / LB : 5.386\nCV : 3.295 / LB : 5.147\nCV : 3.186 / LB : 4.919\nCV : 2.928 / LB : 4.570\n</code></pre>\n<p>Unlike everyone else, the CV/LB gap increases as the score improves… 😣<br>\nDoes anyone have any idea what's going on?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1398717,
      "author_name": "shimacos",
      "author_url": "",
      "post_date": "07/24/2021 12:31:37",
      "content": "<p>CV: 2.999 / LB: 4.169</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1398761,
      "author_name": "saitodevel01",
      "author_url": "",
      "post_date": "07/24/2021 13:19:54",
      "content": "<p>These values are calculated by haversine distance, but the difference should be negligible.</p>\n<p>train : 2.341<br>\nLB    : 3.772</p>\n<p>My model shouldn't overfit to the public LB, but it is highly overfitted to the training data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1399649,
      "author_name": "kuto0633",
      "author_url": "",
      "post_date": "07/25/2021 14:42:39",
      "content": "<p>CV: 3.049 / LB: 4.519</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1399740,
      "author_name": "hashimotoryuichi",
      "author_url": "",
      "post_date": "07/25/2021 16:23:07",
      "content": "<p>CV: 3.467 / LB: 4.735</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1400157,
      "author_name": "shu421",
      "author_url": "",
      "post_date": "07/26/2021 05:35:52",
      "content": "<p>CV: 3.410 / LB: 4.972</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1449647,
      "author_name": "dmitrykalashnikov",
      "author_url": "",
      "post_date": "08/04/2021 23:40:22",
      "content": "<p>CV 5.11 vs LB 6.530</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1396825": "## What is the relationship between your score on the training data and your score on LB?\nOur results are as follows.\n\n- train_score 2.465\n- LB 3.481\n\n\n\n#### The train_score is calculated in the following way\n\n```\nfrom vincenty import vincenty\n\ndef vincenty_meter(r, lat='latDeg', lng='lngDeg', tlat='ground_truth_latDeg', tlng='ground_truth_lngDeg'):\n    return vincenty((r[lat], r[lng]), (r[tlat], r[tlng])) * 1000\n\ndef calc_score(input_df: pd.DataFrame):\n    output_df = input_df.copy()\n    output_df['meter'] = input_df[['latDeg','lngDeg','ground_truth_latDeg','ground_truth_lngDeg']].apply(vincenty_meter, axis=1)\n    meter_score = output_df['meter'].mean()\n\n    scores = []\n    for phone in output_df['phone'].unique():\n        p_50 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 50)\n        p_95 = np.percentile(output_df.loc[output_df['phone']==phone, 'meter'], 95)\n        scores.append(p_50)\n        scores.append(p_95)\n\n    score = sum(scores) / len(scores)\n    return score\n```",
    "1398216": "train : 3.080 / LB : 4.085",
    "1398254": "From what I have observed, the difference after crossing the 5.00 is always close to 1\n\nupd. LB is ~4 and CV is ~3",
    "1398648": "train_score(CV): 3.262\nLB: 4.821",
    "1398649": "```\nCV : 3.770 / LB : 5.386\nCV : 3.295 / LB : 5.147\nCV : 3.186 / LB : 4.919\nCV : 2.928 / LB : 4.570\n```\n\nUnlike everyone else, the CV/LB gap increases as the score improves... 😣\nDoes anyone have any idea what's going on?",
    "1398717": "CV: 2.999 / LB: 4.169",
    "1398761": "These values are calculated by haversine distance, but the difference should be negligible.\n\ntrain : 2.341\nLB    : 3.772\n\nMy model shouldn't overfit to the public LB, but it is highly overfitted to the training data?",
    "1399649": "CV: 3.049 / LB: 4.519",
    "1399740": "CV: 3.467 / LB: 4.735",
    "1400157": "CV: 3.410 / LB: 4.972",
    "1449647": "CV 5.11 vs LB 6.530"
  },
  "source": "meta"
}