{
  "id": 78221,
  "title": "Help needed. Does my code of local test set make sense?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78221",
  "author_name": "",
  "post_date": "2019-01-21T10:42:06.730384600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>My local test set score is much different from LB score. Here is my code for local test set:</p>\n\n<pre><code>train_raw_df = pd.read_csv(train_path)\nmsk = np.random.rand(len(train_raw_df)) &amp;lt; 0.95\ntrain_df = train_raw_df[msk].copy().reset_index(drop=True)\ntest_df = train_raw_df[~msk].copy().reset_index(drop=True)\ntest_y = test_df['target'].values\ntest_df = test_df.drop(columns=['target'])\ndel train_raw_df\n\n# train the model .... no code use test_y\n\nprediction = test_preds &amp;gt; search_result['threshold']\nfinall_res = sklearn.metrics.f1_score(test_y, prediction)\n</code></pre>\n\n<p>The <code>finall_res</code> is my local test score. For LB, to get the same model, I didn't even use full train set to keep <code>train_df</code> set the same (I have set seed):</p>\n\n<pre><code>train_raw_df = pd.read_csv(train_path)\nmsk = np.random.rand(len(train_raw_df)) &amp;lt; 0.95\ntrain_df = train_raw_df[msk].copy().reset_index(drop=True)\ndel train_raw_df\ntest_df = pd.read_csv(test_path)\n\n# train the model .... \n# I get exactly the same loss in every epoch, so I should get the same model.\n\nsub = test_df[['qid']].copy()\nsub['prediction'] = (test_preds &amp;gt; search_result['threshold']).astype('int')\nsub.to_csv(\"submission.csv\", index=False)\n</code></pre>\n\n<p>However, the difference between local and LB still can be 0.001~0.015. It looks like that the LB score is always about 0.694-0.699, no matter what my local score is.  I also tried other seed and the variation of local test set score is about 0.001-0.003.\nShould I trust my local test set score?  Is my code correct?  </p>",
  "messages": [
    {
      "id": "459175",
      "postDate": "01/21/2019 10:42:06",
      "content": "<p>My local test set score is much different from LB score. Here is my code for local test set:</p>\n\n<pre><code>train_raw_df = pd.read_csv(train_path)\nmsk = np.random.rand(len(train_raw_df)) &amp;lt; 0.95\ntrain_df = train_raw_df[msk].copy().reset_index(drop=True)\ntest_df = train_raw_df[~msk].copy().reset_index(drop=True)\ntest_y = test_df['target'].values\ntest_df = test_df.drop(columns=['target'])\ndel train_raw_df\n\n# train the model .... no code use test_y\n\nprediction = test_preds &amp;gt; search_result['threshold']\nfinall_res = sklearn.metrics.f1_score(test_y, prediction)\n</code></pre>\n\n<p>The <code>finall_res</code> is my local test score. For LB, to get the same model, I didn't even use full train set to keep <code>train_df</code> set the same (I have set seed):</p>\n\n<pre><code>train_raw_df = pd.read_csv(train_path)\nmsk = np.random.rand(len(train_raw_df)) &amp;lt; 0.95\ntrain_df = train_raw_df[msk].copy().reset_index(drop=True)\ndel train_raw_df\ntest_df = pd.read_csv(test_path)\n\n# train the model .... \n# I get exactly the same loss in every epoch, so I should get the same model.\n\nsub = test_df[['qid']].copy()\nsub['prediction'] = (test_preds &amp;gt; search_result['threshold']).astype('int')\nsub.to_csv(\"submission.csv\", index=False)\n</code></pre>\n\n<p>However, the difference between local and LB still can be 0.001~0.015. It looks like that the LB score is always about 0.694-0.699, no matter what my local score is.  I also tried other seed and the variation of local test set score is about 0.001-0.003.\nShould I trust my local test set score?  Is my code correct?  </p>",
      "rawMarkdown": "My local test set score is much different from LB score. Here is my code for local test set:\n\n    train_raw_df = pd.read_csv(train_path)\n    msk = np.random.rand(len(train_raw_df)) &lt; 0.95\n    train_df = train_raw_df[msk].copy().reset_index(drop=True)\n    test_df = train_raw_df[~msk].copy().reset_index(drop=True)\n    test_y = test_df['target'].values\n    test_df = test_df.drop(columns=['target'])\n    del train_raw_df\n    \n    # train the model .... no code use test_y\n    \n    prediction = test_preds &gt; search_result['threshold']\n    finall_res = sklearn.metrics.f1_score(test_y, prediction)\n    \nThe `finall_res` is my local test score. For LB, to get the same model, I didn't even use full train set to keep `train_df` set the same (I have set seed):\n\n    train_raw_df = pd.read_csv(train_path)\n    msk = np.random.rand(len(train_raw_df)) &lt; 0.95\n    train_df = train_raw_df[msk].copy().reset_index(drop=True)\n    del train_raw_df\n    test_df = pd.read_csv(test_path)\n\n    # train the model .... \n    # I get exactly the same loss in every epoch, so I should get the same model.\n    \n    sub = test_df[['qid']].copy()\n    sub['prediction'] = (test_preds &gt; search_result['threshold']).astype('int')\n    sub.to_csv(\"submission.csv\", index=False)\n\nHowever, the difference between local and LB still can be 0.001~0.015. It looks like that the LB score is always about 0.694-0.699, no matter what my local score is.  I also tried other seed and the variation of local test set score is about 0.001-0.003.\nShould I trust my local test set score?  Is my code correct?",
      "votes": null
    },
    {
      "id": "459594",
      "postDate": "01/22/2019 03:49:44",
      "content": "<p>I got stuck with the problem about two weeks. So frustrated.</p>",
      "rawMarkdown": "I got stuck with the problem about two weeks. So frustrated.",
      "votes": null
    },
    {
      "id": "459624",
      "postDate": "01/22/2019 05:11:42",
      "content": "<p>I can't find any problems in your codes.Seems all should work.Maybe you can set a larger test set or just use kfold cv to evaluate your model.</p>",
      "rawMarkdown": "I can't find any problems in your codes.Seems all should work.Maybe you can set a larger test set or just use kfold cv to evaluate your model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 459594,
      "author_name": "wxytalent",
      "author_url": "",
      "post_date": "01/22/2019 03:49:44",
      "content": "<p>I got stuck with the problem about two weeks. So frustrated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459624,
      "author_name": "noxuslol",
      "author_url": "",
      "post_date": "01/22/2019 05:11:42",
      "content": "<p>I can't find any problems in your codes.Seems all should work.Maybe you can set a larger test set or just use kfold cv to evaluate your model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "459175": "My local test set score is much different from LB score. Here is my code for local test set:\n\n    train_raw_df = pd.read_csv(train_path)\n    msk = np.random.rand(len(train_raw_df)) &lt; 0.95\n    train_df = train_raw_df[msk].copy().reset_index(drop=True)\n    test_df = train_raw_df[~msk].copy().reset_index(drop=True)\n    test_y = test_df['target'].values\n    test_df = test_df.drop(columns=['target'])\n    del train_raw_df\n    \n    # train the model .... no code use test_y\n    \n    prediction = test_preds &gt; search_result['threshold']\n    finall_res = sklearn.metrics.f1_score(test_y, prediction)\n    \nThe `finall_res` is my local test score. For LB, to get the same model, I didn't even use full train set to keep `train_df` set the same (I have set seed):\n\n    train_raw_df = pd.read_csv(train_path)\n    msk = np.random.rand(len(train_raw_df)) &lt; 0.95\n    train_df = train_raw_df[msk].copy().reset_index(drop=True)\n    del train_raw_df\n    test_df = pd.read_csv(test_path)\n\n    # train the model .... \n    # I get exactly the same loss in every epoch, so I should get the same model.\n    \n    sub = test_df[['qid']].copy()\n    sub['prediction'] = (test_preds &gt; search_result['threshold']).astype('int')\n    sub.to_csv(\"submission.csv\", index=False)\n\nHowever, the difference between local and LB still can be 0.001~0.015. It looks like that the LB score is always about 0.694-0.699, no matter what my local score is.  I also tried other seed and the variation of local test set score is about 0.001-0.003.\nShould I trust my local test set score?  Is my code correct?",
    "459594": "I got stuck with the problem about two weeks. So frustrated.",
    "459624": "I can't find any problems in your codes.Seems all should work.Maybe you can set a larger test set or just use kfold cv to evaluate your model."
  },
  "source": "meta"
}