{
  "id": 123781,
  "title": "Dataset conflict",
  "url": "/competitions/tensorflow2-question-answering/discussion/123781",
  "author_name": "",
  "post_date": "2019-12-30T10:51:26.000326800Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In the dataset for this competition, <code>-1011141123527297803_long</code> is the first <code>example_id</code> in <code>sample_submission.csv</code>. This perfectly matches the test set in the 4G download of the dataset. However, in the dataset attached to a newly created notebook kernel, <code>simplified-nq-test.jsonl</code> has different examples. The closest <code>example_id</code> to the one above is <code>-1011141123527297792_long</code>.</p>\n\n<p>So the ids in  <code>simplified-nq-test.jsonl</code> do not match the ids in <code>sample_submission.csv</code> for a newly created notebook. This causes confusion and submission errors.</p>",
  "messages": [
    {
      "id": "706411",
      "postDate": "12/30/2019 10:51:26",
      "content": "<p>In the dataset for this competition, <code>-1011141123527297803_long</code> is the first <code>example_id</code> in <code>sample_submission.csv</code>. This perfectly matches the test set in the 4G download of the dataset. However, in the dataset attached to a newly created notebook kernel, <code>simplified-nq-test.jsonl</code> has different examples. The closest <code>example_id</code> to the one above is <code>-1011141123527297792_long</code>.</p>\n\n<p>So the ids in  <code>simplified-nq-test.jsonl</code> do not match the ids in <code>sample_submission.csv</code> for a newly created notebook. This causes confusion and submission errors.</p>",
      "rawMarkdown": "In the dataset for this competition, `-1011141123527297803_long` is the first `example_id` in `sample_submission.csv`. This perfectly matches the test set in the 4G download of the dataset. However, in the dataset attached to a newly created notebook kernel, `simplified-nq-test.jsonl` has different examples. The closest `example_id` to the one above is `-1011141123527297792_long`.\n\nSo the ids in  `simplified-nq-test.jsonl` do not match the ids in `sample_submission.csv` for a newly created notebook. This causes confusion and submission errors.",
      "votes": null
    },
    {
      "id": "706482",
      "postDate": "12/30/2019 12:58:47",
      "content": "<p>Hi, Ken! Good to see you here. </p>\n\n<p>This shall not be a problem, the eval script operates with dictionaries, and the order of example ids doesn’t matter much. </p>",
      "rawMarkdown": "Hi, Ken! Good to see you here. \n\nThis shall not be a problem, the eval script operates with dictionaries, and the order of example ids doesn’t matter much.",
      "votes": null
    },
    {
      "id": "706514",
      "postDate": "12/30/2019 13:50:04",
      "content": "<p>Hi Yury. Good to hear from you again. I hope you are well.</p>\n\n<p>I am not concerned about the order of the examples. I am concerned that the examples in <code>simplified-nq-test.jsonl</code> are different to the examples in <code>sample_submission.csv</code>.</p>",
      "rawMarkdown": "Hi Yury. Good to hear from you again. I hope you are well.\n\nI am not concerned about the order of the examples. I am concerned that the examples in `simplified-nq-test.jsonl` are different to the examples in `sample_submission.csv`.",
      "votes": null
    },
    {
      "id": "706515",
      "postDate": "12/30/2019 13:52:48",
      "content": "<p>They are actually the same. Looks like you read them somewhere with Pandas, and the int type is not read correctly. Specify string type then for example ids. </p>",
      "rawMarkdown": "They are actually the same. Looks like you read them somewhere with Pandas, and the int type is not read correctly. Specify string type then for example ids.",
      "votes": null
    },
    {
      "id": "706540",
      "postDate": "12/30/2019 14:33:27",
      "content": "<p>Hi Yury. I am indeed using Pandas, but even if I explicitly convert to string they are different as <a href=\"https://www.kaggle.com/kenkrige/dataset-conflict\">this</a> very  short kernel illustrates.</p>",
      "rawMarkdown": "Hi Yury. I am indeed using Pandas, but even if I explicitly convert to string they are different as [this](https://www.kaggle.com/kenkrige/dataset-conflict) very  short kernel illustrates.",
      "votes": null
    },
    {
      "id": "706625",
      "postDate": "12/30/2019 16:19:06",
      "content": "<p>My apologies. <a href=\"/kashnitsky\">@kashnitsky</a> you are absolutely right. I made a mistake with reading as int then converting to string, which gave a different example_id in the least significant digits.</p>",
      "rawMarkdown": "My apologies. @kashnitsky you are absolutely right. I made a mistake with reading as int then converting to string, which gave a different example_id in the least significant digits.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 706482,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "12/30/2019 12:58:47",
      "content": "<p>Hi, Ken! Good to see you here. </p>\n\n<p>This shall not be a problem, the eval script operates with dictionaries, and the order of example ids doesn’t matter much. </p>",
      "votes": null,
      "replies": [
        {
          "id": 706514,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "12/30/2019 13:50:04",
          "content": "<p>Hi Yury. Good to hear from you again. I hope you are well.</p>\n\n<p>I am not concerned about the order of the examples. I am concerned that the examples in <code>simplified-nq-test.jsonl</code> are different to the examples in <code>sample_submission.csv</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706515,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/30/2019 13:52:48",
          "content": "<p>They are actually the same. Looks like you read them somewhere with Pandas, and the int type is not read correctly. Specify string type then for example ids. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706540,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "12/30/2019 14:33:27",
          "content": "<p>Hi Yury. I am indeed using Pandas, but even if I explicitly convert to string they are different as <a href=\"https://www.kaggle.com/kenkrige/dataset-conflict\">this</a> very  short kernel illustrates.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706625,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "12/30/2019 16:19:06",
          "content": "<p>My apologies. <a href=\"/kashnitsky\">@kashnitsky</a> you are absolutely right. I made a mistake with reading as int then converting to string, which gave a different example_id in the least significant digits.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "706411": "In the dataset for this competition, `-1011141123527297803_long` is the first `example_id` in `sample_submission.csv`. This perfectly matches the test set in the 4G download of the dataset. However, in the dataset attached to a newly created notebook kernel, `simplified-nq-test.jsonl` has different examples. The closest `example_id` to the one above is `-1011141123527297792_long`.\n\nSo the ids in  `simplified-nq-test.jsonl` do not match the ids in `sample_submission.csv` for a newly created notebook. This causes confusion and submission errors.",
    "706482": "Hi, Ken! Good to see you here. \n\nThis shall not be a problem, the eval script operates with dictionaries, and the order of example ids doesn’t matter much.",
    "706514": "Hi Yury. Good to hear from you again. I hope you are well.\n\nI am not concerned about the order of the examples. I am concerned that the examples in `simplified-nq-test.jsonl` are different to the examples in `sample_submission.csv`.",
    "706515": "They are actually the same. Looks like you read them somewhere with Pandas, and the int type is not read correctly. Specify string type then for example ids.",
    "706540": "Hi Yury. I am indeed using Pandas, but even if I explicitly convert to string they are different as [this](https://www.kaggle.com/kenkrige/dataset-conflict) very  short kernel illustrates.",
    "706625": "My apologies. @kashnitsky you are absolutely right. I made a mistake with reading as int then converting to string, which gave a different example_id in the least significant digits."
  },
  "source": "meta"
}