{
  "id": 139653,
  "title": "what test dataset to use?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/139653",
  "author_name": "",
  "post_date": "2020-03-29T18:59:29.453195900Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all,\nThe competition has two test dataset files . test.csv and test-processed-seqlen128.csv . there are many kernels (notebook) with test-en-df/test_en.csv test dataset. so, which one should we use?</p>",
  "messages": [
    {
      "id": "790659",
      "postDate": "03/29/2020 18:59:29",
      "content": "<p>Hi all,\nThe competition has two test dataset files . test.csv and test-processed-seqlen128.csv . there are many kernels (notebook) with test-en-df/test_en.csv test dataset. so, which one should we use?</p>",
      "rawMarkdown": "Hi all,\nThe competition has two test dataset files . test.csv and test-processed-seqlen128.csv . there are many kernels (notebook) with test-en-df/test_en.csv test dataset. so, which one should we use?",
      "votes": null
    },
    {
      "id": "790710",
      "postDate": "03/29/2020 19:46:45",
      "content": "<p>Well, test-processed-seqlen128.csv is the processed version of the test set specifically for BERT with a sequence length of 128. If you are planning to use sequence length other than 128 (which definitely you should use), then you can prepare the test set accordingly yourself through given test.csv. For other test set like test_en, these are the translated version others are using as it seems to naturally perform better due to given English only train data.</p>",
      "rawMarkdown": "Well, test-processed-seqlen128.csv is the processed version of the test set specifically for BERT with a sequence length of 128. If you are planning to use sequence length other than 128 (which definitely you should use), then you can prepare the test set accordingly yourself through given test.csv. For other test set like test_en, these are the translated version others are using as it seems to naturally perform better due to given English only train data.",
      "votes": null
    },
    {
      "id": "791044",
      "postDate": "03/30/2020 04:17:34",
      "content": "<p>hi Mayank,\ni am not able to find the  test-en-df/test_en.csv file in the data shared with this competition . but, i can see other notebooks are using this csv file . any idea on where this csv can be found.</p>",
      "rawMarkdown": "hi Mayank,\ni am not able to find the  test-en-df/test_en.csv file in the data shared with this competition . but, i can see other notebooks are using this csv file . any idea on where this csv can be found.",
      "votes": null
    },
    {
      "id": "791202",
      "postDate": "03/30/2020 07:22:34",
      "content": "<p>I guess the test-en_df are self generated csv? Correct me if I'm wrong.</p>",
      "rawMarkdown": "I guess the test-en_df are self generated csv? Correct me if I'm wrong.",
      "votes": null
    },
    {
      "id": "791361",
      "postDate": "03/30/2020 11:04:12",
      "content": "<p>While preparing notebooks/kernel, there is an option to add a custom dataset as well, you can use it to browse publically available or shared dataset or even upload any dataset by yourself privately to prepare your model complying with competition rules. test_en.csv is similarly the self-prepared dataset shared by others, <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138778\">see this</a> If you want to use test_en, you will have to add that dataset while preparing your notebook in Kaggle kernel.</p>",
      "rawMarkdown": "While preparing notebooks/kernel, there is an option to add a custom dataset as well, you can use it to browse publically available or shared dataset or even upload any dataset by yourself privately to prepare your model complying with competition rules. test_en.csv is similarly the self-prepared dataset shared by others, [see this](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138778) If you want to use test_en, you will have to add that dataset while preparing your notebook in Kaggle kernel.",
      "votes": null
    },
    {
      "id": "791362",
      "postDate": "03/30/2020 11:05:57",
      "content": "<p>Yes, right. do see above reply</p>",
      "rawMarkdown": "Yes, right. do see above reply",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 790710,
      "author_name": "mk9440",
      "author_url": "",
      "post_date": "03/29/2020 19:46:45",
      "content": "<p>Well, test-processed-seqlen128.csv is the processed version of the test set specifically for BERT with a sequence length of 128. If you are planning to use sequence length other than 128 (which definitely you should use), then you can prepare the test set accordingly yourself through given test.csv. For other test set like test_en, these are the translated version others are using as it seems to naturally perform better due to given English only train data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 791044,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "03/30/2020 04:17:34",
          "content": "<p>hi Mayank,\ni am not able to find the  test-en-df/test_en.csv file in the data shared with this competition . but, i can see other notebooks are using this csv file . any idea on where this csv can be found.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 791361,
          "author_name": "mk9440",
          "author_url": "",
          "post_date": "03/30/2020 11:04:12",
          "content": "<p>While preparing notebooks/kernel, there is an option to add a custom dataset as well, you can use it to browse publically available or shared dataset or even upload any dataset by yourself privately to prepare your model complying with competition rules. test_en.csv is similarly the self-prepared dataset shared by others, <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138778\">see this</a> If you want to use test_en, you will have to add that dataset while preparing your notebook in Kaggle kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 791202,
      "author_name": "sukyee",
      "author_url": "",
      "post_date": "03/30/2020 07:22:34",
      "content": "<p>I guess the test-en_df are self generated csv? Correct me if I'm wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 791362,
          "author_name": "mk9440",
          "author_url": "",
          "post_date": "03/30/2020 11:05:57",
          "content": "<p>Yes, right. do see above reply</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "790659": "Hi all,\nThe competition has two test dataset files . test.csv and test-processed-seqlen128.csv . there are many kernels (notebook) with test-en-df/test_en.csv test dataset. so, which one should we use?",
    "790710": "Well, test-processed-seqlen128.csv is the processed version of the test set specifically for BERT with a sequence length of 128. If you are planning to use sequence length other than 128 (which definitely you should use), then you can prepare the test set accordingly yourself through given test.csv. For other test set like test_en, these are the translated version others are using as it seems to naturally perform better due to given English only train data.",
    "791044": "hi Mayank,\ni am not able to find the  test-en-df/test_en.csv file in the data shared with this competition . but, i can see other notebooks are using this csv file . any idea on where this csv can be found.",
    "791202": "I guess the test-en_df are self generated csv? Correct me if I'm wrong.",
    "791361": "While preparing notebooks/kernel, there is an option to add a custom dataset as well, you can use it to browse publically available or shared dataset or even upload any dataset by yourself privately to prepare your model complying with competition rules. test_en.csv is similarly the self-prepared dataset shared by others, [see this](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/138778) If you want to use test_en, you will have to add that dataset while preparing your notebook in Kaggle kernel.",
    "791362": "Yes, right. do see above reply"
  },
  "source": "meta"
}