{
  "id": 122420,
  "title": "Is the number of test images 12?",
  "url": "/competitions/bengaliai-cv19/discussion/122420",
  "author_name": "",
  "post_date": "2019-12-20T02:40:17.880197100Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p><code>df_test_img = pd.read_parquet('../test_image_data_0.parquet')</code>\n<code>df_test = pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')</code>\n<code>print('shape of df_test:', df_test.shape)</code>\n<code>print('shape of df_test_img_0:', df_test_img.shape)</code></p>\n\n<h3>Output</h3>\n\n<p><code>shape of df_test: (36, 3)</code>\n<code>shape of df_test_img_0: (3, 32333)</code></p>\n\n<p>Each image requires 3 rows for grapheme root, vowel diacritic, and consonant diacritic. So does that mean there are only 12 images in the test set? </p>\n\n<p>137x236 = 32332, so it looks like there are also 3 images in <code>test_image_data_0.parquet</code>? It looks like it's the same for the other 3 test parquet files as well? </p>\n\n<p>Thanks for the help.</p>",
  "messages": [
    {
      "id": "699062",
      "postDate": "12/20/2019 02:40:17",
      "content": "<p><code>df_test_img = pd.read_parquet('../test_image_data_0.parquet')</code>\n<code>df_test = pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')</code>\n<code>print('shape of df_test:', df_test.shape)</code>\n<code>print('shape of df_test_img_0:', df_test_img.shape)</code></p>\n\n<h3>Output</h3>\n\n<p><code>shape of df_test: (36, 3)</code>\n<code>shape of df_test_img_0: (3, 32333)</code></p>\n\n<p>Each image requires 3 rows for grapheme root, vowel diacritic, and consonant diacritic. So does that mean there are only 12 images in the test set? </p>\n\n<p>137x236 = 32332, so it looks like there are also 3 images in <code>test_image_data_0.parquet</code>? It looks like it's the same for the other 3 test parquet files as well? </p>\n\n<p>Thanks for the help.</p>",
      "rawMarkdown": "`df_test_img = pd.read_parquet('../test_image_data_0.parquet')`\n`df_test = pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')`\n`print('shape of df_test:', df_test.shape)`\n`print('shape of df_test_img_0:', df_test_img.shape)`\n\n###Output \n`shape of df_test: (36, 3)`\n`shape of df_test_img_0: (3, 32333)`\n\nEach image requires 3 rows for grapheme root, vowel diacritic, and consonant diacritic. So does that mean there are only 12 images in the test set? \n\n137x236 = 32332, so it looks like there are also 3 images in `test_image_data_0.parquet`? It looks like it's the same for the other 3 test parquet files as well? \n\nThanks for the help.",
      "votes": null
    },
    {
      "id": "699165",
      "postDate": "12/20/2019 05:42:28",
      "content": "<p>In the data description page it says</p>\n\n<blockquote>\n  <p>Only the first few rows/images in the test set and sample submission files can be downloaded. These samples provided so you can review the basic structure of the files and to ensure consistency between the publicly available set of file names and those your code will have access to while it is being rerun for scoring.</p>\n</blockquote>",
      "rawMarkdown": "In the data description page it says\n&gt; Only the first few rows/images in the test set and sample submission files can be downloaded. These samples provided so you can review the basic structure of the files and to ensure consistency between the publicly available set of file names and those your code will have access to while it is being rerun for scoring.",
      "votes": null
    },
    {
      "id": "699317",
      "postDate": "12/20/2019 10:01:14",
      "content": "<p>Hi Bo, you can access only 12 test samples when you try to access it through the kernels. When you commit and submit your prediction the server replaces the partial test sets with the full version and calculates the metric on that. I have tried to explain it in this <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">notebook.</a> Let me know if you have any confusion!</p>",
      "rawMarkdown": "Hi Bo, you can access only 12 test samples when you try to access it through the kernels. When you commit and submit your prediction the server replaces the partial test sets with the full version and calculates the metric on that. I have tried to explain it in this [notebook.](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set) Let me know if you have any confusion!",
      "votes": null
    },
    {
      "id": "699386",
      "postDate": "12/20/2019 11:28:40",
      "content": "<p>Will the public LB also be calculated on those 12 samples?</p>",
      "rawMarkdown": "Will the public LB also be calculated on those 12 samples?",
      "votes": null
    },
    {
      "id": "699622",
      "postDate": "12/20/2019 16:45:18",
      "content": "<p>The public LB uses something in the ballpark of 100,000 samples.</p>",
      "rawMarkdown": "The public LB uses something in the ballpark of 100,000 samples.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 699165,
      "author_name": "dhananjay3",
      "author_url": "",
      "post_date": "12/20/2019 05:42:28",
      "content": "<p>In the data description page it says</p>\n\n<blockquote>\n  <p>Only the first few rows/images in the test set and sample submission files can be downloaded. These samples provided so you can review the basic structure of the files and to ensure consistency between the publicly available set of file names and those your code will have access to while it is being rerun for scoring.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 699317,
      "author_name": "reasat",
      "author_url": "",
      "post_date": "12/20/2019 10:01:14",
      "content": "<p>Hi Bo, you can access only 12 test samples when you try to access it through the kernels. When you commit and submit your prediction the server replaces the partial test sets with the full version and calculates the metric on that. I have tried to explain it in this <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">notebook.</a> Let me know if you have any confusion!</p>",
      "votes": null,
      "replies": [
        {
          "id": 699386,
          "author_name": "vzaguskin",
          "author_url": "",
          "post_date": "12/20/2019 11:28:40",
          "content": "<p>Will the public LB also be calculated on those 12 samples?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699622,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "12/20/2019 16:45:18",
          "content": "<p>The public LB uses something in the ballpark of 100,000 samples.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "699062": "`df_test_img = pd.read_parquet('../test_image_data_0.parquet')`\n`df_test = pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')`\n`print('shape of df_test:', df_test.shape)`\n`print('shape of df_test_img_0:', df_test_img.shape)`\n\n###Output \n`shape of df_test: (36, 3)`\n`shape of df_test_img_0: (3, 32333)`\n\nEach image requires 3 rows for grapheme root, vowel diacritic, and consonant diacritic. So does that mean there are only 12 images in the test set? \n\n137x236 = 32332, so it looks like there are also 3 images in `test_image_data_0.parquet`? It looks like it's the same for the other 3 test parquet files as well? \n\nThanks for the help.",
    "699165": "In the data description page it says\n&gt; Only the first few rows/images in the test set and sample submission files can be downloaded. These samples provided so you can review the basic structure of the files and to ensure consistency between the publicly available set of file names and those your code will have access to while it is being rerun for scoring.",
    "699317": "Hi Bo, you can access only 12 test samples when you try to access it through the kernels. When you commit and submit your prediction the server replaces the partial test sets with the full version and calculates the metric on that. I have tried to explain it in this [notebook.](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set) Let me know if you have any confusion!",
    "699386": "Will the public LB also be calculated on those 12 samples?",
    "699622": "The public LB uses something in the ballpark of 100,000 samples."
  },
  "source": "meta"
}