{
  "id": 176392,
  "title": "Is it possible to make tf records for the private sets?",
  "url": "/competitions/landmark-recognition-2020/discussion/176392",
  "author_name": "",
  "post_date": "2020-08-21T15:11:53.723522900Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Sup, i know that this comp has a private train and test set. There is no problem in making a training tf record dataset and then build a data pipeline to train your models and then just use the trained model. My question is for the private dataset, how can we make a tf record dataset and use the pipeline that we build for training the model?. The only option for inference is to use the '/test' directory??</p>",
  "messages": [
    {
      "id": "980406",
      "postDate": "08/21/2020 15:11:53",
      "content": "<p>Sup, i know that this comp has a private train and test set. There is no problem in making a training tf record dataset and then build a data pipeline to train your models and then just use the trained model. My question is for the private dataset, how can we make a tf record dataset and use the pipeline that we build for training the model?. The only option for inference is to use the '/test' directory??</p>",
      "rawMarkdown": "Sup, i know that this comp has a private train and test set. There is no problem in making a training tf record dataset and then build a data pipeline to train your models and then just use the trained model. My question is for the private dataset, how can we make a tf record dataset and use the pipeline that we build for training the model?. The only option for inference is to use the '/test' directory??",
      "votes": null
    },
    {
      "id": "980707",
      "postDate": "08/21/2020 19:41:12",
      "content": "<p>The private rerun has no internet access so if you want to do inference <em>and</em> training on the 100k test set you're correct: you'll have to store them on disk. However the 12hr limit is not enough for training anyways, so I'd recommend only doing inference on the private sets.</p>",
      "rawMarkdown": "The private rerun has no internet access so if you want to do inference _and_ training on the 100k test set you're correct: you'll have to store them on disk. However the 12hr limit is not enough for training anyways, so I'd recommend only doing inference on the private sets.",
      "votes": null
    },
    {
      "id": "982035",
      "postDate": "08/23/2020 01:24:21",
      "content": "<p>Good question. Note that you can build a <code>tf.data.Dataset</code> pipeline from JPEGs. There is an example notebook <a href=\"https://www.kaggle.com/truonghoang/fork-of-siim-isic-melanoma-starter\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Good question. Note that you can build a `tf.data.Dataset` pipeline from JPEGs. There is an example notebook [here][1]\n\n[1]: https://www.kaggle.com/truonghoang/fork-of-siim-isic-melanoma-starter",
      "votes": null
    },
    {
      "id": "982642",
      "postDate": "08/23/2020 14:43:52",
      "content": "<p>Ty, but I believe that for inference we are obligated to use the private /test directory so transforming that data to other formats is not possible. For training you can do whatever you want, maybee build a TF record dataset and train your model, later build a jpeg pipeline to do Inference with the already trained model.</p>\n<p>I will try and build a TF record dataset with all the train images and share it using that example, happy kaggling😌</p>",
      "rawMarkdown": "Ty, but I believe that for inference we are obligated to use the private /test directory so transforming that data to other formats is not possible. For training you can do whatever you want, maybee build a TF record dataset and train your model, later build a jpeg pipeline to do Inference with the already trained model.\n\nI will try and build a TF record dataset with all the train images and share it using that example, happy kaggling😌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 980707,
      "author_name": "usmannkhan",
      "author_url": "",
      "post_date": "08/21/2020 19:41:12",
      "content": "<p>The private rerun has no internet access so if you want to do inference <em>and</em> training on the 100k test set you're correct: you'll have to store them on disk. However the 12hr limit is not enough for training anyways, so I'd recommend only doing inference on the private sets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 982035,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/23/2020 01:24:21",
      "content": "<p>Good question. Note that you can build a <code>tf.data.Dataset</code> pipeline from JPEGs. There is an example notebook <a href=\"https://www.kaggle.com/truonghoang/fork-of-siim-isic-melanoma-starter\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 982642,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "08/23/2020 14:43:52",
          "content": "<p>Ty, but I believe that for inference we are obligated to use the private /test directory so transforming that data to other formats is not possible. For training you can do whatever you want, maybee build a TF record dataset and train your model, later build a jpeg pipeline to do Inference with the already trained model.</p>\n<p>I will try and build a TF record dataset with all the train images and share it using that example, happy kaggling😌</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "980406": "Sup, i know that this comp has a private train and test set. There is no problem in making a training tf record dataset and then build a data pipeline to train your models and then just use the trained model. My question is for the private dataset, how can we make a tf record dataset and use the pipeline that we build for training the model?. The only option for inference is to use the '/test' directory??",
    "980707": "The private rerun has no internet access so if you want to do inference _and_ training on the 100k test set you're correct: you'll have to store them on disk. However the 12hr limit is not enough for training anyways, so I'd recommend only doing inference on the private sets.",
    "982035": "Good question. Note that you can build a `tf.data.Dataset` pipeline from JPEGs. There is an example notebook [here][1]\n\n[1]: https://www.kaggle.com/truonghoang/fork-of-siim-isic-melanoma-starter",
    "982642": "Ty, but I believe that for inference we are obligated to use the private /test directory so transforming that data to other formats is not possible. For training you can do whatever you want, maybee build a TF record dataset and train your model, later build a jpeg pipeline to do Inference with the already trained model.\n\nI will try and build a TF record dataset with all the train images and share it using that example, happy kaggling😌"
  },
  "source": "meta"
}