{
  "id": 613016,
  "title": "Can we use pretrained multimodal llm like QwenVL-2.5 and DeepSeek-OCR ?",
  "url": "/competitions/physionet-ecg-image-digitization/discussion/613016",
  "author_name": "muaz",
  "post_date": "2025-10-23T16:45:18.225000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Are we allowed to use and finetune pretrained multimodal LLM that can read line plots then convert to x y values? Like QwenVL-2.5 and DeepSeek-OCR ?</p>",
  "messages": [
    {
      "id": 3305918,
      "postDate": "2025-10-23T18:09:19.933Z",
      "content": "<p>You may use any pretrained models you like (assuming the license allows you to do so), but please note that internet access is disabled on the competition notebooks, so you cannot call external APIs.  You could upload weights to your Kaggle notebook (or to a Kaggle dataset, which you can use as an input to your Kaggle notebook).</p>",
      "rawMarkdown": "You may use any pretrained models you like (assuming the license allows you to do so), but please note that internet access is disabled on the competition notebooks, so you cannot call external APIs.  You could upload weights to your Kaggle notebook (or to a Kaggle dataset, which you can use as an input to your Kaggle notebook).",
      "votes": 4
    },
    {
      "id": 3306318,
      "postDate": "2025-10-24T10:21:22.623Z",
      "content": "<p>i suggest start off with PhysioNet 2024 winner solution. it has a gitub repo.</p>",
      "rawMarkdown": "i suggest start off with PhysioNet 2024 winner solution. it has a gitub repo.",
      "votes": 1
    },
    {
      "id": 3307775,
      "postDate": "2025-10-27T18:47:20.930Z",
      "content": "<p>didn't run any test on the challenge data , but deepseek-ocr is quite promising since tested on line plots . Ig winner solution would be fine-tuning an vlm based OCR model on k TB of data points</p>",
      "rawMarkdown": "didn't run any test on the challenge data , but deepseek-ocr is quite promising since tested on line plots . Ig winner solution would be fine-tuning an vlm based OCR model on k TB of data points"
    },
    {
      "id": 3305894,
      "postDate": "2025-10-23T16:45:18.227Z",
      "content": "<p>Are we allowed to use and finetune pretrained multimodal LLM that can read line plots then convert to x y values? Like QwenVL-2.5 and DeepSeek-OCR ?</p>",
      "rawMarkdown": "Are we allowed to use and finetune pretrained multimodal LLM that can read line plots then convert to x y values? Like QwenVL-2.5 and DeepSeek-OCR ?"
    },
    {
      "id": 3307773,
      "postDate": "2025-10-27T18:45:03.013Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3305918,
      "author_name": "GDClifford",
      "author_url": "",
      "post_date": "2025-10-23T18:09:19.933000",
      "content": "<p>You may use any pretrained models you like (assuming the license allows you to do so), but please note that internet access is disabled on the competition notebooks, so you cannot call external APIs.  You could upload weights to your Kaggle notebook (or to a Kaggle dataset, which you can use as an input to your Kaggle notebook).</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3306318,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-10-24T10:21:22.623000",
      "content": "<p>i suggest start off with PhysioNet 2024 winner solution. it has a gitub repo.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3307775,
      "author_name": "Flouka",
      "author_url": "",
      "post_date": "2025-10-27T18:47:20.930000",
      "content": "<p>didn't run any test on the challenge data , but deepseek-ocr is quite promising since tested on line plots . Ig winner solution would be fine-tuning an vlm based OCR model on k TB of data points</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3307773,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-10-27T18:45:03.013000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3305918": "You may use any pretrained models you like (assuming the license allows you to do so), but please note that internet access is disabled on the competition notebooks, so you cannot call external APIs.  You could upload weights to your Kaggle notebook (or to a Kaggle dataset, which you can use as an input to your Kaggle notebook).",
    "3306318": "i suggest start off with PhysioNet 2024 winner solution. it has a gitub repo.",
    "3307775": "didn't run any test on the challenge data , but deepseek-ocr is quite promising since tested on line plots . Ig winner solution would be fine-tuning an vlm based OCR model on k TB of data points",
    "3305894": "Are we allowed to use and finetune pretrained multimodal LLM that can read line plots then convert to x y values? Like QwenVL-2.5 and DeepSeek-OCR ?",
    "3307773": ""
  }
}