--- license: apache-2.0 pipeline_tag: image-text-to-text tags: - PaddleOCR - OCR - Manga base_model: PaddlePaddle/PaddleOCR-VL language: - ja - multilingual library_name: PaddleOCR datasets: - hal-utokyo/Manga109-s --- # PaddleOCR-VL-For-Manga
## Model Description PaddleOCR-VL-For-Manga is an OCR model enhanced for Japanese manga text recognition. It is fine-tuned from [PaddleOCR-VL](https://huggingface.co/PaddlePaddle/PaddleOCR-VL) and achieves much higher accuracy on manga speech bubbles and stylized fonts. This model was fine-tuned on a combination of the [Manga109-s dataset](http://www.manga109.org/) and 1.5 million synthetic data samples. It showcases the potential of Supervised Fine-Tuning (SFT) to create highly accurate, domain-specific VLMs for OCR tasks from a powerful, general-purpose base like [PaddleOCR-VL](https://huggingface.co/PaddlePaddle/PaddleOCR-VL), which supports 109 languages. This project serves as a practical guide for developers looking to build their own custom OCR solutions. You can find the training code at the [Github Repository](https://github.com/jzhang533/PaddleOCR-VL-For-Manga), a step by step tutorial is avaiable [here](https://pfcc.blog/posts/paddleocr-vl-for-manga). ## Performance The model achieves a **70% full-sentence accuracy** on a test set of Manga109-s crops (representing a 10% split of the dataset). For comparison, the original PaddleOCR-VL on the same test dataset achieves 27% full sentence accuracy. Common errors involve discrepancies between visually similar characters that are often used interchangeably, such as: - `!?` vs. `!?` (Full-width vs. half-width punctuation) - `OK` vs. `ok` (Full-width vs. half-width letters) - `1205` vs. `1205` (Full-width vs. half-width numbers) - “人” (U+4EBA) vs. “⼈” (U+2F08) (Standard CJK Unified Ideograph vs. CJK Radical) The prevalence of these character types highlights a limitation of standard metrics like Character Error Rate (CER). These metrics may not fully capture the model's practical accuracy, as they penalize semantically equivalent variations that are common in stylized text. ## Examples | # | Image | Prediction | |---|---|---| | 1 |  | 心拍呼吸正常値